[DEV]■ STORY TIMELINE
20B MoE MODEL RUNS AT 120 TOK/S ON IPHONE
Researchers have demonstrated Maple-Preview, a 20 billion parameter mixture-of-experts model, executing at 120 tokens per second on an iPhone. The achievement signals substantial progress in running large language models on mobile devices.
Hacker News+0m
Article URL: https://deepgrove.ai/maple-preview Comments URL: https://news.ycombinator.com/item?id=49173984 Points: 120…