:
[DEV]■ STORY TIMELINE

20B MoE MODEL RUNS AT 120 TOK/S ON IPHONE

Researchers have demonstrated Maple-Preview, a 20 billion parameter mixture-of-experts model, executing at 120 tokens per second on an iPhone. The achievement signals substantial progress in running large language models on mobile devices.

1 SOURCEFIRST SEEN AUG 4, 07:44 PM► READ THE ARTICLE
Hacker News+0m

Article URL: https://deepgrove.ai/maple-preview Comments URL: https://news.ycombinator.com/item?id=49173984 Points: 120…