:

20B MoE MODEL RUNS AT 120 TOK/S ON IPHONE

INDUSTRY DESK1 MIN READ
WED, AUG 5, 2026

■ AI-SUMMARIZED FROM 1 SOURCE ▸ TIMELINE

Researchers have demonstrated Maple-Preview, a 20 billion parameter mixture-of-experts model, executing at 120 tokens per second on an iPhone. The achievement signals substantial progress in running large language models on mobile devices.

Maple-Preview represents a significant advancement in on-device AI inference. The ternary quantized model achieves competitive performance while maintaining practical speed on consumer smartphones, eliminating the need for cloud processing for certain tasks. The 120 tokens-per-second throughput enables real-time interaction with a model that would traditionally require dedicated server infrastructure. This performance level makes the system viable for privacy-critical applications where keeping data local matters. The mixture-of-experts architecture allows selective activation of model parameters, reducing computational overhead compared to dense models of equivalent size. Ternary quantization further compresses the model by limiting weights to three discrete values. The project, shared on Hacker News, generated 120 points and 34 comments, indicating strong community interest in mobile AI capabilities. The development reflects growing momentum toward practical edge deployment of large language models.

■ SOURCES

Hacker News

■ SUMMARY WRITTEN BY AI FROM THE LINKS ABOVE

■ MORE FROM THE DEV DESK

The core team managing Nixpkgs, the package repository for the NixOS Linux distribution, has officially disbanded. The announcement was made on the NixOS Discourse forum.

JUST NOWIndustry Desk

Oracle has implemented a policy prohibiting AI-generated code contributions to OpenJDK, the open-source Java platform. The ban applies to code created by large language models and similar AI systems.

9H AGOAI Desk

A website operator discovered that nearly all of their traffic came from automated bots rather than human visitors, raising questions about how internet metrics are measured and reported.

11H AGOIndustry Desk

Oracle is reducing its Always Free ARM compute tier effective August 18, cutting capacity from higher previous limits to 2 OCPUs and 12GB of memory. The change affects users relying on Oracle's free tier for development and testing.

16H AGOIndustry Desk

■ SUBSCRIBE TO THE DAILY BRIEF

ONE EMAIL, 5 STORIES, 06:00 UTC. UNSUBSCRIBE ANYTIME.