:

AI STRUGGLES WITH CODE RECREATION IN NEW BENCHMARK TEST

AI DESK1 MIN READ
FRI, JUN 26, 2026

■ AI-SUMMARIZED FROM 1 SOURCE ▸ TIMELINE

Epoch AI's MirrorCode benchmark reveals significant limitations in current AI models' ability to recreate programs without access to original code. Claude Opus 4.7 achieved the highest performance at 56 percent solve rate, yet all tested models failed on complex tasks.

The MirrorCode benchmark measures whether AI systems can reverse-engineer complete programs by analyzing their functionality alone. Claude Opus 4.7 led the field by successfully rebuilding a 16,000-line toolkit in 14 hours. However, the benchmark's most challenging tasks pushed AI to its limits. One task required continuous processing for 19 days and cost $2,600 to run—demonstrating both the computational expense and the fundamental difficulty of complex code recreation. The results highlight a clear gap in AI capabilities: while models excel at building code from scratch or making modifications to existing programs, they struggle significantly when tasked with complete reverse-engineering of software. The benchmark underscores that current AI systems, despite recent advances, remain limited in understanding and reproducing intricate program structures without direct access to source code. Epoch AI's findings suggest developers should approach AI-assisted code recreation with realistic expectations for now.

■ SOURCES

The Decoder

■ SUMMARY WRITTEN BY AI FROM THE LINKS ABOVE

■ MORE FROM THE AI DESK

OpenAI has begun rolling out GPT-6 Astra to Pro plan customers on its $100 and $200 monthly tiers. The release follows OpenAI's typical rollout pattern of prioritizing higher-tier subscribers before broader availability.

5H AGOAI Desk

Current AI systems cannot yet independently design circuit boards, according to research from EEBench. The gap between AI capabilities and the complexity of PCB design remains significant.

6H AGOAI Desk

Anthropic researchers have completed a formal mathematical proof of Fermat's Last Theorem, translating Andrew Wiles' decades-old proof into machine-verifiable code. The achievement marks a milestone in computational mathematics, ensuring the theorem's logical foundations are beyond dispute.

9H AGOIndustry Desk

OpenAI has released GPT-6 Astra, its most advanced model to date, marking a significant stride toward artificial general intelligence. The company has implemented new safety guardrails due to the model's powerful cybersecurity capabilities.

10H AGOAI Desk

■ SUBSCRIBE TO THE DAILY BRIEF

ONE EMAIL, 5 STORIES, 06:00 UTC. UNSUBSCRIBE ANYTIME.