AI STRUGGLES WITH CODE RECREATION IN NEW BENCHMARK TEST
■ AI-SUMMARIZED FROM 1 SOURCE ▸ TIMELINE
Epoch AI's MirrorCode benchmark reveals significant limitations in current AI models' ability to recreate programs without access to original code. Claude Opus 4.7 achieved the highest performance at 56 percent solve rate, yet all tested models failed on complex tasks.
■ MORE FROM THE AI DESK
OpenAI has begun rolling out GPT-6 Astra to Pro plan customers on its $100 and $200 monthly tiers. The release follows OpenAI's typical rollout pattern of prioritizing higher-tier subscribers before broader availability.
Current AI systems cannot yet independently design circuit boards, according to research from EEBench. The gap between AI capabilities and the complexity of PCB design remains significant.
Anthropic researchers have completed a formal mathematical proof of Fermat's Last Theorem, translating Andrew Wiles' decades-old proof into machine-verifiable code. The achievement marks a milestone in computational mathematics, ensuring the theorem's logical foundations are beyond dispute.
OpenAI has released GPT-6 Astra, its most advanced model to date, marking a significant stride toward artificial general intelligence. The company has implemented new safety guardrails due to the model's powerful cybersecurity capabilities.