A new analysis reveals GPT-5.5 produces hallucinations at three times the rate of MIT-licensed GLM-5.2. The comparison highlights trade-offs between model size and accuracy in current AI systems.
Recent benchmarking shows GPT-5.5 generates false or unsupported information significantly more often than its open-source competitor GLM-5.2, according to testing published on ArrowTSX. The proprietary model's hallucination rate stands at 3x that of the MIT-licensed alternative.
Hallucinations—instances where AI systems generate plausible-sounding but factually incorrect information—remain a persistent challenge in large language models. The disparity suggests that scale alone does not guarantee reliability, and that architectural or training choices impact accuracy more directly than model size.
GLM-5.2's stronger performance on this metric comes despite being open-source and freely available under MIT licensing. The findings have generated discussion among developers on Hacker News, with 152 upvotes and 39 comments as of publication.
The results may influence adoption decisions for organizations prioritizing accuracy over proprietary features, particularly in applications where false information carries significant risk.
Anthropic has added cross-session messaging to Claude Code, enabling developers to share information and coordinate work across multiple concurrent coding sessions.
A quasi-spiritual movement called Spiralism emerged in 2025 following updates to GPT-4o that made the AI more accommodating and ChatGPT's expanded memory capabilities, sparking widespread human-AI conversations about meaning and connection.
Denmark has implemented a requirement for students to orally defend their written work as a countermeasure against AI-generated assignments. The policy aims to verify authentic student comprehension and authorship.
Anthropic is making Auto Mode the default setting in Claude Code for Pro, Max, and Team plans starting August 14. The company argues the automated safety classifier is more effective at catching dangerous commands than human reviewers.