A comparative study found Claude Code consumes nearly five times more tokens than OpenCode before even processing user prompts, raising efficiency concerns for developers managing API costs.
Researchers at TextTech conducted an empirical analysis comparing token usage between Claude Code and OpenCode by logging all requests sent to Anthropic's endpoint. The findings revealed a stark disparity: Claude Code sent 33,000 tokens before reading the initial prompt, while OpenCode required only 7,000 tokens for the same task.
The investigation began with anecdotal observations—developers noticed their usage meters climbing faster when switching to Claude Code due to unrelated Meridian issues. To verify this hunch, the team captured detailed usage data from both tools, measuring token consumption at the request level.
The 4.7x difference in pre-prompt token usage suggests Claude Code may include significantly more overhead—potentially larger system prompts, additional context, or initialization data. For users operating under token budgets or paying per-token rates, this efficiency gap could substantially impact operational costs.
The research underscores the importance of measuring actual token consumption rather than relying on assumptions about tool performance.
Uber's weekly AI agent requests have grown nearly tenfold since February, yet the company has held spending flat since April after exhausting its entire 2026 AI budget in Q1.
A recent paper shows artificial intelligence often diagnoses and treats patients better than human physicians. The findings are prompting difficult conversations within the medical community about the profession's evolving role.
The Relay Q, launching next year, represents the latest push to establish voice as the primary interface for human-computer interaction, challenging the keyboard's decades-long dominance.
An Anthropic researcher demonstrated automated systems that can identify and correct misaligned behaviors without compromising overall performance. The systems improved on all 10 tested benchmarks measuring specific problematic outputs.