OpenAI is previewing Ultrafast, a new API tier powered by Cerebras that accelerates its most capable GPT-5.6 Sol model up to 14 times faster, generating 750 output tokens per second.
OpenAI has announced Ultrafast, an API tier designed to dramatically accelerate inference speeds for GPT-5.6 Sol, the company's most advanced model. The new tier leverages infrastructure from Cerebras, a specialist in AI chip design and deployment.
The performance gains are substantial. Ultrafast delivers up to 14× faster speeds compared to standard OpenAI API tiers, while generating up to 750 output tokens per second. This throughput represents a significant leap for real-time applications that demand rapid model responses.
The preview signals OpenAI's push to address latency concerns that have limited enterprise adoption of its most capable models. Real-time AI applications—from customer service to code generation—often require faster response times than standard inference provides.
Cerebras' involvement indicates OpenAI's broader infrastructure strategy. The partnership leverages Cerebras' specialized hardware and software stack optimized for transformer-based models, complementing OpenAI's existing computational resources.
The Ultrafast tier will likely carry premium pricing, positioning it as a specialized offering rather than a replacement for existing API tiers. OpenAI has historically tiered its services by capability and speed, allowing customers to optimize cost versus performance based on their needs.
This development reflects intensifying competition in the inference optimization space. Other AI providers and specialized hardware companies are similarly pursuing faster, more efficient model deployment. Faster inference directly impacts user experience and operational costs for AI applications at scale.
The preview phase suggests the offering remains under evaluation before broader availability. OpenAI typically gathers feedback and optimizes offerings during preview periods before general release.
For developers and enterprises relying on GPT-5.6 Sol for latency-sensitive applications, Ultrafast could enable new use cases previously constrained by inference speed.
Machine-generated songs are climbing the charts despite widespread musician outcry, forcing record labels to adapt to a landscape where hits can be produced instantly at minimal cost.
Multiple tools claiming to remove AI watermarks from Claude-generated text have surfaced since Anthropic began watermarking its output. None of the tools can prove their effectiveness, as Anthropic has not released a public detector.
Google has released a new version of its Gemini Flash AI model, while remaining silent on the launch timeline for its more powerful Gemini 3.5 Pro, which continues to face delays.
Anthropic has widened its lead over OpenAI in enterprise AI spending, capturing 43.5% market share according to Ramp's July AI Index. Fable 5 remains a marginal player at just 6% of business token purchases, hampered by high costs.