OpenAI released GPT-Live-1, a full-duplex speech API enabling simultaneous talking and listening in AI applications. The model achieved 80.1% interactivity scores, nearly doubling its predecessor's 45.4%.
OpenAI has made GPT-Live-1 available as a developer API, introducing full-duplex speech capabilities to its product lineup. The model handles bidirectional audio streams, allowing applications to process speech input and generate responses simultaneously rather than waiting for users to finish speaking.
Performance Gains
The new model significantly improves on earlier iterations. Interactivity test scores climbed to 80.1%, a substantial leap from the previous generation's 45.4%. This improvement suggests more natural conversational flow and reduced latency in back-and-forth exchanges.
Pricing and Access
OpenAI priced GPT-Live-1 at $0.05 per minute of API usage. The rate positions it as a premium offering compared to typical text-based API services, reflecting the computational demands of real-time bidirectional audio processing.
Implications for Developers
The API opens new possibilities for voice-driven applications. Developers can build customer service bots, voice assistants, and interactive applications that more closely mimic human conversation patterns. The full-duplex capability reduces the stop-and-start interactions common in earlier speech models, potentially improving user experience in voice-first applications.
The release follows industry momentum toward multimodal and real-time AI capabilities. Competitors have similarly pushed to enhance speech processing, making this an incremental but meaningful step in AI conversation technology.
Considerations
While the performance improvements are notable, developers will need to evaluate whether the cost structure fits their use cases. At five cents per minute, extended voice sessions accumulate expenses quickly compared to text-based alternatives.
Nvidia co-founder and CEO Jensen Huang identified cybersecurity as the next major market for artificial intelligence, predicting the technology will fundamentally reshape how computer systems are defended.
Cognition has released SWE-2, a new AI model designed to compete with Anthropic's Claude 5.1 and OpenAI's GPT-Astra. The model targets software engineering tasks and represents Cognition's latest push in the generative AI space.
India's audio entertainment platform Pocket FM has doubled its revenue run rate to $500 million annually, with AI now powering 93% of its content library. The shift has reduced production costs by approximately 80 times compared to traditional methods.