Human Archive, founded by UC Berkeley and Stanford researchers, is paying gig workers in India to collect physical training data for AI and robotics labs. Workers wear camera-equipped caps and sensors to capture real-world movement data.
The startup addresses a critical bottleneck in robotics development: the scarcity of high-quality, real-world training data. AI and robotics laboratories need vast amounts of physical interaction data to train machine learning models that can perform complex tasks.
Human Archive's model leverages India's large gig economy workforce to collect this data at scale. Participating workers wear specialized equipment that records their movements and interactions with physical environments, generating datasets that robotics companies need.
The approach offers potential benefits for multiple parties: gig workers gain additional income opportunities, robotics labs access affordable training data, and Human Archive positions itself as a data supply chain intermediary. The strategy reflects a broader trend of using human labor in emerging markets to support AI development globally.
The startup operates in a space where demand for training data outpaces supply, particularly for embodied AI systems that must understand physical space and human interaction patterns.
Uber's weekly AI agent requests have grown nearly tenfold since February, yet the company has held spending flat since April after exhausting its entire 2026 AI budget in Q1.
A recent paper shows artificial intelligence often diagnoses and treats patients better than human physicians. The findings are prompting difficult conversations within the medical community about the profession's evolving role.
The Relay Q, launching next year, represents the latest push to establish voice as the primary interface for human-computer interaction, challenging the keyboard's decades-long dominance.
An Anthropic researcher demonstrated automated systems that can identify and correct misaligned behaviors without compromising overall performance. The systems improved on all 10 tested benchmarks measuring specific problematic outputs.