Google has expanded its Gemini API File Search to support multimodal capabilities, enabling developers to search across text, images, and other file types simultaneously.
The enhancement allows developers to build retrieval-augmented generation (RAG) systems that process mixed content formats within a single search operation. Previously, File Search was limited to text-based queries and documents.
Multimodal support enables more sophisticated use cases, such as searching through documents containing both written content and visual elements like diagrams, charts, or photographs. Developers can now index diverse file types and retrieve relevant results across all formats.
The feature integrates directly into the Gemini API, maintaining Google's minimalist approach to developer tools. This addition addresses growing demand for AI systems that can reason across different data types, particularly in enterprise document analysis and knowledge management applications.
The update reflects broader industry movement toward multimodal AI capabilities, with competitors also expanding their search and retrieval features. Google's implementation aims to simplify how developers incorporate complex document understanding into their applications.
TIME magazine is detecting AI crawlers and serving them a different version of its website that includes advertisements. The practice raises questions about how publishers are monetizing bot traffic.
The American Federation of Teachers is partnering with major tech companies to fund AI literacy training for educators, addressing growing concerns about artificial intelligence in classrooms.
Hark has previewed a new browser use agent designed to automate web tasks. The company claims its solution outperforms competitors on both speed and cost.
Open-source models now outperform OpenAI's frontier GPT-5.6 Sol on retrieval tasks while costing a fraction of the price. Neon and Castform demonstrated the efficiency gap in a new benchmark.