5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve — 1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research — 1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve — 1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research — 1 src
Search APIs Evaluated as Decision Surfaces for Tool-Using AI Agents
A new arXiv paper examines how the ranked snippets, URLs, and metadata returned by search APIs shape the choices made by tool-using AI agents, such as whether to answer, search again, or open a page. The authors evaluate these interfaces as decision surfaces using a fixed set of 100 questions drawn from the 254-question SealQA-Hard benchmark. The work suggests that agents can reach similar accuracy while relying on differing amounts or quality of supporting evidence.