5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve — 1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research — 1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve — 1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research — 1 src
A preprint describes a methodological comparison of two coding agents, Anthropic's Claude Code and OpenAI's Codex, each autonomously running the same matched-filter pipeline on simulated Einstein Telescope data. The authors frame it as an early look at how agentic AI performs on a real gravitational-wave analysis workflow rather than as a model benchmark.
OpenAI developer Eric Provencher warns that verbose skill descriptions, broad mandatory-reading requirements, and inflexible approval gates can slow down GPT-6 Astra when used in Codex. He recommends that developers trim instructions to fit the specific task and clearly state what a completed job looks like, since stronger models require less detailed guidance.
OpenAI has made its Agents API available as a public beta, giving outside developers access to the same infrastructure that powers Codex and ChatGPT agents. The API supports cloud-based agents that can operate autonomously for extended periods, run code, and pass tasks to sub-agents. Pricing is based solely on token consumption, with Cloudflare, Vercel, and Oracle providing optional sandbox environments.
WHY IT MATTERS ↘By exposing the same agent runtime behind Codex and ChatGPT, OpenAI turns agent orchestration into a metered API, pushing competitors to differentiate on reliability, sandboxing, and cost control rather than model quality alone. The token-only pricing also makes long-running autonomous and sub-agent workflows financially variable, raising governance and budget concerns for teams deploying them in production.
A developer shared a free, self-hostable application he uses to run his own business, describing it as combining Claude Code, Cowork, and cloud agent sessions in one install. The system is multi-tenant by design, allowing multiple people to collaborate. Its distinguishing feature is organizing Claude Code and Codex agents into company departments.
OpenAI spotlighted a case study in which an MIT researcher pairs GPT-5.6 Sol with its Codex agent to carry out quantum computing work with minimal human involvement. The system runs lab procedures on its own, interprets the measurement data it collects, and handles qubit recalibration. The writeup serves as an example of agentic AI being applied to hands-on scientific research.
WHY IT MATTERS ↘Agentic AI moving from software tasks into physical lab work shifts the value proposition toward automating scientific labor itself, where validation requirements and error costs are far higher than in code generation. It also signals frontier labs competing for research-automation workloads, forcing labs to define oversight and verification protocols for experiments run with minimal human involvement.
Australian law firm Gilbert + Tobin has detailed how it scaled OpenAI's ChatGPT Enterprise and Codex across its organisation. The firm credits executive sponsorship, formal governance controls and clear human oversight for the deployment. The case study highlights how professional services firms are adopting generative AI while keeping accountability with people.
WHY IT MATTERS ↘It gives a concrete template for regulated professional-services firms: executive sponsorship plus formal governance and human sign-off can unlock firm-wide generative AI deployment where pure productivity arguments stall. That raises the bar for vendors competing on enterprise controls, and shifts the differentiator from model capability to auditable accountability.
Japanese company Polimill is developing AI infrastructure intended for public-sector use, drawing on OpenAI's GPT models and Codex. The system is designed to let local government staff search and apply administrative knowledge, while also speeding up its own development. The effort points to growing adoption of commercial AI models in government service delivery.
WHY IT MATTERS ↘It shows the frontier labs' public-sector strategy is now being executed through small local integrators rather than direct government contracts, which spreads model dependence into municipal workflows that are hard to migrate once entrenched. Using Codex to build the system also suggests the same vendors selling AI deployment are increasingly relying on it internally, compressing delivery timelines and lowering the barrier for small firms to win government work.