5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve — 1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research — 1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve — 1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research — 1 src
Salesforce introduced Koa, a reasoning model developed with Nvidia and built on Nvidia's open-weight Nemotron foundation. The model is tuned for enterprise functions such as sales, marketing and customer support rather than general-purpose use. The companies position it as an alternative to models from the major AI labs.
A new arXiv paper empirically tests whether a general-purpose shell interface outperforms purpose-built tools when AI agents handle enterprise workflows. The authors note that shell-based agents perform well on coding tasks, but enterprise work also requires moving across applications and services and coordinating multiple steps. The study examines these trade-offs to identify which tool interface design suits digital worker agents.
A new arXiv preprint examines how enterprise AI agents are typically given static credentials at deployment that mirror the full set of permissions an employee role could hold. The authors argue this approach, inherited from role-based access control, grants agents far more access than any single task requires. They evaluate an alternative architecture that scopes an agent's permissions to the specific task it is performing.
OpenAI's September enterprise presentation of GPT-6 Astra included safety metrics and administrative controls that were missing when the model first launched. The change followed an incident report and an unauthorized wiki, which prompted OpenAI to revise its published safety commitments. The gap between the two announcements suggests the company adjusted its disclosures within roughly a two-week window.
The authors argue that LLM agents operating on enterprise systems of record are difficult to evaluate because production customer data cannot be used for testing and no existing substitute offers reliable ground truth. Their new benchmark addresses this by generating a synthetic enterprise environment paired with exact ground-truth labels, enabling systematic scoring of agents that use enterprise tools.
A cross-listed arXiv paper introduces IBIB, a protocol for evaluating AI systems as they are actually deployed inside enterprises rather than as bare model checkpoints. The authors argue that real-world capability emerges from the combination of weights, serving configuration, precision, output contract, and harness, so scoring an advertised model name alone is a measurement error. After auditing 18 existing benchmarks and finding that every one grades model identifiers, the paper proposes routing-based measurement instead.
A new research paper introduces EVA-Bench, a benchmark for assessing voice agents across the entire interaction pipeline. It combines simulated conversations that mimic real usage with metrics tailored to voice-specific behaviors, filling a gap left by earlier evaluation tools that handled these aspects separately. The work responds to the growing deployment of voice agents in enterprise applications.
OpenAI has introduced GPT-6 Astra, which it positions as its strongest model for workplace and enterprise applications. The system emphasizes multi-step reasoning, the ability to operate a computer on a user's behalf, and improved writing and design capabilities. The release targets business users rather than the general consumer market.
WHY IT MATTERS ↘Bundling computer use into a flagship enterprise model shifts AI deployments from assisted drafting to autonomous task execution, forcing companies to confront access controls, auditability, and liability gaps that most current governance frameworks don't cover. It also intensifies the enterprise agentic race against Anthropic and Google, where purchasing decisions will increasingly hinge on reliability and safety controls rather than benchmark scores.
OpenAI says advances in AI capability combined with falling costs let individuals and companies take on a broader range of tasks than before. The company positions cheaper, stronger models as a way for organizations to grow without proportionally higher expenses.
WHY IT MATTERS ↘Falling per-token costs paired with rising capability lower the break-even point for automating mid-complexity work, making AI economically viable for smaller firms and lower-margin functions. This shifts vendor competition toward price-performance rather than raw benchmark leads, pressuring model providers' margins and prompting buyers to revisit build-versus-buy economics.