4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems — 1 src1.3 Paper proposes evolving context parameterization for large language models — 1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems — 1 src1.3 Paper proposes evolving context parameterization for large language models — 1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src
Hugging Face urges AI safety systems to refuse harmful subsets, not entire topics
A Hugging Face blog post examines how moderation classifiers and language models often block benign requests simply because they touch a flagged subject. The authors argue that refusal policies should be scoped to the genuinely harmful portion of a topic, and question whose definition of safety gets encoded into today's systems. The piece advocates building more granular safety taxonomies that cut down over-refusal without weakening protection.
WHY IT MATTERS ↘Over-refusal quietly erodes product utility and user trust while inflating eval and support costs, so teams tuning moderation stacks face a concrete trade-off between safety coverage and usability rather than a simple safety-maximizing default. The governance angle — whose definition of harm gets encoded into classifiers — also pressures vendors to document and defend their safety taxonomies as enterprises and regulators scrutinize automated content decisions.