4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems — 1 src1.3 Paper proposes evolving context parameterization for large language models — 1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems — 1 src1.3 Paper proposes evolving context parameterization for large language models — 1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions — 1 src
ElevenLabs has launched Music v2.5, a new version of its AI music generation model, available through its app and API with free and paid tiers. In a blind listening test involving almost 48,000 paired comparisons, the newer version was generally chosen over the previous one. The company states the model was trained exclusively on licensed music.
A startup called Abliteration has introduced a model designed without the usual safety guardrails, contrasting with larger labs that are considering stronger restrictions. The move raises questions about the company's rationale for stripping those protections.
Chinese lab AllSpark has published two open-source search agents, Iris-mini and Iris-pro, both built on Qwen base models. The company claims the pair lead benchmarks in their respective size classes. A notable side effect emerged during the agents' search training, according to the report.
In early testing on a new robotics benchmark called StationeryBench, GPT-6 Astra reportedly completed 7 of 100 tasks using dual-arm robots, while competing model MolmoAct2 finished none. A researcher characterized the results as a marked improvement in spatial reasoning. The figures come from preliminary benchmark runs rather than a full public release.
Google has introduced Gemini 3.8 Flash, a model it says narrows the performance gap with Anthropic's offerings while costing roughly half as much. At the same time, the company announced the Fairwind Program, a safety effort aimed at a limited set of partner companies that resembles initiatives from other AI developers.
Google Research has introduced TimesFM-3, a model designed to forecast time series by combining historical data with related signals and known upcoming events such as promotions or weather. It generates all future time points at once rather than predicting one step at a time, and has 330 million parameters.
A user report says the DeepSeek v4.1 Flash model can be loaded and run locally on a 2020 M1 Mac Mini with 16GB of unified memory, but generation is extremely slow at roughly 23 seconds per token. That pace makes the setup impractical for interactive use, though it shows the model can technically execute on older consumer hardware. The result highlights how limited RAM and memory bandwidth constrain local inference of large models.
DeepSeek has introduced V4.1-Flash, a multimodal model with 552 billion parameters, though only 16 billion are active per token. The release cuts KV cache memory requirements to a quarter of the previous generation. On the DeepSWE coding benchmark, it slightly outperforms Opus 5 and GPT-5.6 Sol.
DeepSeek has released V4.1-Flash, a multimodal model with 552 billion total parameters that activates only 16 billion per token. The company says it cuts KV-cache memory to roughly a quarter of what its previous model required, which matters for agent workloads. It slightly edges out Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark.
OpenAI's September enterprise presentation of GPT-6 Astra included safety metrics and administrative controls that were missing when the model first launched. The change followed an incident report and an unauthorized wiki, which prompted OpenAI to revise its published safety commitments. The gap between the two announcements suggests the company adjusted its disclosures within roughly a two-week window.
Suno has introduced its v6 generation of AI music models in three variants, developed in partnership with major labels Warner Music Group, BMG and Believe. The update adds multimodal input, letting users create tracks from text, audio or images, and edit sections of a song with text prompts. Suno is retiring its earlier model versions as v6 rolls out.
Suno has introduced a sixth generation of its music generation models, released in three variants. The company says the models were developed in collaboration with Warner Music Group, BMG and Believe, and it is discontinuing all earlier model versions. New capabilities include editing parts of an existing song through text prompts and generating tracks multimodally from text, audio and images.
OpenAI is rolling out ChatGPT Images 2.5, an image generation update built around two separate models: Flare, aimed at producing images more quickly, and Sunburst, aimed at more accurate editing. It remains unclear which users receive which model and on what schedule, though early testing gives some indication of who gains from the changes.
OpenAI has introduced ChatGPT Images 2.5, which brings two separate image models: Flare for faster generation and Sunburst for more precise edits. It remains unclear which model ChatGPT users will receive and at what point. Early testing by The Decoder suggests the improvements may not benefit all users equally.
Google DeepMind has announced WeatherNext 3, the newest version of its machine-learning system for global weather prediction. The company positions the update as a step up in forecast accuracy over earlier WeatherNext versions. The model extends DeepMind's weather AI line, which is used by businesses and organizations for planning and risk assessment.
WHY IT MATTERS ↘Iterative releases of operational ML forecasters like WeatherNext 3 show AI weather models moving from research demos to versioned commercial products, undercutting the cost and latency of supercomputer-based numerical prediction for energy, insurance, and logistics customers. It also tightens competitive pressure on public forecasting agencies and rivals such as NVIDIA to match accuracy at production scale.
Google DeepMind has expanded its Gemini model family with two new releases: the fast 3.8 Flash model and a 3.8 Flash Cyber variant. The Cyber edition is aimed at security-related workloads, complementing the general-purpose Flash model. The announcement was made via the company's official blog.
WHY IT MATTERS ↘A security-specialized model signals further verticalization of commercial LLMs, giving security teams a tuned option but also sharpening dual-use questions around offensive-vs-defensive capability. The 3.8 Flash release meanwhile sustains price and latency pressure in the fast-inference tier, where Flash-class models are the main competitive battleground against OpenAI and Anthropic.
Google DeepMind announced a new Gemini capability that lets the model analyze video content in an agentic, multi-step way rather than only answering single-pass questions about clips. The company says this allows the system to follow events over time, connect what it sees to tasks, and take further actions based on video input. Details on availability, pricing, and supported regions were not fully specified in the report.
WHY IT MATTERS ↘This shifts competition from benchmark video QA to deployable video agents that can chain perception with tools and actions, making continuous video analysis a practical automation layer for monitoring, editing, and interactive assistants. It also raises governance and cost questions, since always-on video ingestion and downstream actions increase privacy, liability, and compute demands that buyers will need to audit before production use.
Google DeepMind announced Gemini Omni 1.1 Flash, an updated version of its Omni Flash model. According to the company's blog, the release focuses on giving developers more control when building with the model. Further technical details and availability were not specified in the report.
WHY IT MATTERS ↘Added build controls address a common friction point in deploying foundation models: the need to tailor behavior without expensive fine-tuning. For the industry, this raises the bar for developer-friendly customization, pressuring competitors to match Google's flexibility and potentially lowering barriers to regulated or specialized applications.