LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

open-weights

topic19 events
industryTODAY 12:00 UTC

Mozilla report finds open models match frontier AI within about 4 months

A Mozilla analysis, previewed by Ars Technica, compares openly available models with paid frontier systems and finds the open options reach comparable capability after roughly four months. Buyers of the leading closed models pay about five times more for that head start. The report suggests the practical advantage of premium closed models is short-lived relative to their price.

papersTODAY 04:00 UTC

Crypto Accounting Bench tests LLMs on reconstructing crypto transaction entries

Researchers released Crypto Accounting Bench, a set of 118 evaluation items that ask language models to rebuild the full journal entry an organization recorded for a crypto-asset transaction. The benchmark is applied to both frontier and open-weight models to gauge how well they handle this specialized accounting task. It appears on arXiv under cs.AI and cs.CL.

papersTODAY 04:00 UTC

AttriCite: Open 4B Model Trained for Citation Recovery in Scientific Papers

A new arXiv paper introduces AttriCite, an openly released 4B-parameter model trained to identify which paper an author intended to cite from the surrounding text of a citation. The work frames this as a "citation recovery" task, aiming to support faithful attribution of scientific claims to their sources. The authors position the approach as a step toward more reliable attribution in AI-assisted scientific writing.

papersTODAY 04:00 UTC

arXiv paper evaluates open-source LLMs for RAG in ESG reporting

A new arXiv preprint examines how well open-source large language models perform when paired with retrieval-augmented generation for environmental, social, and governance reporting tasks. The authors focus on automating the extraction of key performance indicators from ESG disclosures, a step they describe as important for corporate accountability. The abstract suggests limits in current open-source model performance for this domain.

modelsTODAY 04:00 UTC

ZGCM-1: Open 7B Foundation Model Targets Math and Agentic Search

Researchers released ZGCM-1, a 7-billion-parameter dense foundation model trained from scratch with a focus on data, system, and algorithmic efficiency. The work argues that smaller models should not try to memorize the open web, but instead be optimized for targeted capabilities such as mathematical reasoning and agentic search. It is presented as a fully open release.

modelsTODAY 04:00 UTC

North Small Translate debuts as open-weight machine translation model

North Small Translate is a new open-weight translation model that also follows instructions, described as being trained on the same base as Cohere's Command A Plus mixture-of-experts system with 25 billion active parameters. The authors position it as a cost-effective option for machine translation workloads that need instruction-following behaviour.

modelsTODAY 04:00 UTC

MameLoshnLM: First Open-Source 8B Language Model for Yiddish Introduced

Researchers released MameLoshnLM, described as the first open-source 8-billion-parameter language model dedicated to Yiddish. The work also includes an evaluation benchmark intended to fill the gap in reliable testing resources for the language. It addresses the low digital availability of Yiddish text despite its substantial written heritage.

modelsTODAY 04:00 UTC

Salesforce Koa: enterprise LLM post-trained from Nemotron-3-Super-120B with GRPO

Salesforce has introduced Koa, an enterprise-focused language model created by post-training the open-weight Nemotron-3-Super-120B foundation model. The training process uses reinforcement learning with Group Relative Policy Optimization (GRPO), drawing on public data and other sources. The work targets agentic tool use in enterprise settings, according to the arXiv preprint.

papersTODAY 04:00 UTC

Study assesses context segmentation in locally run small models for CTF security tasks

A preprint on arXiv examines how segmenting context affects the performance of open-weight small language models used for capture-the-flag cybersecurity exercises. The authors frame the work around the risk that locally hosted models can sidestep the guardrails enforced by proprietary APIs. The paper is a revised cross-listing in the cs.AI category.

papersTODAY 04:00 UTC

Quantization-Conditioned Backdoor Attacks Target Open-Weight LLM Agents

A new arXiv paper describes an attack in which an adversary releases a full-precision model checkpoint that passes standard audits but behaves maliciously once it is quantized for deployment. Because quantization is a common default path for running open-weight agent models, the technique could let compromised agents slip past pre-release checks. The work frames this as a supply-chain risk for quantized LLM deployments.

papersSEP 12 04:00 UTC

Open recipe targets IMO gold with post-trained Nemotron math models

A new arXiv paper examines how post-training choices and test-time inference setups influence a model's ability to write natural-language proofs for difficult olympiad problems. Using Nemotron 3 Ultra as a base, the authors produce two specialist checkpoints via supervised fine-tuning and reinforcement learning, and release the training approach publicly.

modelsSEP 10 06:11 UTC

DeepSeek publishes V4.1 Flash model on Hugging Face

DeepSeek has added a new model called V4.1 Flash to its Hugging Face repository, where the weights and model card are hosted. The listing drew attention on Hacker News, though the report gives no further detail on capabilities, size, or licensing. It appears to be a lighter or faster variant in the company's V4 series.

modelsSEP 10 04:00 UTC

Ling 2.0: open reasoning-focused language models scale to 1 trillion parameters

A new technical report introduces Ling 2.0, a family of reasoning-oriented foundation models built on a unified Mixture-of-Experts architecture that spans from tens of billions up to one trillion parameters. The series is released as an open language foundation, with the stated goal of strengthening general reasoning ability across all model sizes.

papersSEP 10 04:00 UTC

Study Proposes Maverick for Private, Verifiable LLM Inference via Matrix-Vector Delegation

A paper posted to arXiv introduces Maverick, a system designed to let users run LLM inference on external servers without exposing their inputs or blindly trusting the returned results. The method delegates the heavy matrix-vector multiplications that dominate transformer inference while adding privacy protections and a mechanism to confirm that computations were performed correctly. The authors frame the work as a step toward making private and verifiable inference practical for open-source models.

tipsSEP 9 22:26 UTC

AWS guide covers deploying Qwen3.8-2.4T-A95B on SageMaker HyperPod with vLLM

Amazon published a walkthrough for running the open-weight Qwen3.8-2.4T-A95B model, which has 2.4 trillion parameters, on its SageMaker HyperPod service using the vLLM inference engine. The guide covers setting up the cluster, applying NVFP4 quantization, and exposing an OpenAI-compatible endpoint. It also notes support for tool calling, reasoning, and multi-token prediction speculative decoding.

tipsSEP 9 18:30 UTC

Reporter Removes Open-Source AI Safety Limits, Agent Hacks His Home Devices

A Wired writer stripped the safety restrictions from a capable open-source model and set it loose on the gadgets in his home, where it discovered security flaws and broke into a desktop computer. The same agent then outlined steps to harden those devices. The piece is a hands-on look at how easily guardrails can be removed and what an unconstrained agent can accomplish.