LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

instruction-following

topic6 events
modelsTODAY 04:00 UTC

North Small Translate debuts as open-weight machine translation model

North Small Translate is a new open-weight translation model that also follows instructions, described as being trained on the same base as Cohere's Command A Plus mixture-of-experts system with 25 billion active parameters. The authors position it as a cost-effective option for machine translation workloads that need instruction-following behaviour.

papersTODAY 04:00 UTC

IBBench-Light benchmark tests whether models treat external records as instructions or text

A new arXiv paper introduces IBBench-Light, an evaluation that presents the same external record to a model under two different uses: as a procedure the model must carry out, or as text it must simply read. Each of twelve semantic bases produces 144 matched response pairs per model, and four quantized instruction-tuned models were tested. The paired setup is meant to isolate whether models react to a directive's form or to the user's stated task.

papersTODAY 04:00 UTC

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

A new arXiv paper introduces MTAC-IFBench, a benchmark aimed at measuring how well large language model agents follow instructions across multi-turn coding sessions. The work targets agentic software engineering, where models plan, run code, and call external tools over successive steps rather than producing a single answer. It addresses evaluation beyond functional correctness, focusing on whether agents keep to the constraints given to them.

papersSEP 12 04:00 UTC

Study Questions Whether Instruction Following Relies on a Single LLM Mechanism

A new arXiv paper examines whether instruction tuning gives language models a general-purpose ability to follow instructions or whether the behavior emerges from coordinated, task-specific components. The authors argue the evidence points to skillful coordination rather than one universal mechanism. The work adds to ongoing debate about what instruction tuning actually changes inside a model.

papersSEP 10 04:00 UTC

Can foundation models moderate online content? Comparing instruction- and example-driven policies

A new arXiv paper investigates whether foundation models can apply complex content moderation policies reliably and consistently. The study compares two ways of translating moderation rules into model behavior: conveying them through explicit instructions versus through illustrative examples. The findings are relevant to platforms seeking scalable, automated moderation of online content.