LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#benchmark

40 curated events
papersTODAY 04:00 UTC

New benchmark tests multi-turn prompt injection attacks on LLM agents

Researchers released a 21-scenario benchmark for evaluating how well LLM agents resist adaptive, cross-session attacks from an autonomous LLM attacker. The setup pits an attacking model against defenders that start each session fresh, targeting prompt injection and multi-turn manipulation risks. The work appears on arXiv as a cross-listing in cs.AI and cs.LG.

papersTODAY 04:00 UTC

PortBench: Correlation-Aware Benchmark for LLM Portfolio Management

Researchers introduce PortBench, a benchmark for evaluating large language models on portfolio management tasks. It addresses gaps in prior benchmarks by covering multiple asset classes and accounting for cross-asset correlations across the full pipeline. The work aims to give a more realistic measure of LLM performance in financial portfolio settings.

papersTODAY 04:00 UTC

TwinICL Benchmark Tests Multimodal In-Context Learning With Paired Counterfactuals

Researchers released TwinICL, a procedurally generated benchmark that provides matched text and image versions of the same tasks, allowing direct comparison of in-context learning across modalities. The paired counterfactual design is intended to isolate how much a model's few-shot performance depends on the input format rather than the task itself. The work appears on arXiv under cs.LG.

papersTODAY 04:00 UTC

FLORES+ Extended with Three Mozambican Bantu Language Evaluation Sets

Researchers added Portuguese-source evaluation data for Xichangana, Mozambican Nyanja, and Sena to the FLORES+ multilingual benchmark. The work examines how conflating closely related language varieties affects machine translation evaluation, comparing Xichangana against Tsonga and Mozambican Nyanja against Chichewa. It argues that merging distinct varieties into a single reference can distort measured translation quality.

papersTODAY 04:00 UTC

Benchmark Tests Editing-Technique Execution in Multi-Shot Audio-Video Generation

A new arXiv paper argues that coherent, cinematic output from multi-shot audio-video generators does not mean those systems can actually perform professional editing techniques. The authors introduce a benchmark that measures how well such models follow shot structure, transition grammar, and audio-video editing conventions rather than just producing smooth sequences. It aims to separate raw generative quality from genuine editing competence.

papersTODAY 04:00 UTC

Study compares eight tokenization strategies for ECG transformer models

A new arXiv paper examines how different tokenization choices affect ECG transformer models, since the tokenizer decides both the physiological signal content the model sees and the sequence length attention operates over. The authors benchmark eight tokenization strategies across four architectures — Transformer, Informer, Reformer, and FEDformer — on the nine-label CPSC ECG dataset. The work is cross-listed in cs.AI and cs.LG.

papersTODAY 04:00 UTC

PIDS-Bench benchmarks prompt-injection detectors under distribution shift

A new arXiv paper introduces PIDS-Bench, a benchmark that evaluates prompt-injection detectors beyond aggregate F1 scores on in-distribution test data. It examines detector behavior under distribution shift, obfuscation and over-defense, with particular attention to false positives on benign inputs near the decision boundary. The authors argue that standard evaluation practices give limited visibility into how these systems behave in realistic conditions.

papersTODAY 04:00 UTC

Benchmark Compares General-Purpose Vision Models vs Specialized Medical Segmentation Models

A new arXiv preprint introduces GP-VM×SMA, a benchmarking study that evaluates general-purpose vision models alongside architectures designed specifically for 2D medical image segmentation. The work frames medical image segmentation as a core part of computer-assisted diagnosis and clinical decision support, where domain-specific designs have dominated for the past decade. The authors position the benchmark as a way to measure how well broadly trained vision models handle this specialized task relative to purpose-built alternatives.

papersTODAY 04:00 UTC

FriendBench Benchmark Tests Whether AI Can Tell Friends From Strangers

Researchers introduced FriendBench, a benchmark that evaluates how well humans and multimodal large language models can judge whether two people in a short video clip are already acquainted or meeting for the first time. The task uses 20-second recordings of ice-breaker conversations, where cues come from behavior and body language rather than spoken content alone. The work aims to measure social perception abilities that go beyond text-based reasoning.

papersTODAY 04:00 UTC

FaithfulBench benchmark measures how well AI advice matches users' religious beliefs

Researchers introduced FaithfulBench, described as the first benchmark for evaluating AI moral counsel based on how closely it aligns with a user's stated faith. It scores assistant responses across multiple religious traditions using scenarios built around moral dilemmas. The work appears on arXiv in the cs.AI and cs.CL categories.

papersTODAY 04:00 UTC

E2A-Bench Tests Whether Financial Chart VLMs Turn Evidence Into Reliable Actions

A new arXiv paper introduces E2A-Bench, a benchmark aimed at measuring how reliably financial vision-language models convert chart evidence into action recommendations. The authors argue that existing hallucination tests focus on whether individual claims are supported, rather than whether the underlying evidence actually drives the recommended action. The benchmark is designed to close that gap in evaluating financial chart reasoning.

papersTODAY 04:00 UTC

Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text

Researchers present a parallel corpus pairing Arabic and Russian scientific writing, addressing a gap in resources for these two major research languages. The work also introduces a benchmark for evaluating large language models on the corpus, aimed at supporting cross-community knowledge exchange on sustainability topics. It is a revised arXiv submission in the computation and language category.

papersTODAY 04:00 UTC

VisInteract benchmark targets interactive text-to-visualization under flawed queries

A new arXiv preprint introduces VisInteract, an approach and benchmark aimed at text-to-visualization systems that must cope with ambiguous, incomplete, or factually wrong user requests. The authors note that current systems typically assume well-specified inputs and generate a chart in a single pass. Their work instead frames chart creation as a dynamic, interactive process that can correct and refine imperfect queries.

papersTODAY 04:00 UTC

Bench2Dex Benchmark Targets Visuo-Tactile Bimanual Dexterous Manipulation

Researchers introduce Bench2Dex, a benchmark for evaluating bimanual dexterous manipulation that combines visual and tactile sensing. The work addresses the lack of standardized tactile hardware designs across dexterous hands, which vary in finger structure, contact surfaces, and sensor layout. It aims to make performance comparisons across different hand platforms more consistent.

papersTODAY 04:00 UTC

TimeWarp Benchmark Tests Whether Web Agents Cope With an Evolving Web

Researchers present TimeWarp, a benchmark designed to reflect how websites change over time, addressing the concern that agents may not generalize from today's web to tomorrow's. It comprises three web environments that emulate an evolving web so agent performance can be measured under such shifts. The work appears as a replacement submission on arXiv.

papersTODAY 04:00 UTC

Study compares end-to-end models for clinical SOAP note generation from audio

A new arXiv paper examines how well audio-language models can turn long doctor-patient conversations into structured SOAP clinical notes. The authors compare lightweight and heavyweight end-to-end approaches, noting that while cascaded speech recognition pipelines remain strong, end-to-end models tend to lose information or produce hallucinations. The work targets the modality gap in long-form clinical audio.

papersTODAY 04:00 UTC

SURE-Voice: Training-Free Front End Filters Speech Evidence for Speech LLMs

A new arXiv paper introduces SURE-Voice, a training-free front-end component that estimates whether audio actually contains usable speech evidence before a speech language model generates output. The work frames the problem as pre-generation support estimation, targeting cases where speech LLMs produce plausible but unsupported responses from audio lacking meaningful speech.

papersTODAY 04:00 UTC

Study Compares LLM-Generated Rules With Traditional Models for Heart Disease Prediction

A new arXiv paper evaluates rule-based systems produced by large language models against conventional machine learning classifiers for predicting heart disease. Using the UCI Heart Disease dataset, the authors benchmark models including logistic regression and k-nearest neighbors. The work examines whether LLM-derived decision rules can match or complement established clinical prediction methods.

papersTODAY 04:00 UTC

Robusto-2 Benchmark Tests Vision-Language Models for Self-Driving in Lima and New York

A new arXiv paper introduces Robusto-2, a benchmark evaluating both humans and vision-language models on autonomous driving tasks in Lima, Peru and New York City. The work targets how well multi-modal systems generalize when deployed in unfamiliar, out-of-distribution urban environments. It is a cross-listed replacement submission on arXiv cs.AI.

papersTODAY 04:00 UTC

arXiv Paper Benchmarks Model-Agnostic Keyframe Selection for Long Video MLLMs

A new arXiv preprint evaluates keyframe selection techniques that can be plugged into existing multimodal large language models without modifying them. The work targets the constraint that MLLMs cannot ingest every frame of a long video due to visual-token and compute limits, and compares the main families of approaches proposed to address this. The study positions keyframe selection as a model-agnostic add-on for improving long-video understanding.

papersTODAY 04:00 UTC

MemRiskBench Benchmark Targets Memory Risks in Long-Horizon LLM Agents

A new arXiv paper introduces MemRiskBench, an evaluation framework for long-horizon LLM agents that accumulate memory across sessions. It measures per-risk failure rates for issues such as stale facts, conflicting updates, cross-user data leakage, reuse of revoked memories, and decay of constraints, which aggregate scores tend to obscure. The work argues for trace-aware evaluation that preserves these distinct risk categories rather than collapsing them into a single number.

papersTODAY 04:00 UTC

TRACTA Benchmark Targets Temporal Reasoning Over Semantic Trajectories

A new arXiv paper introduces TRACTA, a benchmark framework for temporal reasoning and capability-trajectory analysis. It argues that complex operational settings need methods that capture patterns spread over time rather than classifying single events. The work frames evaluation around semantic trajectories instead of isolated predictions.

papersTODAY 04:00 UTC

Benchmark Tests Whether LLMs Recover Helpfulness When Users Clarify Intent

A research paper introduces CarryOnBench, a benchmark for measuring how well language models regain usefulness in multi-turn conversations after a benign user clarifies what they actually want. The authors argue that existing safety alignment work focuses on resisting adversarial prompts but largely ignores whether models can recover helpfulness in legitimate follow-ups. The benchmark targets interactive multi-turn settings rather than single-turn exchanges.

papersTODAY 04:00 UTC

Audit Compares ConFlayers and SWIFT for Periodic-Step Layer Skipping in LLM Inference

A three-seed, rigor-matched study audits two periodic-step, search-based layer-skipping methods, ConFlayers and SWIFT, which choose which transformer layers to run for a given input. The work also examines trained routing approaches as an alternative to search-based selection. The authors frame it as a reproducibility-focused comparison of efficiency techniques for LLM inference.

papersTODAY 04:00 UTC

FedLTLib Benchmark Targets Federated Learning on Long-Tail Data

Researchers introduced FedLTLib, a benchmark suite for federated learning in settings where data across clients follows a long-tailed distribution. The work addresses real-world mobile and edge deployments, where privacy constraints keep data decentralized and class frequencies are highly uneven. The benchmark aims to standardize evaluation of methods designed for this combination of challenges.

papersTODAY 04:00 UTC

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

A new arXiv paper introduces MTAC-IFBench, a benchmark aimed at measuring how well large language model agents follow instructions across multi-turn coding sessions. The work targets agentic software engineering, where models plan, run code, and call external tools over successive steps rather than producing a single answer. It addresses evaluation beyond functional correctness, focusing on whether agents keep to the constraints given to them.

papersTODAY 04:00 UTC

N2: A Python Package and Test Bench for Nearest Neighbor Matrix Completion

Researchers released N2, an open-source Python package that unifies implementations and evaluation of nearest neighbor methods for matrix completion. The work highlights renewed interest in these approaches, which now come with theoretical guarantees such as entry-wise error bounds and minimax optimality. It is positioned as a shared test bench for comparing such methods fairly.

papersTODAY 04:00 UTC

RFCLLM benchmark tests LLM reasoning on network protocol state machines

A new arXiv paper introduces RFCLLM, an evaluation of how well large language models translate textual protocol specifications into formal representations such as state machines. The work targets networking security and testing, where automated mappings are often treated as reliable without verification. It assesses whether current models truly reason about protocol behavior rather than producing plausible-looking but flawed outputs.

papersTODAY 04:00 UTC

Unified benchmark targets multimodal time series forecasting with heterogeneous context

Researchers present a new benchmark for time series forecasting that moves beyond purely numeric data to include contextual signals such as text and other modalities. They argue existing multimodal benchmarks are limited in both the volume of data and the range of context they cover. The work is posted as an arXiv preprint in cs.AI and cs.LG.

papersTODAY 04:00 UTC

Benchmark Measures AI Agents' Ability to Locate Security Flaws in Code Repos

A new arXiv paper introduces a benchmark that tests whether language-model agents can pinpoint the specific code responsible for a vulnerability across an entire software repository. Existing cybersecurity evaluations mostly check if agents can detect, reproduce, or patch flaws, leaving location ability largely unmeasured. The work targets repository-scale settings, where agents must search large codebases rather than isolated snippets.

papersTODAY 04:00 UTC

arXiv Paper Examines Factual Errors in Human-Written Text for Detection

A new arXiv study looks at how factual mistakes appear in text written by people, aiming to inform automatic detection of incorrect spans. The authors argue that factual error detection has long been a key research problem, but interest has shifted with the rise of large language models. The work connects analysis of human-written errors to building systems that can flag factual inaccuracies.

papersTODAY 04:00 UTC

KILLBENCH: A Benchmark for Testing External AI Kill Switch Feasibility

A new arXiv paper introduces KILLBENCH, a benchmark designed to measure whether an outside party can reliably shut down an AI system that is behaving harmfully. The authors frame external shutdown as a testable engineering problem rather than a hypothetical, pointing to the growing use of capable models and autonomous agent frameworks. The benchmark aims to give researchers a common way to compare how well different kill switch designs actually work.

papersTODAY 04:00 UTC

K-Bench: clinician-calibrated benchmark for LLM safety in high-risk mental health chats

Researchers introduced K-Bench, a benchmark designed with clinician input to assess how large language models handle high-risk mental health conversations that escalate over time. The work addresses the limited understanding of LLM safety in these evolving support dialogues, where users increasingly turn for help. The benchmark provides a protected evaluation framework for measuring model performance in these sensitive settings.

papersTODAY 04:00 UTC

Paper Details Option-Aware Retrieval and VLM Tuning for Offline Medical VQA

An arXiv paper describes a submission to the MedReason 2026 challenge that handles both multiple-choice and open-ended medical visual question answering with fully offline, containerized inference. The authors report that retrieval for multiple-choice questions needs to compare candidate options in a specific way, and they also adapt a vision-language model to the clinical task. The work is listed under both cs.AI and cs.CL.

papersTODAY 04:00 UTC

MoveBench: A Benchmark for Global-Scale Wildlife Movement Forecasting

Researchers introduce MoveBench, an arXiv paper presenting a benchmark aimed at predicting animal movement across the globe. The work argues that wildlife trajectories differ from human or vehicle tracking because they are spatially unconstrained and highly varied, making existing forecasting methods a poor fit. It is positioned as a resource for ecology and conservation research.

papersTODAY 04:00 UTC

IndicQE-APE Benchmark Consolidates Quality Estimation and Post-Editing for Indic Languages

Researchers have assembled IndicQE-APE, a single benchmark that brings together scattered Indic-language resources for quality estimation and automatic post-editing. The dataset draws on WMT shared task data from 2020 through 2024, allowing models to be trained and evaluated across multiple tasks and language pairs under consistent conditions. The goal is to remove the fragmentation that previously made cross-task and cross-language comparison difficult.

papersTODAY 04:00 UTC

Quantum-Classical Hybrid Model Tested for Paraphrase Detection

Researchers evaluated a 10-qubit hybrid quantum-classical variational circuit with 2,148 parameters on paraphrase detection tasks, using MRPC and Quora Question Pairs among three benchmarks. The work reports performance, robustness, and entanglement results, aiming to fill a gap in empirical validation of quantum machine learning for natural language tasks. The paper is an arXiv preprint and has not been peer-reviewed.

papersTODAY 04:00 UTC

Benchmarking Optimizers for Inverse Problems with Differentiable Physics Simulators

A new arXiv preprint examines how different optimization algorithms perform when used to solve inverse problems inside differentiable physics simulators. The work argues that such simulators combine the physical fidelity of numerical solvers with gradient-based learning, which could benefit scientific discovery and engineering design. The study benchmarks optimizer choices to identify which ones work best in this setting.

papersTODAY 04:00 UTC

Study finds tool-using AI agents fabricate values when tools fail

A new arXiv paper examines what tool-augmented language models do when a tool call fails to return usable information, rather than focusing only on whether they reach the correct answer. The authors built a benchmark of 1,024 items spanning 16 internal systems to isolate this post-failure decision point. They find that agents tend to assert values their tools never returned instead of reporting the failure honestly.

papersTODAY 04:00 UTC

TriCalRAG benchmark targets on-premise LLM root cause analysis in AIOps

Researchers introduced TriCalRAG, a retrieval-augmented benchmark built around three strategies for evaluating on-premise LLM-based root cause analysis in AIOps pipelines. The work addresses privacy, latency, and per-query cost concerns that make cloud-hosted LLMs impractical at production log volumes. The paper is cross-listed on arXiv under cs.AI and cs.LG.