LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#benchmarks

40 curated events
papersTODAY 04:00 UTC

Study Examines Limits of Agentic ICD-10-CM Coding Benchmarks

A new arXiv paper analyzes how well agentic systems perform on ICD-10-CM medical coding, the alphanumeric codes used in the US for diagnoses, billing, and epidemiology. The authors argue that standard benchmarks rely on aggregate scores that hide poor performance on harder coding scenarios. The work aims to expose where current evaluation practices fall short for complex cases.

papersTODAY 04:00 UTC

Study Questions Reliability of Reference-Free Speech Quality Metrics for TTS

A new arXiv paper examines whether reference-free speech quality predictors such as UTMOS, DNSMOS and SCOREQ are dependable both as automatic evaluators for text-to-speech systems and as reward signals in preference optimization. The authors argue that these dual roles rest on assumptions about prediction accuracy that may not hold for modern TTS outputs. The work suggests current evaluation practices could mislead comparisons and reward-based training.

papersTODAY 04:00 UTC

arXiv Paper Audits Commonsense Reasoning Benchmarks Used for LLM Evaluation

A new arXiv paper argues that commonsense reasoning in language models is typically measured with multiple-choice benchmarks such as HellaSwag and PIQA, yet these benchmarks themselves are rarely scrutinized. The authors propose a more comprehensive approach to evaluating the benchmarks, questioning how well they actually capture the capability they claim to test. The work is framed as a meta-evaluation of standard commonsense reasoning tests.

papersTODAY 04:00 UTC

CapGeo-Bench separates visual perception from geometric reasoning in multimodal models

A new arXiv paper introduces CapGeo-Bench, a benchmark designed to evaluate geometric understanding in multimodal large language models while distinguishing failures in visual perception from failures in reasoning. The authors note that even strong closed models such as GPT-o3 continue to lag on geometry problems despite success on purely textual math tasks. The benchmark aims to give a clearer picture of where these systems break down.

papersTODAY 04:00 UTC

TestHallVQA benchmark probes document-level reasoning in vision-language models

A new arXiv paper introduces TestHallVQA, a benchmark built from scientific exam material for evaluating large vision-language models on visual question answering over long, multi-page documents. The authors argue that current planar VQA benchmarks tend to test isolated skills rather than document-level reasoning amid redundant context. The benchmark is intended to expose where such models fail when relevant information is buried in lengthy inputs.

papersTODAY 04:00 UTC

Mizan benchmark evaluates LLMs on Iraqi Arabic and civic context

Researchers introduced Mizan, a benchmark designed to test large language models on Iraqi Arabic and on civic topics relevant to Iraq. Existing Arabic evaluation efforts have largely centered on Modern Standard Arabic, leaving regional dialects and country-specific knowledge thinly covered. The work aims to give a national-level measure of model performance beyond aggregated MSA leaderboards.

papersTODAY 04:00 UTC

Paper Argues ASR Transcripts Are a Flawed Yardstick for Audio-LLM Tasks

A research paper examines how speech and audio LLMs are typically evaluated, namely by testing whether a waveform prompt outperforms an automatic speech recognition transcript. The authors argue that for closed-set, known tasks this setup mixes together two distinct things: whether the model actually used acoustic information and whether it simply needed the task spelled out. They propose auditing generative audio calls as a way to separate those factors.

papersTODAY 04:00 UTC

TOPO-Bench offers open-source framework for evaluating topological mapping

Researchers released TOPO-Bench, an open-source evaluation framework for topological mapping aimed at standardizing how navigation systems are compared. It introduces metrics, datasets and protocols, including a way to quantify perceptual aliasing, a common failure mode in place recognition. The work addresses the lack of shared benchmarks that has made results across topological mapping systems hard to compare.

papersTODAY 04:00 UTC

IBBench-Light benchmark tests whether models treat external records as instructions or text

A new arXiv paper introduces IBBench-Light, an evaluation that presents the same external record to a model under two different uses: as a procedure the model must carry out, or as text it must simply read. Each of twelve semantic bases produces 144 matched response pairs per model, and four quantized instruction-tuned models were tested. The paired setup is meant to isolate whether models react to a directive's form or to the user's stated task.

papersTODAY 04:00 UTC

WMT26 Builds Pseudo-References for 10 Reference-Free MT Language Pairs

Researchers describe the process used to create pseudo-references for the WMT26 General Machine Translation task, where ten language pairs lack any human translations or post-edited outputs. The work also covers six additional language pairs that do have references, aiming to give systems a consistent basis for automatic scoring. The paper details how these synthetic references were constructed and evaluated.

papersTODAY 04:00 UTC

arXiv paper examines robustness, cost and governance trade-offs in VLM document extraction

A new arXiv preprint argues that evaluations of vision-language models for extracting structured fields from business documents focus too heavily on accuracy against clean benchmarks. The authors propose assessing approaches along additional dimensions such as robustness, cost, and governance considerations, aiming to help practitioners pick a method suited to a given task complexity. No specific model or tool is released with the work.

papersTODAY 04:00 UTC

KnowBench proposes effort-reduction benchmark for clinical AI evaluation

A new arXiv preprint introduces KnowBench, a benchmark that assesses clinical AI systems by how much work they save clinicians instead of how closely their outputs match reference texts or expert rubrics. The authors argue that existing evaluation methods were built for research settings and measure resemblance to an artifact rather than reduction of a real-world burden. The paper frames deployment-grounded effort reduction as a unified metric for clinical AI.

papersTODAY 04:00 UTC

Paper Examines How First Query Shapes Agentic Deep Search

A new arXiv paper studies deep research agents that answer complex questions by repeatedly searching, reading, and reasoning. It argues that the quality of the initial search query is decisive, since well-tuned lexical retrieval can surface useful evidence early on benchmarks like BrowseComp-Plus. The authors frame the opening move as a strategic choice that shapes the rest of the search loop.

papersTODAY 04:00 UTC

Paper applies minimum description length to temporal misalignment in multichannel time-series classification

A new arXiv preprint examines how unsynchronized sensor streams — caused by latency, clock drift, or preprocessing — degrade multichannel time-series classification. The authors propose using the minimum description length principle to diagnose and characterize these relative delays. The work targets a setting where most existing methods assume channels are already aligned.

papersTODAY 04:00 UTC

Paper Argues Policy Ambiguity Skews Agent Benchmark Results

A new arXiv paper contends that agent benchmarks assume each policy implies one correct action, an assumption that natural-language policies often break through silence, ambiguity, or contradiction. The authors describe these as policy loopholes, where multiple defensible readings exist but evaluations still count a single behavior as an agent error. The work suggests such ambiguous cases should be separated from genuine policy-compliance failures in benchmark scoring.

papersTODAY 04:00 UTC

arXiv Paper Proposes Human-Grounded Calibration for Long-Text Image-Text Matching

A new arXiv preprint addresses the difficulty of judging whether lengthy descriptive text actually matches an image, a task relevant to vision-language systems. The authors note that raw similarity scores from dual-encoder models are hard to interpret and propose calibrating them against human judgments. The work targets more reliable long-text image-text congruence scoring.

papersTODAY 04:00 UTC

Study Compares Shell Commands and Specialized Tools for Enterprise AI Agents

A new arXiv paper empirically tests whether a general-purpose shell interface outperforms purpose-built tools when AI agents handle enterprise workflows. The authors note that shell-based agents perform well on coding tasks, but enterprise work also requires moving across applications and services and coordinating multiple steps. The study examines these trade-offs to identify which tool interface design suits digital worker agents.

papersTODAY 04:00 UTC

Study questions LLM-as-a-judge validity for psychological depth evaluations

A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.

papersTODAY 04:00 UTC

Audio encoders detect managerial evasiveness in earnings calls

A new arXiv preprint presents an approach that uses conversational audio encoders to spot evasive language from managers during earnings conference calls. Rather than aggregating vocal and lexical features across an entire call, the method analyzes the conversational dynamics between analysts and executives. Prior research has tied such cues to later negative outcomes for firms, and this work aims to capture them more precisely.

papersTODAY 04:00 UTC

Crypto Accounting Bench tests LLMs on reconstructing crypto transaction entries

Researchers released Crypto Accounting Bench, a set of 118 evaluation items that ask language models to rebuild the full journal entry an organization recorded for a crypto-asset transaction. The benchmark is applied to both frontier and open-weight models to gauge how well they handle this specialized accounting task. It appears on arXiv under cs.AI and cs.CL.

papersTODAY 04:00 UTC

Study Uses Activation Patching to Trace How VLMs Read Bar Chart Values

A new arXiv paper examines how vision-language models arrive at exact values when reading vertical bar charts. The authors apply counterfactual activation patching to trace where and how chart evidence is combined across space and depth in the network. The work argues that correct answers alone do not reveal the underlying mechanisms models use.

papersTODAY 04:00 UTC

Paper Probes LLM Benchmark Success Using Token-Level Perplexity

A new arXiv paper argues that standard task-performance evaluations of large language models reveal little about whether correct answers stem from the mechanisms researchers assume, which can encourage confirmation bias. The authors propose a simple, principled method that uses token-level perplexity to contrast how models behave on benchmarks with how they distribute probability internally. The work is a replacement submission to arXiv's computation and language section.

papersTODAY 04:00 UTC

VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition

A new arXiv paper introduces VoiceCodeBench, a benchmark that checks whether speech recognition transcripts reproduce exact written values rather than just scoring well on word error rate. It focuses on structured tokens such as identifiers, file paths, and commands, which voice-driven workflows often need verbatim. The work argues that WER alone does not capture whether these exact tokens survive transcription.

papersTODAY 04:00 UTC

Study Analyzes 160,000 Training Runs to Improve Offline Policy Learning Baselines

A new arXiv paper examines how reporting choices, hyperparameter tuning, and dataset characteristics affect offline policy learning results. Drawing on roughly 160,000 training runs, the authors argue that reliable progress requires careful reporting, well-tuned baselines, and evaluation across varied conditions. The work offers practical guidance for making policy-learning benchmarks more reproducible and comparable.

papersTODAY 04:00 UTC

arXiv Study Tests Shared KV Cache Across Two 27B vLLM Replicas

Researchers examined what happens when two single-GPU 27B vLLM inference replicas share a 256 GiB host-memory cache pool (LMCache), aiming to skip repeated prefill work as requests move between replicas. The paper reports both correctness problems in transferring cached state and performance limits tied to lost prefix locality. The authors argue that shared caching only pays off when state handoff is reliable and locality loss stays small.

papersTODAY 04:00 UTC

Turkish MMLU Pro Benchmark Examines Limits of Adding Answer Options

A new arXiv paper introduces Turkish MMLU Pro, a benchmark built from 12,000 Turkish-language questions spanning 58 sections. Each item keeps its original question stem and five answer choices, allowing researchers to test whether adding more options actually improves measurement quality. The authors argue that extra options can reduce scores without making the assessment more valid.

papersTODAY 04:00 UTC

Paper Proposes EventGraph and EventField Pipeline for Interpretable Temporal Video Reasoning

A new arXiv preprint describes a video reasoning approach that pairs a discrete event graph with a continuous event field, plus a human-readable glyph view, so intermediate reasoning steps can be inspected. The authors evaluate the pipeline on a curated EPIC-KITCHENS subset containing 10 videos and 50 questions about temporal relationships. The work sits in the interpretability and video-language research space rather than announcing a product or model release.

papersTODAY 04:00 UTC

Search APIs Evaluated as Decision Surfaces for Tool-Using AI Agents

A new arXiv paper examines how the ranked snippets, URLs, and metadata returned by search APIs shape the choices made by tool-using AI agents, such as whether to answer, search again, or open a page. The authors evaluate these interfaces as decision surfaces using a fixed set of 100 questions drawn from the 254-question SealQA-Hard benchmark. The work suggests that agents can reach similar accuracy while relying on differing amounts or quality of supporting evidence.

papersTODAY 04:00 UTC

RAMP Framework Rates Repository AI Maturity After Coding-Agent Adoption

A new arXiv paper introduces RAMP (Repository AI Maturity Profile), a four-level cumulative scale for describing how deeply coding agents are integrated into a software repository. The authors argue that prior studies report only average outcomes across adopters, which masks large variation between teams, and that greater agent adoption can come with higher quality costs and technical debt.

papersTODAY 04:00 UTC

Study Compares SmolVLA Task Success and Latency Across PyTorch and ONNX Deployments

A new arXiv paper examines how deploying the SmolVLA vision-language-action model in different runtime formats affects both inference speed and closed-loop task performance. The authors benchmark HuggingFaceVLA/smolvla_libero on a 6 GB RTX 2060 across the LIBERO Spatial and Object suites using MuJoCo and LeRobot with a fixed seed. The results indicate that cutting latency through optimized deployment can shift task behavior, so faster inference does not automatically mean better outcomes.

papersTODAY 04:00 UTC

Paper Proposes Diagnostics for LLM-Based Synthetic Consumer Panels

A new arXiv preprint examines how large language models are used as stand-ins for survey respondents, a practice that can cut costs dramatically compared with traditional polling. The authors argue that aggregate validation scores hide systematic problems such as compressed variance and flipped coefficient signs, and they offer diagnostic and correction methods to address them.

papersTODAY 04:00 UTC

PhysMent benchmark evaluates LLM physics reasoning through interactive experiments

Researchers introduced PhysMent, a benchmark designed to test how well large language models reason about physical systems by running experiments rather than answering static questions. The work argues that strong scores on existing science benchmarks do not show whether models can actively probe the physical world. The abstract notes that this ability remains poorly understood.

papersTODAY 04:00 UTC

Study proposes vulnerability modeling and execution-based benchmark for secure code generation

A new arXiv paper addresses the gap between code that runs correctly and code that is secure when generated by large language models. The authors argue that progress has been limited by existing benchmarks that are small and not executable, making security flaws hard to measure reliably. Their approach combines task-adaptive modeling of vulnerabilities with an execution-based benchmark intended to evaluate both functional correctness and security.

papersTODAY 04:00 UTC

Study Measures Citation Attribution Across RAG Context Compression Methods

A new arXiv paper examines how context compression in retrieval-augmented generation affects citation attribution, not just answer quality. The authors benchmark reranking and extractive compression approaches under varying compression budgets to map what they call the attribution-compression frontier. The work suggests that evaluating answer accuracy alone misses important degradation in source attribution.

papersTODAY 04:00 UTC

GroundBench benchmark aims to pinpoint where vision-language models fail on affordance tasks

A new arXiv paper introduces GroundBench, described as a factorized, counterfactual benchmark for identifying the specific points at which vision-language models break down on affordance tasks. The work cites a companion evaluation in which explicitly naming the target part in a manipulation prompt improved action accuracy by 0.32 to 0.63 across eight vision-language models, and no model exceeded a constant baseline before that part was named. The benchmark is intended to isolate these failures rather than report only aggregate scores.

papersTODAY 04:00 UTC

ChartAnno Benchmark Tests Multimodal LLMs on Chart Annotation Generation

A new research benchmark called ChartAnno evaluates how well multimodal large language models can generate annotations for charts, a task that helps explain data and highlight key findings in visualizations. The work examines whether these models can automate annotation authoring, which is normally done by hand. It is presented as an arXiv paper revision.

papersSEP 10 04:00 UTC

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

A revised arXiv paper introduces a benchmark designed to test AI agents on tasks outside the well-known applications that dominate current evaluations. The authors contend that testing in familiar, comparatively simple settings can mask how poorly agents generalize to novel situations. The work aims to give a more accurate picture of how agentic systems will behave in real-world deployment.

papersTODAY 04:00 UTC

FairFund-Bench Benchmark Tests Distributive Bias in LLM Resource Allocation

A new arXiv paper introduces FairFund-Bench, a benchmark for measuring how large language models distribute scarce resources and whether those allocations vary by race, gender, or similar traits. The authors note that prior audits of LLM bias have yielded conflicting findings, and position their benchmark as a way to standardize such evaluations. The work targets fairness in settings where models take part in allocating limited goods or funds.

papersTODAY 04:00 UTC

arXiv paper evaluates open-source LLMs for RAG in ESG reporting

A new arXiv preprint examines how well open-source large language models perform when paired with retrieval-augmented generation for environmental, social, and governance reporting tasks. The authors focus on automating the extraction of key performance indicators from ESG disclosures, a step they describe as important for corporate accountability. The abstract suggests limits in current open-source model performance for this domain.