LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#evaluation

40 curated events
papersTODAY 04:00 UTC

Mizan benchmark evaluates LLMs on Iraqi Arabic and civic context

Researchers introduced Mizan, a benchmark designed to test large language models on Iraqi Arabic and on civic topics relevant to Iraq. Existing Arabic evaluation efforts have largely centered on Modern Standard Arabic, leaving regional dialects and country-specific knowledge thinly covered. The work aims to give a national-level measure of model performance beyond aggregated MSA leaderboards.

papersTODAY 04:00 UTC

Study Questions Whether Consistent Local LLM Judges Match Human Ratings

A new arXiv paper examines the use of local large language models as automated judges of other models' outputs, a practice meant to cut the cost and time of human evaluation. The authors investigate whether a judge that returns stable, consistent scores can still be unreliable when compared against human ratings. The work argues that consistency alone is not sufficient evidence of a trustworthy evaluator.

papersTODAY 04:00 UTC

Sparse Autoencoders Can Preserve Different Readouts at Equal Reconstruction Error

A new arXiv paper argues that matching reconstruction error and sparsity levels does not guarantee two sparse autoencoders capture the same linearly decodable information from model activations. The authors formalize this gap as a matrix-valued distortion between optimal ridge readouts and propose decoder-preserving training objectives. The work offers a way to evaluate which downstream signals survive sparse compression.

papersTODAY 04:00 UTC

VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition

A new arXiv paper introduces VoiceCodeBench, a benchmark that checks whether speech recognition transcripts reproduce exact written values rather than just scoring well on word error rate. It focuses on structured tokens such as identifiers, file paths, and commands, which voice-driven workflows often need verbatim. The work argues that WER alone does not capture whether these exact tokens survive transcription.

papersTODAY 04:00 UTC

Crypto Accounting Bench tests LLMs on reconstructing crypto transaction entries

Researchers released Crypto Accounting Bench, a set of 118 evaluation items that ask language models to rebuild the full journal entry an organization recorded for a crypto-asset transaction. The benchmark is applied to both frontier and open-weight models to gauge how well they handle this specialized accounting task. It appears on arXiv under cs.AI and cs.CL.

papersTODAY 04:00 UTC

Paper Argues ASR Transcripts Are a Flawed Yardstick for Audio-LLM Tasks

A research paper examines how speech and audio LLMs are typically evaluated, namely by testing whether a waveform prompt outperforms an automatic speech recognition transcript. The authors argue that for closed-set, known tasks this setup mixes together two distinct things: whether the model actually used acoustic information and whether it simply needed the task spelled out. They propose auditing generative audio calls as a way to separate those factors.

papersTODAY 04:00 UTC

Paper Proposes Framework for Judging When Synthetic Survey Data Is Trustworthy

A new arXiv paper argues that the debate over synthetic data in marketing research has been stuck between two extremes: treating large language models as a replacement for human survey respondents, or rejecting them outright. The authors say the more useful question is when synthetic respondents can be trusted, and they outline how that reliability should be evaluated. The work focuses on marketing research but touches on broader issues of validating model-generated data.

papersTODAY 04:00 UTC

Study questions LLM-as-a-judge validity for psychological depth evaluations

A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.

papersTODAY 04:00 UTC

arXiv Paper Proposes Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

A new arXiv paper argues that judging multi-agent LLM systems only by their final answers misses how those systems actually reach their results. The authors propose a diagnostic method grounded in human group research to distinguish process losses from assembly bonuses when LLM teams collaborate. This matters both for building better agent pipelines and for using LLM groups as stand-ins for human group behavior.

papersTODAY 04:00 UTC

arXiv paper evaluates LoRA fine-tuning scale and rank for control-systems Q&A

A new arXiv preprint examines how LoRA fine-tuning performs on question answering for a control-systems university course. The study measures results across model sizes and LoRA rank settings, since such questions demand consistent terminology, notation, derivations, and step-by-step reasoning. It appears to be a multidimensional evaluation of whether parameter-efficient tuning can handle specialized technical coursework.

papersTODAY 04:00 UTC

Paper Argues Policy Ambiguity Skews Agent Benchmark Results

A new arXiv paper contends that agent benchmarks assume each policy implies one correct action, an assumption that natural-language policies often break through silence, ambiguity, or contradiction. The authors describe these as policy loopholes, where multiple defensible readings exist but evaluations still count a single behavior as an agent error. The work suggests such ambiguous cases should be separated from genuine policy-compliance failures in benchmark scoring.

papersTODAY 04:00 UTC

Paper Probes LLM Benchmark Success Using Token-Level Perplexity

A new arXiv paper argues that standard task-performance evaluations of large language models reveal little about whether correct answers stem from the mechanisms researchers assume, which can encourage confirmation bias. The authors propose a simple, principled method that uses token-level perplexity to contrast how models behave on benchmarks with how they distribute probability internally. The work is a replacement submission to arXiv's computation and language section.

papersTODAY 04:00 UTC

KnowBench proposes effort-reduction benchmark for clinical AI evaluation

A new arXiv preprint introduces KnowBench, a benchmark that assesses clinical AI systems by how much work they save clinicians instead of how closely their outputs match reference texts or expert rubrics. The authors argue that existing evaluation methods were built for research settings and measure resemblance to an artifact rather than reduction of a real-world burden. The paper frames deployment-grounded effort reduction as a unified metric for clinical AI.

papersTODAY 04:00 UTC

FaithfulBench benchmark measures how well AI advice matches users' religious beliefs

Researchers introduced FaithfulBench, described as the first benchmark for evaluating AI moral counsel based on how closely it aligns with a user's stated faith. It scores assistant responses across multiple religious traditions using scenarios built around moral dilemmas. The work appears on arXiv in the cs.AI and cs.CL categories.

papersTODAY 04:00 UTC

E2A-Bench Tests Whether Financial Chart VLMs Turn Evidence Into Reliable Actions

A new arXiv paper introduces E2A-Bench, a benchmark aimed at measuring how reliably financial vision-language models convert chart evidence into action recommendations. The authors argue that existing hallucination tests focus on whether individual claims are supported, rather than whether the underlying evidence actually drives the recommended action. The benchmark is designed to close that gap in evaluating financial chart reasoning.

papersTODAY 04:00 UTC

Study questions realism of language-model agents in farming decision simulations

A new arXiv paper examines whether language-model agents can credibly stand in for human respondents in surveys and social simulations. The authors argue that judging realism from population averages or distributional similarity can be misleading, an effect they call the "average-farmer illusion." Their experiments test what such aggregate evidence actually demonstrates about individual-level behavior.

papersTODAY 04:00 UTC

Reference-Free Metric Targets Lexical Tone Evaluation in Multilingual TTS

A new arXiv paper proposes a reference-free way to measure whether text-to-speech systems render lexical tone correctly in languages where pitch changes word meaning. The authors note that standard character error rate scoring misses such errors, using Yorùbá words like "ọkọ" (husband), "ọkọ̀" (vehicle) and "ọkọ́" (hoe), which differ only by tone, as an illustration. The work aims to give a budget-friendly evaluation option for massively multilingual speech systems.

papersTODAY 04:00 UTC

arXiv Paper Proposes Human-Grounded Calibration for Long-Text Image-Text Matching

A new arXiv preprint addresses the difficulty of judging whether lengthy descriptive text actually matches an image, a task relevant to vision-language systems. The authors note that raw similarity scores from dual-encoder models are hard to interpret and propose calibrating them against human judgments. The work targets more reliable long-text image-text congruence scoring.

papersTODAY 04:00 UTC

Study examines how editorial routing affects qualification of AI-assisted computational results

This arXiv paper investigates how large language models are used to interpret computational results and draft scientific manuscripts. Holding the underlying computational evidence fixed, the authors tested whether spreading comparisons across different modeling choices changes how findings are described and qualified. The results indicate that editorial routing decisions influence the hedging and qualification of reported results in AI-assisted writing.

papersTODAY 04:00 UTC

arXiv paper proposes document topic alignment metrics for topic models of health social media

A new arXiv preprint argues that existing evaluations of topic models for short social media texts rely almost entirely on measures of the generated topics themselves. The author proposes metrics that also assess how well documents align with their assigned topics. The work targets public health communication data drawn from social media.

papersTODAY 04:00 UTC

Paper argues compliance data is often misused as evaluation data for AI systems

A new arXiv paper claims that a common mistake in assessing deployed AI systems is treating data gathered for operational monitoring or regulatory compliance as though it were collected for comparative evaluation. Using automated driving as its main example, the work calls for clearer measurement validity standards so that compliance-oriented datasets are not used to make comparative performance claims. The authors frame this as a recurring evaluation failure rather than an isolated incident.

papersTODAY 04:00 UTC

Study examines hindsight bias in clinical LLM temporal reasoning

A new arXiv paper argues that clinical language models are frequently assessed on retrospective patient records that already contain the eventual diagnosis, treatment response and outcome. Because those records expose information a real prospective decision-maker would not have, such benchmarks may reward models for exploiting future data instead of genuine reasoning. The work examines how this exposure shapes model judgments in clinical temporal tasks.

papersTODAY 04:00 UTC

arXiv paper examines deductive, inductive and abductive reasoning in language models

A revised arXiv preprint analyzes how language models handle three forms of reasoning: deduction, induction, and abduction. The authors compare ways tasks are specified to models, such as explicit instructions versus few-shot examples, and argue that current evaluations leave parts of the reasoning picture unresolved. The work is a research paper rather than a product or model release.

papersTODAY 04:00 UTC

Turkish MMLU Pro Benchmark Examines Limits of Adding Answer Options

A new arXiv paper introduces Turkish MMLU Pro, a benchmark built from 12,000 Turkish-language questions spanning 58 sections. Each item keeps its original question stem and five answer choices, allowing researchers to test whether adding more options actually improves measurement quality. The authors argue that extra options can reduce scores without making the assessment more valid.

papersTODAY 04:00 UTC

Study Measures Generation Gap Between Speech-Only and Speech-Text Models

A new arXiv paper proposes a way to quantify the quality gap between speech-only language models and those that also handle text. The authors note that the gap is hard to measure because speech and text systems are usually trained on different data and judged with different metrics. The method aims to put the two modalities on a comparable footing.

papersTODAY 04:00 UTC

arXiv paper tests staged prompts across six frontier AI models

A newly posted arXiv paper describes experiments in which the same three-part prompt sequence was run ten times for each of six frontier AI models from OpenAI, Anthropic, xAI and Google DeepMind. The prompts move from asking about architectural preferences toward a fuller task, suggesting the study compares how different systems respond as questioning gets more demanding. The abstract is truncated, so final findings and conclusions are not yet visible.

papersTODAY 04:00 UTC

arXiv Paper Proposes Slice-Wise Non-Regression Checks for Model Upgrades

A new preprint addresses the problem of selecting a model checkpoint that improves overall performance without hurting specific data slices that matter to downstream users. It frames checkpoint selection as a comparison against a retained incumbent model, subject to tolerance thresholds for acceptable degradation. The work separates different kinds of failure relative to those tolerances and offers certification-style guarantees with fallback to the incumbent.

papersTODAY 04:00 UTC

DynSTEER: Dynamic Stage-wise Evaluation and Review for LLM Agents

A new arXiv paper introduces DynSTEER, a framework for evaluating large language model agents that operate over long-horizon tasks. It targets gaps in existing evaluation methods, which typically judge only final outcomes and struggle to pinpoint where errors occur. The approach combines stage-wise trajectory assessment with review during execution rather than after the fact.

papersTODAY 04:00 UTC

Framework Tests Whether AI Agents Can Predict A/B Test Outcomes

A new arXiv paper proposes a validation framework for using AI agents to simulate the results of A/B tests, which normally require real user traffic, engineering time, and weeks of waiting. The approach conditions agents on behavioral profiles to estimate experiment outcomes before a rollout. The work focuses on how such simulations should be checked for accuracy rather than on a specific product.

papersSEP 10 04:00 UTC

No Free Checker: A Survey of Verifiers for Robot Policies

A new survey paper catalogues methods that score robot behaviors, ranging from success detectors and reward models to runtime monitors. It examines how these verification approaches serve two roles: evaluating vision-language-action policies and providing training signals for them. The work appears on arXiv, cross-listed between the AI and machine learning categories.

papersTODAY 04:00 UTC

TRACTA Benchmark Targets Temporal Reasoning Over Semantic Trajectories

A new arXiv paper introduces TRACTA, a benchmark framework for temporal reasoning and capability-trajectory analysis. It argues that complex operational settings need methods that capture patterns spread over time rather than classifying single events. The work frames evaluation around semantic trajectories instead of isolated predictions.

papersTODAY 04:00 UTC

arXiv primer surveys evaluation methods for LLMs in healthcare

A new arXiv paper reviews how large language models used in clinical and medical settings should be assessed. It argues that evaluating these systems is harder than conventional machine learning evaluation for a variety of reasons. The work is framed as an introductory guide to evaluation approaches for healthcare LLMs.

papersTODAY 04:00 UTC

Search APIs Evaluated as Decision Surfaces for Tool-Using AI Agents

A new arXiv paper examines how the ranked snippets, URLs, and metadata returned by search APIs shape the choices made by tool-using AI agents, such as whether to answer, search again, or open a page. The authors evaluate these interfaces as decision surfaces using a fixed set of 100 questions drawn from the 254-question SealQA-Hard benchmark. The work suggests that agents can reach similar accuracy while relying on differing amounts or quality of supporting evidence.

papersTODAY 04:00 UTC

GroundBench benchmark aims to pinpoint where vision-language models fail on affordance tasks

A new arXiv paper introduces GroundBench, described as a factorized, counterfactual benchmark for identifying the specific points at which vision-language models break down on affordance tasks. The work cites a companion evaluation in which explicitly naming the target part in a manipulation prompt improved action accuracy by 0.32 to 0.63 across eight vision-language models, and no model exceeded a constant baseline before that part was named. The benchmark is intended to isolate these failures rather than report only aggregate scores.

papersSEP 10 04:00 UTC

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

A revised arXiv paper introduces a benchmark designed to test AI agents on tasks outside the well-known applications that dominate current evaluations. The authors contend that testing in familiar, comparatively simple settings can mask how poorly agents generalize to novel situations. The work aims to give a more accurate picture of how agentic systems will behave in real-world deployment.

papersTODAY 04:00 UTC

STAGE Diagnoses Semantic-Action Gap in Embodied Agents

A new arXiv paper examines why embodied agents can correctly identify what an instruction refers to yet still fail to act on it correctly, a problem the authors call the semantic-action gap. The proposed STAGE framework is designed to diagnose how well recovered instruction meaning transfers into the actions an agent actually executes. The work targets grounded execution in embodied language agents rather than reference resolution alone.

papersTODAY 04:00 UTC

GAVEL: LLM Judge Protocol for Comparing Extracted Clinical Timelines Against Case Reports

Researchers introduce GAVEL, a protocol that uses a large language model as a judge to compare two clinical timelines extracted from case reports, rather than relying on a single expert reference annotation. The approach aims to address limitations in existing extraction pipelines, where evaluation is constrained by imperfect reference labels and imprecise event alignment. The work is described in a new arXiv preprint.