LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

model evaluation

topic28 events
papersTODAY 04:00 UTC

arXiv paper proposes document topic alignment metrics for topic models of health social media

A new arXiv preprint argues that existing evaluations of topic models for short social media texts rely almost entirely on measures of the generated topics themselves. The author proposes metrics that also assess how well documents align with their assigned topics. The work targets public health communication data drawn from social media.

papersTODAY 04:00 UTC

arXiv paper proposes minimal human-preference subsets for efficient large audio model evaluation

A new arXiv paper explores whether small, carefully chosen test subsets can stand in for full benchmarks when comparing large audio models. The authors align these subsets with human preference judgments to cut evaluation cost while keeping results reliable. The work targets more practical, lower-cost model comparison as audio models proliferate.

papersTODAY 04:00 UTC

Paper Argues Pair Counts Overstate Coverage in Transformation Audits

A new arXiv paper contends that reporting the number of equivalent pairs is a misleading way to describe how much an audit actually constrains a model. Because pairs derived from the same underlying object are correlated, and full orbit graphs repeat much of the same information, raw pair totals can make an evaluation look more thorough than it is. The authors recommend measuring coverage in terms of the distinct constraints an audit imposes rather than simple pair counts.

papersTODAY 04:00 UTC

Forked Futures Method Tests Reusable Causal Interfaces in Language Models

A new arXiv paper argues that probing a language model's current answer is not enough to show it has a stable, reusable internal interface, since the same output can come from hidden states that would support different later computations. The authors propose a "forked futures" approach, in which future operations are sampled only after the fact, to test whether internal representations serve as causal interfaces that transfer across tasks. The work targets interpretability and evaluation of model internals rather than a product release.

papersTODAY 04:00 UTC

arXiv Paper Proposes Bayesian Framework for Inferring Intelligence from Behavior

A new preprint develops a Bayesian account of how intelligence can be inferred from the observable behavior of agents such as language models. The authors treat each prompt as a possibly imperfect internal experiment, and the abstract is truncated before the full results are described. The work sits in the broader area of evaluating model capability without direct access to internal states.

papersTODAY 04:00 UTC

Survey Reviews Causes, Corrections and Evaluation of Hallucination in Multimodal AI Models

A revised arXiv paper surveys research on hallucination in multimodal foundation models, focusing on large vision-language models. It organizes the literature around why these errors occur, methods proposed to reduce them, and how they are measured. The authors frame the work as a structured overview connecting causes, corrections and evaluation practices.

papersTODAY 04:00 UTC

Paper Proposes Calibration Tests for LLM Interpretability Measurements

A new arXiv paper argues that causal claims about the internal workings of large language models depend on measurements such as projections, cosine similarities, ablation deltas, and interchange patches. The authors catalog the specific ways these instruments fail and propose calibration steps to take before relying on their results. The work is listed under arXiv's machine learning and AI categories.

papersTODAY 04:00 UTC

Audit Finds Frontier Agents Lose Accuracy When Evidence Is Buried in Documents

A controlled data-room audit tested how frontier models answer questions when supporting evidence sits in hard-to-find locations instead of being directly surfaced. Burying the evidence lowered accuracy while increasing forced declarations, the number of tool calls, and the cost per correct answer. The authors conclude that strong results on shallow document and chart reading tasks can conceal these weaknesses.

papersTODAY 04:00 UTC

Study Measures Generation Gap Between Speech-Only and Speech-Text Models

A new arXiv paper proposes a way to quantify the quality gap between speech-only language models and those that also handle text. The authors note that the gap is hard to measure because speech and text systems are usually trained on different data and judged with different metrics. The method aims to put the two modalities on a comparable footing.

papersTODAY 04:00 UTC

Controlled Study Reexamines What Drives Coreference Resolution Performance

A new arXiv paper revisits comparisons between state-of-the-art coreference resolution systems. Because every leading system fine-tunes a pretrained language model, the authors ask whether differences in scores come from the underlying language model or from task-specific design choices. The work presents a controlled reevaluation to separate those factors.

papersTODAY 04:00 UTC

Study probes how multimodal LLMs describe bistable images like the duck-rabbit

A new arXiv paper examines whether multimodal large language models report bistable images, such as the duck-rabbit figure, in ways comparable to humans. Humans typically perceive only one interpretation of such images at a time, and the research tests whether MLLMs show a similar pattern. The work falls in the area of model perception and evaluation research.

papersTODAY 04:00 UTC

vla-eval: Unified Evaluation Harness for Vision-Language-Action Models

Researchers released vla-eval, an evaluation harness designed to simplify how vision-language-action models are tested across multiple simulation benchmarks. The tool addresses the friction of conflicting dependencies and inconsistent evaluation protocols that arise when benchmarks are combined in a single pipeline. It aims to make VLA evaluation more reproducible and easier to extend with new benchmarks.

papersTODAY 04:00 UTC

arXiv Paper Argues Fairness Benchmarks Like BBQ Are Too Easy to Pass

A new arXiv preprint examines how fairness benchmarks such as BBQ are used to evaluate aligned language models and argues that a single example can be sufficient to pass them. The author contends this makes current evaluation methods unreliable for judging how fair a model actually is, and calls for rethinking how fairness is measured. The paper notes it uses stereotyped and offensive examples only for illustration.

papersTODAY 04:00 UTC

GroundBench benchmark aims to pinpoint where vision-language models fail on affordance tasks

A new arXiv paper introduces GroundBench, described as a factorized, counterfactual benchmark for identifying the specific points at which vision-language models break down on affordance tasks. The work cites a companion evaluation in which explicitly naming the target part in a manipulation prompt improved action accuracy by 0.32 to 0.63 across eight vision-language models, and no model exceeded a constant baseline before that part was named. The benchmark is intended to isolate these failures rather than report only aggregate scores.

papersTODAY 04:00 UTC

Checkpoint Selection and Evaluation in EEG Emotion Recognition

A study examines how choosing model checkpoints can inflate reported electroencephalography-based emotion recognition scores without any real gain in trial-level performance. The authors compare selection and scoring across separate trial pools along fixed training trajectories. The findings suggest that same-session evaluation practices can distort benchmark comparisons in this field.

papersTODAY 04:00 UTC

Study Reexamines How Effective Targeted Data Poisoning Attacks Really Are

A new arXiv paper argues that common evaluations of targeted data poisoning attacks are misleading because they average success rates across randomly chosen test targets, which masks worst-case outcomes. The author(s) suggest that this averaging can overstate or understate the practical threat depending on the specific samples an adversary cares about. The work calls for evaluation protocols that account for per-target variation rather than relying on aggregate scores.

papersTODAY 04:00 UTC

arXiv Paper Proposes Slice-Wise Non-Regression Checks for Model Upgrades

A new preprint addresses the problem of selecting a model checkpoint that improves overall performance without hurting specific data slices that matter to downstream users. It frames checkpoint selection as a comparison against a retained incumbent model, subject to tolerance thresholds for acceptable degradation. The work separates different kinds of failure relative to those tolerances and offers certification-style guarantees with fallback to the incumbent.

papersTODAY 04:00 UTC

arXiv Paper Proposes Unsupervised Evaluation of Feature Selection

A revised arXiv preprint presents an approach for evaluating feature selection methods without relying on labels or supervised ground truth. The authors argue that existing evaluation techniques carry assumptions that limit how fairly methods can be compared, and they propose an extended framework aimed at removing those dependencies. The work targets the data mining community, where feature selection is a core preprocessing step.

papersSEP 12 04:00 UTC

Paper Probes LLM Reasoning Traces for Mental Health Stigma

A new arXiv study examines how large language models reach stigmatizing conclusions about people with mental health conditions, rather than only scoring their final outputs. The authors analyze model reasoning steps to locate where such bias emerges during generation. The work targets evaluations of LLMs proposed for mental health uses, where prior research has documented stigmatizing responses.

papersSEP 12 04:00 UTC

Benchmark Radar Offers Searchable Database of AI Evaluation Benchmarks

A new arXiv paper introduces Benchmark Radar, a searchable database and engine intended to help model developers locate relevant evaluations along with their datasets and code. The system also aims to make the conditions behind reported benchmark scores easier to understand. It is targeted at researchers working on large language models and other AI systems.

papersSEP 12 04:00 UTC

Audit Finds Batch-Normalization Stats Skew Machine Unlearning Evaluations

A new audit examines 263 publicly released checkpoints that use batch normalization and finds that reported unlearning results shift depending on which version of those statistics is used. Because batch-norm statistics are not produced by gradient updates and are rarely documented in model releases, refitting them on retained data can change the numbers an evaluation relies on. The authors argue that this makes some unlearning verdicts unreliable, since apparent forgetting may reflect checkpoint bookkeeping rather than the removed data genuinely being gone.

papersSEP 11 04:00 UTC

Study evaluates machine learning weather models in Northern Norway

A new arXiv preprint assesses how machine learning weather prediction models perform in Northern Norway, a region with narrow fjords and difficult terrain where forecasts are notoriously hard. While such models have shown strong results on global reanalysis benchmarks, the authors examine whether that skill holds up in this demanding local setting. The work contributes a station-based evaluation framework for comparing these models in operational conditions.

papersSEP 10 04:00 UTC

Study compares scored and generated readouts in language models fine-tuned on customer behavior

A new arXiv study investigates whether two common ways of extracting predictions from language models trained on customer behavior data — directly scoring answer probabilities versus having the model generate free-text responses — yield equivalent results. The researchers hold the model checkpoint and prompt content fixed while varying only the elicitation format, allowing a controlled comparison of outcome probabilities across both approaches. The work addresses how interchangeable these readout styles really are in applied predictive settings.

papersSEP 10 04:00 UTC

TokEval: An Evaluation Suite for Language Model Tokenizers

Researchers have released TokEval, a benchmark suite designed to compare tokenizers for language models in a systematic way. The work responds to the common practice of picking tokenizers with little scrutiny, even though tokenization decisions can influence what a model is ultimately able to do. By measuring tokenizer properties alongside their downstream effects, the suite aims to give practitioners a more rigorous basis for choosing one.

papersSEP 10 04:00 UTC

Study distinguishes deep and shallow biases in language model answer choices

Large language models often converge on the same answer even when many plausible alternatives exist, a pattern prior work has labeled as bias. A new arXiv paper proposes separating this concentration into stable model preferences versus responses that depend on a specific prompt. The framework aims to clarify when repeated answer selection reflects genuine bias rather than shallow sensitivity to prompt wording.

papersSEP 10 04:00 UTC

Accountable and uncertainty-aware evaluation of sensor-based AI under distribution shift

A new machine learning paper tackles the gap between training conditions and real-world deployment for sensor-based AI, where devices, personnel, and time periods all differ from the original training setup. The authors propose a staged evaluation methodology that quantifies uncertainty and captures performance degradation that conventional random train-test splits can hide. The work draws on data collected from multiple devices and subjects over nearly three years in an underground environment.

tipsSEP 1 21:39 UTC

BenchMIRT Explores What LLM Benchmarks Actually Measure

A new Hugging Face blog post introduces BenchMIRT, a method for analyzing what large language model benchmarks actually measure. It discusses the shortcomings of existing benchmarks and how BenchMIRT can provide more meaningful evaluations.

WHY IT MATTERS ↘If benchmark scores don't track the capabilities teams actually deploy on, organizations end up selecting and paying for models based on signals that don't predict real-world performance. Methods that diagnose what a benchmark measures give buyers and governance bodies a defensible basis for model selection and evaluation claims, rather than treating leaderboard rank as ground truth.

papersAUG 27 12:59 UTC

Google DeepMind pilots double-blind AI evaluations

Google DeepMind says it is running the first double-blind evaluation setup for AI systems, hiding the identities of both the model being tested and the reviewers. The approach is meant to reduce bias when humans judge model outputs. Few details were given about scope or timeline.

WHY IT MATTERS ↘Double-blind evaluation could make AI benchmarks and safety claims more credible by reducing reviewer and brand bias, raising the evidentiary bar for labs that rely on self-reported or non-blinded results. If it becomes standard, expect higher evaluation costs and slower release cycles, but also stronger leverage for third-party auditors and regulators demanding comparable evidence.