LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

LLM evaluation

topic37 events
papersTODAY 04:00 UTC

CLQT benchmark targets diagnostic evaluation of LLM portfolio-management agents

A new arXiv paper introduces CLQT, a closed-loop, cost-aware and strategy-consistent benchmark for evaluating LLM agents that manage investment portfolios. The authors argue that ranking agents by returns over a fixed window fails to show whether their process is sound or their performance durable, and propose a diagnostic alternative. The work is cross-listed in cs.AI and cs.LG as a replacement submission.

papersTODAY 04:00 UTC

Study Finds Rubrics Can Be Exploited to Shift LLM Judge Preferences

A new arXiv paper identifies a vulnerability in evaluation pipelines that use LLM-based judges guided by natural-language rubrics. The authors show that rubrics can serve as an attack surface, allowing subtle preference drift in judge behavior that may go unnoticed by standard benchmarks. The work highlights the need for more robust validation of rubric-driven evaluation and alignment setups.

papersTODAY 04:00 UTC

IBBench-Light benchmark tests whether models treat external records as instructions or text

A new arXiv paper introduces IBBench-Light, an evaluation that presents the same external record to a model under two different uses: as a procedure the model must carry out, or as text it must simply read. Each of twelve semantic bases produces 144 matched response pairs per model, and four quantized instruction-tuned models were tested. The paired setup is meant to isolate whether models react to a directive's form or to the user's stated task.

papersTODAY 04:00 UTC

Crypto Accounting Bench tests LLMs on reconstructing crypto transaction entries

Researchers released Crypto Accounting Bench, a set of 118 evaluation items that ask language models to rebuild the full journal entry an organization recorded for a crypto-asset transaction. The benchmark is applied to both frontier and open-weight models to gauge how well they handle this specialized accounting task. It appears on arXiv under cs.AI and cs.CL.

papersTODAY 04:00 UTC

arXiv Paper Proposes Reward-Based Outcome Evaluation for Grammatical Error Correction

A new arXiv preprint argues that grammatical error correction systems are typically judged by how closely their edits or outputs match reference corrections, which can unfairly penalize valid alternative rewrites. The author proposes evaluating GEC by the outcome instead, using reward-based scoring that judges whether the resulting text is fluent and correct rather than counting edits. This reduces reliance on gold references and aims to better reflect real-world usefulness.

papersTODAY 04:00 UTC

arXiv paper extends adaptive testing to continuous-score LLM evaluation

A revised arXiv paper proposes an adaptive evaluation method that applies computerized adaptive testing ideas to generation tasks, where model outputs receive continuous scores instead of binary or multiple-choice marks. The approach aims to reduce the number of items needed while maintaining confident ranking of models. It targets LLM benchmarking beyond traditional multiple-choice setups.

papersTODAY 04:00 UTC

DynSTEER: Dynamic Stage-wise Evaluation and Review for LLM Agents

A new arXiv paper introduces DynSTEER, a framework for evaluating large language model agents that operate over long-horizon tasks. It targets gaps in existing evaluation methods, which typically judge only final outcomes and struggle to pinpoint where errors occur. The approach combines stage-wise trajectory assessment with review during execution rather than after the fact.

papersTODAY 04:00 UTC

arXiv Paper Audits Commonsense Reasoning Benchmarks Used for LLM Evaluation

A new arXiv paper argues that commonsense reasoning in language models is typically measured with multiple-choice benchmarks such as HellaSwag and PIQA, yet these benchmarks themselves are rarely scrutinized. The authors propose a more comprehensive approach to evaluating the benchmarks, questioning how well they actually capture the capability they claim to test. The work is framed as a meta-evaluation of standard commonsense reasoning tests.

papersTODAY 04:00 UTC

Study Examines How Unnecessary Tool Access Affects LLM Answers

A new arXiv paper investigates how giving large language models access to external tools they do not actually need changes the way they answer questions. The authors find that mere availability of tools can shift model behavior even when no external information is required. The work suggests tool provisioning should be matched to the task rather than offered by default.

papersTODAY 04:00 UTC

Study finds biomedical reference generation unreliable across 26 LLMs

Researchers tested 26 large language models from eight developers on their ability to produce accurate biomedical citations. The authors report that fabricated or incorrect references remain a persistent problem across the models evaluated. The work is a preprint and characterises the frequency of this failure mode rather than proposing a fix.

papersTODAY 04:00 UTC

Paper Argues ASR Transcripts Are a Flawed Yardstick for Audio-LLM Tasks

A research paper examines how speech and audio LLMs are typically evaluated, namely by testing whether a waveform prompt outperforms an automatic speech recognition transcript. The authors argue that for closed-set, known tasks this setup mixes together two distinct things: whether the model actually used acoustic information and whether it simply needed the task spelled out. They propose auditing generative audio calls as a way to separate those factors.

papersTODAY 04:00 UTC

arXiv paper tests staged prompts across six frontier AI models

A newly posted arXiv paper describes experiments in which the same three-part prompt sequence was run ten times for each of six frontier AI models from OpenAI, Anthropic, xAI and Google DeepMind. The prompts move from asking about architectural preferences toward a fuller task, suggesting the study compares how different systems respond as questioning gets more demanding. The abstract is truncated, so final findings and conclusions are not yet visible.

modelsTODAY 04:00 UTC

MameLoshnLM: First Open-Source 8B Language Model for Yiddish Introduced

Researchers released MameLoshnLM, described as the first open-source 8-billion-parameter language model dedicated to Yiddish. The work also includes an evaluation benchmark intended to fill the gap in reliable testing resources for the language. It addresses the low digital availability of Yiddish text despite its substantial written heritage.

papersTODAY 04:00 UTC

arXiv Paper Tests Reliability of LLM Factuality Metrics Using Answer Perturbation

A new arXiv study examines whether the metrics used to judge large language models' factual accuracy are themselves dependable. The authors probe how sensitive these evaluation methods are by perturbing model answers and observing whether scores change as expected. They argue that current factuality benchmarks need stronger validation before their results are trusted.

papersTODAY 04:00 UTC

Study questions LLM-as-a-judge validity for psychological depth evaluations

A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.

papersTODAY 04:00 UTC

PortBench: Correlation-Aware Benchmark for LLM Portfolio Management

Researchers introduce PortBench, a benchmark for evaluating large language models on portfolio management tasks. It addresses gaps in prior benchmarks by covering multiple asset classes and accounting for cross-asset correlations across the full pipeline. The work aims to give a more realistic measure of LLM performance in financial portfolio settings.

papersTODAY 04:00 UTC

Deliberative Diagnostic Framework for Evaluating LLM Opinion Simulation

A new arXiv paper proposes a diagnostic method for testing whether large language models genuinely reason about newly presented information or simply reproduce opinions they absorbed during training. The authors argue this distinction matters for "silicon sampling," where LLM personas stand in for human respondents in public-opinion research. Existing evaluations, they note, only check whether a simulated persona gives a plausible answer rather than how it handles fresh evidence.

papersTODAY 04:00 UTC

PhysMent benchmark evaluates LLM physics reasoning through interactive experiments

Researchers introduced PhysMent, a benchmark designed to test how well large language models reason about physical systems by running experiments rather than answering static questions. The work argues that strong scores on existing science benchmarks do not show whether models can actively probe the physical world. The abstract notes that this ability remains poorly understood.

papersTODAY 04:00 UTC

arXiv primer surveys evaluation methods for LLMs in healthcare

A new arXiv paper reviews how large language models used in clinical and medical settings should be assessed. It argues that evaluating these systems is harder than conventional machine learning evaluation for a variety of reasons. The work is framed as an introductory guide to evaluation approaches for healthcare LLMs.

papersSEP 12 04:00 UTC

NovGauge Benchmark Targets LLM Weakness in Judging Paper Novelty

A new arXiv preprint introduces NovGauge, a benchmark designed to test how well large language models assess the novelty of research papers. Unlike earlier benchmarks that reduce novelty to a single overall score, it breaks the task into separate dimensions so researchers can pinpoint where a model fails. The work is motivated by the growing use of LLMs in peer review at major AI conferences, where novelty judgments remain unreliable.

papersSEP 12 04:00 UTC

Study Examines How Anonymizing Input Data Affects Large Language Model Performance

A new arXiv paper investigates how removing personally identifiable information from inputs changes the usefulness of large language models. The authors note that anonymization is now common practice in sensitive deployments, but its effect on model performance has not been thoroughly characterized. The work aims to clarify the trade-off between privacy protection and model utility.

papersSEP 12 04:00 UTC

Study Frames LLM Political Stance as Context-Dependent, Not Fixed

A new arXiv paper argues that a language model's political leanings are better described as a probability distribution that shifts with the prompt and surrounding context rather than a single stable viewpoint. The authors report empirical tests across nine current LLMs to support this framing of ideology as conditional on context. The work is positioned as a measurement approach for studying political behavior in models.

papersSEP 11 04:00 UTC

Study Measures Reliability of Automated Jailbreak Evaluators

A new arXiv paper examines how well automated evaluators judge whether jailbreak attacks on language models succeed, noting that human expert review is expensive and hard to scale. The authors argue that jailbreak research often fails to properly validate the evaluators it depends on, and they empirically measure those evaluators' behavior.

papersSEP 11 04:00 UTC

Study Finds LLM Simulators Can Circumvent Automated Explanation Tests

A new arXiv paper examines automated simulatability, a protocol that scores explanations by how well they let a user predict a model's outputs without relying on costly human evaluation. The authors report that when LLMs stand in for human explainees, they can bypass the explanations themselves, undermining the validity of the metric. The work suggests automated simulatability may overstate how useful an explanation really is.

papersSEP 10 04:00 UTC

KernelGenBench Tests Whether LLMs and Agents Can Write Efficient Kernels Across Hardware

Researchers introduced KernelGenBench, a benchmark that evaluates how well large language models and agentic systems can produce specialized accelerator kernels. The benchmark assesses code generation across diverse operator sources and hardware platforms, aiming to fill a gap left by earlier evaluations of kernel-writing capability.

papersSEP 10 04:00 UTC

Formal context operations and evaluation criteria for LLM use in systems engineering design

A newly listed arXiv preprint lays out a framework of formal operations for building the context supplied to large language models when they generate architecture models in engineering design. The authors also define criteria for judging whether LLM-produced outputs are fit for systems engineering tasks. The work aims to make generative AI a dependable accelerator in design workflows.

papersSEP 10 04:00 UTC

Study finds LLMs degrade as error auditors with batch size, hallucinating confidently

Researchers assembled a corpus of 150 academic papers with deliberately planted errors to test how well large language models can act as automated document-quality auditors. They report that detection reliability worsens as processing batch sizes increase, and that models sometimes fabricate audit findings with high confidence. The results cast doubt on deploying LLMs unsupervised for contamination-detection tasks.

papersSEP 10 04:00 UTC

GANDR Introduces Claim-Level Auditing for Verifiable Legal LLM Answers

Researchers have proposed GANDR, an approach that audits legal question-answering systems at the level of individual claims instead of scoring whole answers. Each statement generated by the model is checked against the specific source it cites, so readers can verify a grounded answer piece by piece. The work targets a shortcoming of existing grounded-generation pipelines, which typically evaluate answers only in aggregate.

papersSEP 10 04:00 UTC

Paper shows fixed-rollout pass@k evaluations identify only limited information

A new paper examines the common practice of extrapolating pass@k benchmark results to attempt counts larger than the number of samples actually collected per problem. Under a pooled conditional-Binomial model, the authors show that success counts from fixed-size rollouts determine only a finite number of distribution moments. The result implies such evaluations cannot fully characterize model performance well beyond the sampled regime.

papersSEP 10 04:00 UTC

Researchers propose grounded evaluation and repair for LLM-generated PDDL planning problems

A new arXiv paper examines how large language models convert natural-language planning descriptions into PDDL problem instances, arguing that common checks like syntactic validity or planner success can overstate actual quality. The authors introduce an evaluation and repair framework that grounds assessment more firmly in the underlying planning task to better catch and fix flawed outputs.

papersSEP 10 04:00 UTC

Divergence-based approach proposed to evaluate fidelity loss in quantized LLMs

A new arXiv paper argues that zero-shot task accuracy is an inadequate yardstick for quantized large language models, because it relies only on argmax predictions and hides changes in output distributions. The authors introduce a divergence-based method for measuring how much behavioral fidelity is lost when models undergo aggressive post-training compression for memory-constrained edge devices.

papersSEP 10 04:00 UTC

Psychometric audit finds MMLU aggregate scores mainly measure factual retrieval, not reasoning

A new arXiv paper applies psychometric methods to the MMLU benchmark, analyzing how question difficulty is distributed across its aggregate score. The authors conclude that the headline number primarily reflects a model's ability to recall facts, providing limited signal about reasoning skill. The results caution against relying on MMLU alone as a measure of general AI capability.

papersSEP 10 04:00 UTC

Decomposing LLM-Judge Uncertainty to Target Expert Labels

A research paper addresses how to decide which LLM-judged outputs actually need human expert review. It separates the judge's uncertainty into aleatoric uncertainty, which reflects genuine disagreement among experts and cannot be reduced by more labels, and epistemic uncertainty, which signals where expert annotation would help. The goal is to spend limited expert labeling effort on the cases where it is most useful.

papersSEP 10 04:00 UTC

Study Links LLM-as-a-Judge Scoring Inconsistency to Internal Judge Circuits

A new arXiv paper investigates why the same large language model gives systematically different verdicts when serving as an automated evaluator, depending on the required output format such as a 1-5 rating versus a true/false label. The authors trace these discrepancies to specific internal 'judge circuits' within the model, providing a mechanistic account of the phenomenon. The work aims to improve the reliability of LLM-based evaluation pipelines across different output formats.

papersSEP 10 04:00 UTC

SWORD benchmark probes how consistently LLMs reject false facts across languages

Researchers present SWORD, a benchmark that systematically distorts facts from Wikidata and tests whether large language models notice the resulting errors in different languages. Their experiments reveal that models frequently fail to reject distorted statements consistently across languages, even when they perform well on standard multilingual question-answering benchmarks. The findings suggest existing evaluations can overstate a model's genuine factual understanding outside English.

papersSEP 10 04:00 UTC

ActTraitBench: benchmark quantifies the knowledge-decision gap in LLM persona behavior

Researchers introduced ActTraitBench, a benchmark that checks whether large language models actually behave in line with the personas they describe in explicit self-reports. Using human-grounded behavioral validation, it measures a knowledge-decision gap between what models state and the choices they make implicitly. The arXiv posting is a revised v2 of the paper.

papersSEP 10 04:00 UTC

YallaMorph benchmark evaluates Arabic morphological generation in LLMs

Researchers have released YallaMorph, a benchmark for measuring how well large language models generate morphologically accurate Arabic. It addresses a gap in current Arabic evaluation, which focuses on downstream tasks rather than directly testing whether models can control grammatical forms like inflection and derivation. The work highlights that producing fluent Arabic text does not guarantee correct morphosyntactic output.