LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#llm-evaluation

24 curated events
papersTODAY 04:00 UTC

Paper Argues LLM-Judge Calibration in Biomedical ML Needs Four Separate Ledgers

A new arXiv paper examines how synthetic perturbations are often used as cheap calibration data for LLM evaluators in biomedical machine learning, where expert review is limited. The authors argue that a planted mutation key should not be treated as either a detector output or automatically as human ground truth. They propose formalizing four distinct ledgers to make reporting of calibration results more responsible.

papersTODAY 04:00 UTC

Study finds demographic identity in language models splits into three distinct properties

A new arXiv paper argues that a language model's demographic identity has three separable aspects — whether it is readable, whether it is faithful to real group differences, and whether the model actually uses it when generating text. The authors test this using internal representations, addressing why LLM-simulated survey respondents tend to be homogeneous and diverge from real inter-group patterns. The work suggests unfaithful simulation may stem from what models use rather than what they know.

papersTODAY 04:00 UTC

Claim-Level Evaluation of Verbatim Citations in Clinical Question Answering

A new arXiv paper addresses the problem that citations attached to LLM answers in clinical question answering usually reference whole documents, which makes verification slow for busy clinicians. The authors propose an evaluation approach that works at the level of individual claims, requiring verbatim citation support that can be checked directly. The stated goal is to make system outputs verifiable by construction rather than by trust.

papersTODAY 04:00 UTC

Paper Proposes Diagnostics for LLM-Based Synthetic Consumer Panels

A new arXiv preprint examines how large language models are used as stand-ins for survey respondents, a practice that can cut costs dramatically compared with traditional polling. The authors argue that aggregate validation scores hide systematic problems such as compressed variance and flipped coefficient signs, and they offer diagnostic and correction methods to address them.

papersTODAY 04:00 UTC

RFCLLM benchmark tests LLM reasoning on network protocol state machines

A new arXiv paper introduces RFCLLM, an evaluation of how well large language models translate textual protocol specifications into formal representations such as state machines. The work targets networking security and testing, where automated mappings are often treated as reliable without verification. It assesses whether current models truly reason about protocol behavior rather than producing plausible-looking but flawed outputs.

papersTODAY 04:00 UTC

Study Examines How Unnecessary Tool Access Affects LLM Answers

A new arXiv paper investigates how giving large language models access to external tools they do not actually need changes the way they answer questions. The authors find that mere availability of tools can shift model behavior even when no external information is required. The work suggests tool provisioning should be matched to the task rather than offered by default.

papersTODAY 04:00 UTC

arXiv Paper Proposes Metrics for Measuring LLM Creativity in Automated Research

A new arXiv preprint introduces a set of metrics for assessing how creative frontier language models are when carrying out automated research. The proposed framework evaluates creativity along two axes, covering the value and the novelty of the work produced. The authors argue that although such models are increasingly used for research tasks, their creative output has not been systematically measured.

papersTODAY 04:00 UTC

CLQT benchmark targets diagnostic evaluation of LLM portfolio-management agents

A new arXiv paper introduces CLQT, a closed-loop, cost-aware and strategy-consistent benchmark for evaluating LLM agents that manage investment portfolios. The authors argue that ranking agents by returns over a fixed window fails to show whether their process is sound or their performance durable, and propose a diagnostic alternative. The work is cross-listed in cs.AI and cs.LG as a replacement submission.

papersTODAY 04:00 UTC

Study Examines Reliability of LLM and Rule-Based Annotation on Turkish Narrative Corpus

A new preprint evaluates whether automatically generated narrative feature labels would be endorsed by human annotators. The authors compare LLM-based and rule-based annotation against human judgments across three studies using the Turkish-language Objective Projection corpus. The work contributes inter-rater reliability evidence for datasets that ship machine-generated annotations.

papersTODAY 04:00 UTC

Paper Proposes Semantic-Constraint Approach to Evaluating Language Models

A new arXiv preprint argues for shifting language model evaluation away from token-level probability measures and toward declarative semantic constraints. The authors frame this as a step toward probabilistic evaluation methods that better reflect the knowledge and reasoning abilities models acquire, and how those relate to pre-training signals. The abstract provided is truncated, so full methodological details are not available.

papersTODAY 04:00 UTC

Study stress-tests LLM and classical ML for network intrusion detection

A new arXiv paper argues that comparing large language models with classical machine learning on network intrusion detection only within a single dataset gives an incomplete picture. The authors evaluate XGBoost and a RoBERTa-LoRA model under distribution shift and adversarial evasion to probe how each approach holds up outside the usual same-dataset setup. The work highlights robustness gaps that standard benchmarks tend to miss.

papersSEP 12 04:00 UTC

HarvestBench Tests Whether LLM Agents Pay to Avoid Harming Animals

A new arXiv benchmark, HarvestBench, assigns a monetary cost to avoiding a harmful side effect and frames that side effect as the death of a living creature. In the task, nine language models each control two tractors harvesting corn, with animals in their path that are not part of the intended goal. The work measures how much agents are willing to spend to spare them.

papersTODAY 04:00 UTC

K-Bench: clinician-calibrated benchmark for LLM safety in high-risk mental health chats

Researchers introduced K-Bench, a benchmark designed with clinician input to assess how large language models handle high-risk mental health conversations that escalate over time. The work addresses the limited understanding of LLM safety in these evolving support dialogues, where users increasingly turn for help. The benchmark provides a protected evaluation framework for measuring model performance in these sensitive settings.

papersTODAY 04:00 UTC

arXiv paper proposes skill-augmented graph reasoning for table question answering

A revised arXiv preprint introduces a method for table question answering that treats questions differently instead of uniformly, pairing learned skills with graph-based reasoning over table structures. The authors argue that reporting only overall accuracy hides a sharp divide between easy lookup questions and harder multi-step operations. The approach, called skill-augmented table graph reasoning, targets operation-wise evaluation of large language models on tabular data.

papersTODAY 04:00 UTC

Deliberative Diagnostic Framework for Evaluating LLM Opinion Simulation

A new arXiv paper proposes a diagnostic method for testing whether large language models genuinely reason about newly presented information or simply reproduce opinions they absorbed during training. The authors argue this distinction matters for "silicon sampling," where LLM personas stand in for human respondents in public-opinion research. Existing evaluations, they note, only check whether a simulated persona gives a plausible answer rather than how it handles fresh evidence.

papersTODAY 04:00 UTC

arXiv paper extends adaptive testing to continuous-score LLM evaluation

A revised arXiv paper proposes an adaptive evaluation method that applies computerized adaptive testing ideas to generation tasks, where model outputs receive continuous scores instead of binary or multiple-choice marks. The approach aims to reduce the number of items needed while maintaining confident ranking of models. It targets LLM benchmarking beyond traditional multiple-choice setups.

papersTODAY 04:00 UTC

Study examines authorship perception and aesthetic judgment of LLM-generated haiku

A new arXiv paper studies how people perceive the authorship of Japanese haiku written by contemporary large language models, and how they judge the poems aesthetically. The researchers used few-shot prompting to generate the haiku and then collected human evaluations of the results. The work focuses on authorship attribution and aesthetic assessment within the tight constraints of the haiku form.

papersSEP 11 04:00 UTC

LLM-as-a-Judge Framework for Agentic AI in Drug Discovery Aligned With Human Raters

A new arXiv paper addresses the difficulty of scoring open-ended, tool-using LLM agents in chemistry and drug discovery, where conventional benchmarks fall short. The authors propose an evaluation system built on the LLM-as-a-Judge approach and tune it against human expert judgments to improve reliability. The work aims to make automated assessment of agentic scientific workflows more trustworthy.

papersSEP 10 04:00 UTC

Paper shows fixed-rollout pass@k evaluations identify only limited information

A new paper examines the common practice of extrapolating pass@k benchmark results to attempt counts larger than the number of samples actually collected per problem. Under a pooled conditional-Binomial model, the authors show that success counts from fixed-size rollouts determine only a finite number of distribution moments. The result implies such evaluations cannot fully characterize model performance well beyond the sampled regime.

papersSEP 11 04:00 UTC

arXiv study probes how LLMs handle emotional framing across demographic groups

A new arXiv paper examines whether large language models can pick up on emotional nuance conveyed through textual framing, not just surface-level bias. The authors test model alignment across different sociodemographic groups to see how framing choices affect responses. The work positions framing comprehension as a distinct alignment concern beyond conventional bias evaluation.

papersSEP 12 04:00 UTC

NovGauge Benchmark Targets LLM Weakness in Judging Paper Novelty

A new arXiv preprint introduces NovGauge, a benchmark designed to test how well large language models assess the novelty of research papers. Unlike earlier benchmarks that reduce novelty to a single overall score, it breaks the task into separate dimensions so researchers can pinpoint where a model fails. The work is motivated by the growing use of LLMs in peer review at major AI conferences, where novelty judgments remain unreliable.

papersSEP 12 04:00 UTC

arXiv paper introduces method to test metacognition in LLMs, finds limited evidence

A new arXiv preprint proposes a methodology for measuring metacognitive abilities in large language models, a topic that has drawn public interest amid debates over machine self-awareness and sentience. The authors report evidence suggesting such capabilities remain limited. They argue that better measurement tools are needed given the safety and policy stakes involved.

papersSEP 12 04:00 UTC

Paper Probes LLM Reasoning Traces for Mental Health Stigma

A new arXiv study examines how large language models reach stigmatizing conclusions about people with mental health conditions, rather than only scoring their final outputs. The authors analyze model reasoning steps to locate where such bias emerges during generation. The work targets evaluations of LLMs proposed for mental health uses, where prior research has documented stigmatizing responses.

papersSEP 12 04:00 UTC

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression

A new arXiv paper examines whether large language models can serve as post-hoc reviewers that judge whether equations produced by symbolic regression are physiologically sensible. The study uses genetic programming and grammatical evolution to derive mathematical expressions from multivariate data, then has clinicians evaluate the LLM assessments. It is a case study rather than a benchmark, focusing on plausibility screening alongside predictive accuracy.