LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

LLM-as-a-judge

topic18 events
papersTODAY 04:00 UTC

GAVEL: LLM Judge Protocol for Comparing Extracted Clinical Timelines Against Case Reports

Researchers introduce GAVEL, a protocol that uses a large language model as a judge to compare two clinical timelines extracted from case reports, rather than relying on a single expert reference annotation. The approach aims to address limitations in existing extraction pipelines, where evaluation is constrained by imperfect reference labels and imprecise event alignment. The work is described in a new arXiv preprint.

papersTODAY 04:00 UTC

IROH Retrieval System Leads JOKER 2026 Humor Ranking Task

A team called VANGUARD describes IROH, a three-stage retrieval pipeline that blends sparse and dense retrieval with cross-encoder reranking and LLM judges distilled from rationales. The system placed first on the English JOKER 2026 Task 1 leaderboard with a MAP score of 0.6347. The work appears as an arXiv paper under the cs.CL category.

papersTODAY 04:00 UTC

Paper Argues LLM-Judge Calibration in Biomedical ML Needs Four Separate Ledgers

A new arXiv paper examines how synthetic perturbations are often used as cheap calibration data for LLM evaluators in biomedical machine learning, where expert review is limited. The authors argue that a planted mutation key should not be treated as either a detector output or automatically as human ground truth. They propose formalizing four distinct ledgers to make reporting of calibration results more responsible.

papersTODAY 04:00 UTC

Paper Extends Condorcet's Jury Theorem to Panels of AI Advisers

A new arXiv paper examines how Condorcet's jury theorem applies when the same question is posed to several AI models, as happens in self-consistency sampling and LLM-as-a-judge setups. The theorem holds that adding independent, competent voters makes a majority more reliable, but the author argues this breaks down for AI advisers. The work introduces a latent-dimension framing to characterize when aggregating multiple model outputs actually improves accuracy.

papersTODAY 04:00 UTC

Study questions LLM-as-a-judge validity for psychological depth evaluations

A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.

papersTODAY 04:00 UTC

Study Questions Whether Consistent Local LLM Judges Match Human Ratings

A new arXiv paper examines the use of local large language models as automated judges of other models' outputs, a practice meant to cut the cost and time of human evaluation. The authors investigate whether a judge that returns stable, consistent scores can still be unreliable when compared against human ratings. The work argues that consistency alone is not sufficient evidence of a trustworthy evaluator.

papersTODAY 04:00 UTC

Study Examines Reliability of LLM Judges for Patent-Drafting Agents

Researchers introduce Vibe Patenting, a testbed that evaluates whether LLM judges can reliably assess AI agents performing professional patent-drafting work. The work probes how dependable automated evaluation is when applied to complex, specialized tasks rather than general benchmarks. It highlights open questions about using LLMs as evaluators in high-stakes professional settings.

papersTODAY 04:00 UTC

Study Finds LLM Judges Underuse Non-Directional Verdicts Allowed by Task Contracts

A new arXiv paper examines how large language models act as judges in evidence-based fact verification, converting supporting material into final verdicts. The authors report that even when task instructions explicitly permit non-directional outcomes such as "Conflicting" or "Not Enough Evidence," models tend to favor directional verdicts instead. The work suggests a mismatch between stated judging criteria and the labels models actually produce.

papersTODAY 04:00 UTC

Study Finds Rubrics Can Be Exploited to Shift LLM Judge Preferences

A new arXiv paper identifies a vulnerability in evaluation pipelines that use LLM-based judges guided by natural-language rubrics. The authors show that rubrics can serve as an attack surface, allowing subtle preference drift in judge behavior that may go unnoticed by standard benchmarks. The work highlights the need for more robust validation of rubric-driven evaluation and alignment setups.

papersYESTERDAY 16:29 UTC

Hacker News thread debates whether agreement between LLM judges signals reliability

A Hacker News discussion examines the practice of using one large language model to grade another's output, and asks whether consensus among several such judges actually indicates a correct verdict. Commenters raise concerns that models can share the same blind spots or biases, so agreement may reflect correlated error rather than genuine quality. The thread touches on how evaluation setups should be validated, for example against human raters or adversarial examples.

papersSEP 12 04:00 UTC

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression

A new arXiv paper examines whether large language models can serve as post-hoc reviewers that judge whether equations produced by symbolic regression are physiologically sensible. The study uses genetic programming and grammatical evolution to derive mathematical expressions from multivariate data, then has clinicians evaluate the LLM assessments. It is a case study rather than a benchmark, focusing on plausibility screening alongside predictive accuracy.

papersSEP 11 04:00 UTC

LLM-as-a-Judge Framework for Agentic AI in Drug Discovery Aligned With Human Raters

A new arXiv paper addresses the difficulty of scoring open-ended, tool-using LLM agents in chemistry and drug discovery, where conventional benchmarks fall short. The authors propose an evaluation system built on the LLM-as-a-Judge approach and tune it against human expert judgments to improve reliability. The work aims to make automated assessment of agentic scientific workflows more trustworthy.

papersSEP 10 04:00 UTC

Reference-based method audits LLM bias via relative representations of hidden states

An arXiv paper in cs.AI introduces a technique for auditing bias in large language models by analyzing internal hidden states instead of relying on generated outputs. By comparing a model's representations against those of a reference model using relative representations, the approach aims to detect internal bias shifts that output-based benchmarks or judge models could miss. The authors frame it as a cheaper alternative to benchmark-heavy or judge-dependent auditing pipelines.

papersSEP 10 04:00 UTC

Decomposing LLM-Judge Uncertainty to Target Expert Labels

A research paper addresses how to decide which LLM-judged outputs actually need human expert review. It separates the judge's uncertainty into aleatoric uncertainty, which reflects genuine disagreement among experts and cannot be reduced by more labels, and epistemic uncertainty, which signals where expert annotation would help. The goal is to spend limited expert labeling effort on the cases where it is most useful.

papersSEP 10 04:00 UTC

Evidence-Grounded Text Evaluation with LLM Judges Aims to Make Rubric Scoring Reliable

A research paper on arXiv introduces a method for scoring text against evaluation rubrics using large language models, addressing how black-box judge models can apply identical criteria in inconsistent ways. The approach ties each score to concrete evidence drawn from the evaluated text, making the reasoning behind judgments easier to audit and reproduce. The work is cross-listed under arXiv categories for artificial intelligence, computational linguistics, and machine learning.

papersSEP 10 04:00 UTC

Study Reveals Position Bias in Rubric-Based LLM-as-a-Judge Evaluations

A new arXiv paper examines large language models acting as evaluators under rubric-based protocols, a setting that has received less attention than pointwise and pairwise comparison methods. The authors find that the ordering of responses systematically influences the judge's verdicts, exposing position bias in this evaluation setup. The results suggest that pipelines relying on LLM judges may need safeguards or reordering strategies to produce reliable assessments.

papersSEP 10 04:00 UTC

Study Links LLM-as-a-Judge Scoring Inconsistency to Internal Judge Circuits

A new arXiv paper investigates why the same large language model gives systematically different verdicts when serving as an automated evaluator, depending on the required output format such as a 1-5 rating versus a true/false label. The authors trace these discrepancies to specific internal 'judge circuits' within the model, providing a mechanistic account of the phenomenon. The work aims to improve the reliability of LLM-based evaluation pipelines across different output formats.

papersSEP 10 04:00 UTC

XAI-Arena: Testing whether LLMs can judge the quality of explainable AI explanations

A new arXiv paper introduces XAI-Arena, a study of whether large language models can reliably evaluate explanations produced by explainable AI methods. The authors note that current evaluation relies heavily on subjective human judgment, which hurts reproducibility, scalability, and comparability across studies. The work explores automated, LLM-based assessment as a potential alternative to manual expert reviews.