LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#llm-as-a-judge

4 curated events
papersTODAY 04:00 UTC

Study questions LLM-as-a-judge validity for psychological depth evaluations

A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.

papersTODAY 04:00 UTC

Paper Extends Condorcet's Jury Theorem to Panels of AI Advisers

A new arXiv paper examines how Condorcet's jury theorem applies when the same question is posed to several AI models, as happens in self-consistency sampling and LLM-as-a-judge setups. The theorem holds that adding independent, competent voters makes a majority more reliable, but the author argues this breaks down for AI advisers. The work introduces a latent-dimension framing to characterize when aggregating multiple model outputs actually improves accuracy.

papersTODAY 04:00 UTC

Study Questions Whether Consistent Local LLM Judges Match Human Ratings

A new arXiv paper examines the use of local large language models as automated judges of other models' outputs, a practice meant to cut the cost and time of human evaluation. The authors investigate whether a judge that returns stable, consistent scores can still be unreliable when compared against human ratings. The work argues that consistency alone is not sufficient evidence of a trustworthy evaluator.

papersSEP 10 04:00 UTC

Study Links LLM-as-a-Judge Scoring Inconsistency to Internal Judge Circuits

A new arXiv paper investigates why the same large language model gives systematically different verdicts when serving as an automated evaluator, depending on the required output format such as a 1-5 rating versus a true/false label. The authors trace these discrepancies to specific internal 'judge circuits' within the model, providing a mechanistic account of the phenomenon. The work aims to improve the reliability of LLM-based evaluation pipelines across different output formats.