LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

human evaluation

topic6 events
papersTODAY 04:00 UTC

Study Questions Whether Consistent Local LLM Judges Match Human Ratings

A new arXiv paper examines the use of local large language models as automated judges of other models' outputs, a practice meant to cut the cost and time of human evaluation. The authors investigate whether a judge that returns stable, consistent scores can still be unreliable when compared against human ratings. The work argues that consistency alone is not sufficient evidence of a trustworthy evaluator.

papersTODAY 04:00 UTC

FriendBench Benchmark Tests Whether AI Can Tell Friends From Strangers

Researchers introduced FriendBench, a benchmark that evaluates how well humans and multimodal large language models can judge whether two people in a short video clip are already acquainted or meeting for the first time. The task uses 20-second recordings of ice-breaker conversations, where cues come from behavior and body language rather than spoken content alone. The work aims to measure social perception abilities that go beyond text-based reasoning.

papersYESTERDAY 16:29 UTC

Hacker News thread debates whether agreement between LLM judges signals reliability

A Hacker News discussion examines the practice of using one large language model to grade another's output, and asks whether consensus among several such judges actually indicates a correct verdict. Commenters raise concerns that models can share the same blind spots or biases, so agreement may reflect correlated error rather than genuine quality. The thread touches on how evaluation setups should be validated, for example against human raters or adversarial examples.

papersSEP 11 04:00 UTC

Study Finds LLM Simulators Can Circumvent Automated Explanation Tests

A new arXiv paper examines automated simulatability, a protocol that scores explanations by how well they let a user predict a model's outputs without relying on costly human evaluation. The authors report that when LLMs stand in for human explainees, they can bypass the explanations themselves, undermining the validity of the metric. The work suggests automated simulatability may overstate how useful an explanation really is.

papersSEP 10 04:00 UTC

XAI-Arena: Testing whether LLMs can judge the quality of explainable AI explanations

A new arXiv paper introduces XAI-Arena, a study of whether large language models can reliably evaluate explanations produced by explainable AI methods. The authors note that current evaluation relies heavily on subjective human judgment, which hurts reproducibility, scalability, and comparability across studies. The work explores automated, LLM-based assessment as a potential alternative to manual expert reviews.

papersAUG 27 12:59 UTC

Google DeepMind pilots double-blind AI evaluations

Google DeepMind says it is running the first double-blind evaluation setup for AI systems, hiding the identities of both the model being tested and the reviewers. The approach is meant to reduce bias when humans judge model outputs. Few details were given about scope or timeline.

WHY IT MATTERS ↘Double-blind evaluation could make AI benchmarks and safety claims more credible by reducing reviewer and brand bias, raising the evidentiary bar for labs that rely on self-reported or non-blinded results. If it becomes standard, expect higher evaluation costs and slower release cycles, but also stronger leverage for third-party auditors and regulators demanding comparable evidence.