LIVE PULSE
4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems1 src1.3 Paper proposes evolving context parameterization for large language models1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems1 src1.3 Paper proposes evolving context parameterization for large language models1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

LLM benchmarking

topic4 events
papersTODAY 04:00 UTC

arXiv paper extends adaptive testing to continuous-score LLM evaluation

A revised arXiv paper proposes an adaptive evaluation method that applies computerized adaptive testing ideas to generation tasks, where model outputs receive continuous scores instead of binary or multiple-choice marks. The approach aims to reduce the number of items needed while maintaining confident ranking of models. It targets LLM benchmarking beyond traditional multiple-choice setups.

papersTODAY 04:00 UTC

SALUTE benchmark evaluates and adapts LLMs for defense-domain tasks

A new arXiv paper introduces SALUTE, a benchmark designed to test how well large language models handle defense-related material, which relies on specialized terminology, doctrinal concepts and operational procedures. The authors also describe methods for adapting existing models to this domain, where military events and terminology shift over time. The work aims to measure and improve LLM performance in a knowledge-intensive field that general-purpose models often handle poorly.

papersSEP 10 04:00 UTC

MetroLLM-Bench: New Benchmark Tests LLMs as Public Transit Kiosk Assistants

Researchers have released MetroLLM-Bench, a set of 955 test cases that measures how well language models can act as the decision-making layer of a metro station kiosk. The benchmark draws on six real subway networks of varying sizes and covers eleven task types, including route planning, fare computation, and responding to service disruptions. The paper was posted to arXiv and cross-listed in the AI, computation and language, and machine learning categories.

papersSEP 10 04:00 UTC

New benchmark tests if LLMs can engineer the AI infrastructure that powers them

A new arXiv paper introduces Φ-Bench, a benchmark that measures how well large language models can help develop and optimize the computing infrastructure used to run AI systems. The authors argue that existing benchmarks do not adequately cover these infrastructure-engineering tasks, which go beyond typical code generation. The work aims to gauge whether LLMs can realistically contribute to the specialized systems engineering that underpins their own operation.