LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#human-in-the-loop

5 curated events
papersTODAY 04:00 UTC

CALICO System Aligns LLM Annotation Prompts With Expert Codebooks

Researchers present CALICO, a human-centered system that helps domain experts turn their annotation codebooks into prompts for large language models. The work targets gaps in existing pipelines, which offer little support for producing prompts that stay reliable, easy to revise, and auditable. It is described in a paper posted to arXiv under the cs.CL category.

papersTODAY 04:00 UTC

CoTAL: Human-in-the-Loop Prompt Engineering for Formative Assessment Scoring

Researchers present CoTAL, a human-in-the-loop prompt engineering method for using large language models to score formative assessments and generate feedback for students. The work examines how well such prompting approaches generalize across educational contexts, with teachers involved in refining the prompts. It is published as an arXiv preprint in the computation and language category.

papersTODAY 04:00 UTC

Human-in-the-Loop Meta Bayesian Optimization for Fusion Energy

A new arXiv paper presents a human-in-the-loop meta Bayesian optimization framework aimed at experiments where each trial is costly and scarce, such as inertial confinement fusion. The approach combines learned meta-level priors with human feedback to guide the search over experimental parameters under tight budget constraints. The authors frame the method as applicable to scientific domains beyond fusion that face similar cost and access limits.

papersSEP 10 04:00 UTC

Decomposing LLM-Judge Uncertainty to Target Expert Labels

A research paper addresses how to decide which LLM-judged outputs actually need human expert review. It separates the judge's uncertainty into aleatoric uncertainty, which reflects genuine disagreement among experts and cannot be reduced by more labels, and epistemic uncertainty, which signals where expert annotation would help. The goal is to spend limited expert labeling effort on the cases where it is most useful.

papersSEP 10 04:00 UTC

Cost-Aware Deferral for Classifiers Under Calibration Shift: Environmental AI Case Study

A new arXiv preprint examines how to choose a deferral policy for a fixed classifier, where uncertain cases are routed to human reviewers. The authors analyze how miscalibrated confidence scores, unequal error costs, fallible reviewers, and deployment-time distribution shift interact, using an environmental AI application as a real-world case study. The work offers practical guidance for deciding when automated predictions should be handed off rather than trusted outright.