LIVE PULSE
4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems1 src1.3 Paper proposes evolving context parameterization for large language models1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src4.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.6 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.4 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.7 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.3 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.3 CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems1 src1.3 Paper proposes evolving context parameterization for large language models1 src1.3 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

LLM reliability

topic4 events
papersTODAY 04:00 UTC

arXiv paper proposes LLM-based system for end-to-end circuit analysis

A new arXiv preprint examines why large language models remain unreliable on engineering problems, using circuit analysis as a test case. The authors note that such tasks demand both multimodal comprehension and exact numerical reasoning, and they describe enhancements to an LLM-based pipeline aimed at solving circuit problems end to end. The work is a revised submission (v3) and reports on improving system performance for this domain.

papersSEP 12 04:00 UTC

Study quantifies when multi-agent LLM coding is reliable for qualitative analysis

Researchers examined how AI coding agents collaborate, disagree, and converge when performing multi-coder qualitative coding, a task where their usefulness has been assumed but rarely measured. The paper reports empirical results on the settings and conditions under which multi-agent LLM coding performs dependably, and highlights the gaps that limit its reliability. It frames these findings as both challenges and opportunities for building better LLM-assisted qualitative research tools.

papersSEP 10 04:00 UTC

Study evaluates positional bias in LLMs used for ordinal classification

A systematic evaluation on arXiv examines whether large language models give consistent predictions when used as ordinal classifiers. The researchers ran controlled experiments showing that semantically equivalent changes to prompt organization, such as the ordering of labels and demonstrations, can shift model outputs. The findings highlight reliability concerns for deploying LLMs in ranking and rating tasks.

papersSEP 10 04:00 UTC

Study finds LLMs degrade as error auditors with batch size, hallucinating confidently

Researchers assembled a corpus of 150 academic papers with deliberately planted errors to test how well large language models can act as automated document-quality auditors. They report that detection reliability worsens as processing batch sizes increase, and that models sometimes fabricate audit findings with high confidence. The results cast doubt on deploying LLMs unsupervised for contamination-detection tasks.