LIVE PULSE
5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.1 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.8 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.6 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.3 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 Fine-Tuning Vision-Language Models with Listener Gaze for Referring Expressions1 src1.4 Perceptual Reality Transformer Explores What Illustrations Must Preserve1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

hallucination

topic6 events
papersTODAY 04:00 UTC

Survey Reviews Causes, Corrections and Evaluation of Hallucination in Multimodal AI Models

A revised arXiv paper surveys research on hallucination in multimodal foundation models, focusing on large vision-language models. It organizes the literature around why these errors occur, methods proposed to reduce them, and how they are measured. The authors frame the work as a structured overview connecting causes, corrections and evaluation practices.

papersTODAY 04:00 UTC

Study Finds Chemical Chain-of-Thought in Reasoning Models Prone to Hallucination

A new arXiv paper examines how language models trained for chemical reasoning use chain-of-thought steps, and finds that the intermediate reasoning frequently contains fabricated content. Testing four reasoning model families across twelve chemistry tasks, the authors report that hallucination is widespread and largely disconnected from the final answer. The work suggests chain-of-thought traces in this domain act more like an unreliable scratchpad than a faithful record of the model's reasoning.

papersTODAY 04:00 UTC

UniCAR-RL targets fine-grained perception failures in multimodal math reasoning

A new arXiv paper introduces UniCAR-RL, a reinforcement learning approach aimed at improving how multimodal large language models handle math problems involving diagrams and figures. The authors argue that weak fine-grained visual perception leads models to hallucinate details early, which then causes errors to compound through the rest of the reasoning chain. The method is framed as improving perception before deeper reasoning steps are attempted.

papersSEP 12 04:00 UTC

Paper Proposes Adaptive Perturbation Selection for Contrastive Audio Decoding

A new arXiv paper addresses hallucination in large audio-language models, where models sometimes let language priors override what is actually heard in the audio. The authors propose a method that adaptively selects perturbations for contrastive decoding, a training-free approach, arguing that existing techniques rely on crude perturbations such as masking or added noise. The work aims to improve how reliably these models ground their outputs in acoustic evidence.

papersSEP 12 04:00 UTC

arXiv Paper Probes Knowledge Attribution to Distinguish Hallucination Types in LLMs

A new arXiv preprint proposes probing methods to trace where large language models source their knowledge, aiming to separate two kinds of hallucination. The authors distinguish faithfulness violations, where a model mishandles context it was given, from factuality violations, where its answers stem from incorrect stored knowledge. The work targets better attribution of model outputs to internal knowledge versus supplied context.

papersSEP 10 04:00 UTC

New arXiv Study Links LLM Faithfulness to Input Data Plausibility

A newly posted computational linguistics paper examines whether large language models become less faithful to a provided context when the input data seems implausible. The authors analyse how a model's tendency to hallucinate or misinterpret facts varies with the plausibility of what it is given, a question with direct consequences for retrieval-augmented generation and data-to-text systems. The work seeks to clarify when models follow supplied evidence versus falling back on their own priors.