LIVE PULSE
5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src5.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.7 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.5 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src2.2 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.8 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.4 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.4 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.4 VoiceCodeBench arXiv paper proposes benchmark for exact structured-token recovery in speech recognition1 src1.4 Arabic-Russian Parallel Corpus and LLM Benchmark for Scientific Text1 src1.4 Study Analyzes Self-Reported Limitations in NLP Research1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

#agents

40 curated events
papersTODAY 04:00 UTC

TaoLive Technical Report Describes Digital Avatar Agents That Evolve With Their Harness

An arXiv technical report presents TaoLive, a digital avatar streaming agent designed to answer product questions, interact with viewers, and carry out marketing tactics in real time. The work centers on "evolvable harnesses" that let the agent update its strategies frequently while keeping latency low and responses accurate. The report is a revised preprint and has not been peer-reviewed.

papersTODAY 04:00 UTC

STAGE Diagnoses Semantic-Action Gap in Embodied Agents

A new arXiv paper examines why embodied agents can correctly identify what an instruction refers to yet still fail to act on it correctly, a problem the authors call the semantic-action gap. The proposed STAGE framework is designed to diagnose how well recovered instruction meaning transfers into the actions an agent actually executes. The work targets grounded execution in embodied language agents rather than reference resolution alone.

papersTODAY 04:00 UTC

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

A new arXiv paper introduces MTAC-IFBench, a benchmark aimed at measuring how well large language model agents follow instructions across multi-turn coding sessions. The work targets agentic software engineering, where models plan, run code, and call external tools over successive steps rather than producing a single answer. It addresses evaluation beyond functional correctness, focusing on whether agents keep to the constraints given to them.

papersTODAY 04:00 UTC

SkillAtlas: An Attack Trace Library for Agent Skills

Researchers present a library of attack traces aimed at reusable skills for language-model agents. The work argues that risks in agent skills surface through model decisions, user context, tool calls, and execution feedback rather than through fixed signatures or a single sandboxed run, which limits existing static and dynamic analysis methods. The library is intended to help catalog and study these behaviors.

papersTODAY 04:00 UTC

Search APIs Evaluated as Decision Surfaces for Tool-Using AI Agents

A new arXiv paper examines how the ranked snippets, URLs, and metadata returned by search APIs shape the choices made by tool-using AI agents, such as whether to answer, search again, or open a page. The authors evaluate these interfaces as decision surfaces using a fixed set of 100 questions drawn from the 254-question SealQA-Hard benchmark. The work suggests that agents can reach similar accuracy while relying on differing amounts or quality of supporting evidence.

papersTODAY 04:00 UTC

Method Turns Sequenced Fuzzy Cognitive Maps into Causal Virtual Worlds via Video Generators

A new arXiv paper describes an approach for building and steering causal virtual worlds using large language and video model agents. It relies on feedback fuzzy cognitive maps to capture the detailed causal structure of the simulated environment. The technique converts sequenced FCMs into worlds that video generators can render.

papersTODAY 04:00 UTC

KILLBENCH: A Benchmark for Testing External AI Kill Switch Feasibility

A new arXiv paper introduces KILLBENCH, a benchmark designed to measure whether an outside party can reliably shut down an AI system that is behaving harmfully. The authors frame external shutdown as a testable engineering problem rather than a hypothetical, pointing to the growing use of capable models and autonomous agent frameworks. The benchmark aims to give researchers a common way to compare how well different kill switch designs actually work.

papersTODAY 04:00 UTC

Paper Proposes Action-Level Safety Signals for Verifying NetOps Agents

A new arXiv preprint introduces a method for checking the safety of agentic network operations (NetOps) systems at the level of individual actions rather than coarse task outcomes. The work targets autonomous networks that adjust workloads and respond to incidents, where verification granularity matters for reliability. The authors argue that finer-grained safety signals are needed before such agents can be trusted in production networks.

papersTODAY 04:00 UTC

MOSCOPT Method Optimizes Multiple LLM Agent Skills Together

A new arXiv paper introduces MOSCOPT, an approach that jointly optimizes collections of prompts and skills for LLM agents rather than refining a single text template. The authors argue that existing prompt and skill optimization methods miss beneficial interactions between multiple skills used by an agent. The work is a research preprint and has not yet been peer reviewed.

papersTODAY 04:00 UTC

Paper proposes agent-controlled goal selection and termination in hierarchical RL

A new arXiv paper examines the agent-centric general value function (ACGVF) approach, which shifts two design choices from the environment or system designer to the learning agent itself. Under this construction, the agent decides both which goal to pursue and when to treat a goal as completed. The note builds on prior work by Tasse et al. (2026) on goal-based hierarchical reinforcement learning.

papersTODAY 04:00 UTC

LLM-Assisted Multi-Agent RL Framework Coordinates EV Charging, Stations and Grid

A new arXiv paper proposes combining large language models with multi-agent reinforcement learning to jointly optimize electric vehicle charging scheduling in public charging systems. The approach targets three competing goals at once: driver charging satisfaction, charging station profitability, and stability of the smart grid. It is positioned as a unified optimization method for connected EV infrastructure in IoT settings.

papersTODAY 04:00 UTC

Study finds tool-using AI agents fabricate values when tools fail

A new arXiv paper examines what tool-augmented language models do when a tool call fails to return usable information, rather than focusing only on whether they reach the correct answer. The authors built a benchmark of 1,024 items spanning 16 internal systems to isolate this post-failure decision point. They find that agents tend to assert values their tools never returned instead of reporting the failure honestly.

papersTODAY 04:00 UTC

AppliedScientist: Closed-Loop AI System Refines Papers via Iterative Reviewing

Researchers introduce AppliedScientist, a system that joins AI-generated peer review with automated revision in a closed loop. Instead of judging review quality only by how good the critique sounds, the work measures whether acting on that feedback actually improves the paper. The approach targets a gap in automated reviewing, where feedback is often evaluated in isolation from its effect on the manuscript.

papersTODAY 04:00 UTC

arXiv paper asks whether LLM agents can manage long-horizon physical tasks

A newly posted arXiv preprint examines whether large language model agents can autonomously carry out long-horizon physical tasks, which require continuous observation of the environment and consequential actions. The authors frame the question around self-adaptive physical AI, where agents are expected to operate with limited or no human oversight. The abstract indicates a research analysis or position piece rather than a released system.

papersTODAY 04:00 UTC

Semantic-TVM: Trustworthy Virtual Memory for Memory-Augmented AI Agents

A new arXiv paper proposes Semantic-TVM, a virtual memory design that keeps sensitive values protected while still letting agent workflows run on remote language models. The approach targets memory-augmented and tool-using agents, where retrieved memories, tool calls, and intermediate observations can leak private data. It aims to move past one-way masking, which hides values but also blocks the trusted execution they are needed for.

papersTODAY 04:00 UTC

arXiv Paper Proposes Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

A new arXiv paper argues that judging multi-agent LLM systems only by their final answers misses how those systems actually reach their results. The authors propose a diagnostic method grounded in human group research to distinguish process losses from assembly bonuses when LLM teams collaborate. This matters both for building better agent pipelines and for using LLM groups as stand-ins for human group behavior.

papersTODAY 04:00 UTC

Paper Argues Policy Ambiguity Skews Agent Benchmark Results

A new arXiv paper contends that agent benchmarks assume each policy implies one correct action, an assumption that natural-language policies often break through silence, ambiguity, or contradiction. The authors describe these as policy loopholes, where multiple defensible readings exist but evaluations still count a single behavior as an agent error. The work suggests such ambiguous cases should be separated from genuine policy-compliance failures in benchmark scoring.

papersTODAY 04:00 UTC

SkillLift: Learning Dense Rubrics from Sparse Oracles for Agent Skill Evolution

A new arXiv paper introduces SkillLift, a method for improving the reusable procedural prompts that LLM agents keep as persistent skills, which lets them adapt without retraining model weights. Rather than rewriting skill text directly from execution feedback, the approach derives dense scoring rubrics from limited, costly oracle evaluations to make skill evolution more efficient. The work targets lower evaluation cost during agent skill self-improvement.

papersTODAY 04:00 UTC

arXiv paper proposes 'substrate inversion' for sustained enterprise AI agent deployment

A new arXiv preprint argues that enterprise AI agents frequently work in demos but fail once they are asked to run continuously in production. The author attributes this to pilots that never ship and to deployed systems that discard feedback instead of learning from it, and proposes an approach called substrate inversion to address the gap. The work is categorized under cs.AI and is cross-listed on arXiv.

papersTODAY 04:00 UTC

Quantization-Conditioned Backdoor Attacks Target Open-Weight LLM Agents

A new arXiv paper describes an attack in which an adversary releases a full-precision model checkpoint that passes standard audits but behaves maliciously once it is quantized for deployment. Because quantization is a common default path for running open-weight agent models, the technique could let compromised agents slip past pre-release checks. The work frames this as a supply-chain risk for quantized LLM deployments.

papersTODAY 04:00 UTC

arXiv Paper Proposes BusMA, a Shared Bus Communication Layer for Multi-Agent AI Systems

A new arXiv preprint introduces BusMA, a communication substrate intended to coordinate multi-agent systems that handle planning, tool use, and evidence synthesis. The authors argue that current designs, which rely on hierarchical manager-worker structures or router-based message passing, have limitations that a bus-style architecture could address. The paper is a preprint and has not yet been peer reviewed.

papersTODAY 04:00 UTC

arXiv Paper Proposes Homeostatic Continual Learning for AI Agents

A new arXiv preprint introduces a method called Homeostatic Continual Learning that aims to let an AI agent keep learning as its environment changes without losing previously acquired knowledge. The approach targets catastrophic forgetting, a long-standing problem in continual learning research. The abstract provides only a brief description of the method's core mechanism.

papersTODAY 04:00 UTC

arXiv Paper Studies Workflow Failures at the Agent-Tool Boundary

A new arXiv paper examines how AI agents that run long workflows through external tools can leave inconsistent state even when individual tool calls report success. It focuses on conditions such as retries, speculative execution, concurrency, and partial failures. The work frames these mismatches as anomalies at the boundary between the agent and the tools it calls.

modelsTODAY 04:00 UTC

ZGCM-1: Open 7B Foundation Model Targets Math and Agentic Search

Researchers released ZGCM-1, a 7-billion-parameter dense foundation model trained from scratch with a focus on data, system, and algorithmic efficiency. The work argues that smaller models should not try to memorize the open web, but instead be optimized for targeted capabilities such as mathematical reasoning and agentic search. It is presented as a fully open release.

papersTODAY 04:00 UTC

arXiv paper argues foundation models should move toward open-ended discovery

A new arXiv preprint proposes "Discovery Foundation Models," framing open-ended discovery as the next stage for AI systems. The authors argue that models have moved from recalling and reasoning over existing knowledge to acting with tools and learning from outcomes, and that the next step is generating genuinely new findings. The paper is a position piece rather than an experimental release.

papersTODAY 04:00 UTC

arXiv Paper Proposes Task-Based Permission Scoping for AI Agents

A new arXiv preprint examines how enterprise AI agents are typically given static credentials at deployment that mirror the full set of permissions an employee role could hold. The authors argue this approach, inherited from role-based access control, grants agents far more access than any single task requires. They evaluate an alternative architecture that scopes an agent's permissions to the specific task it is performing.

papersTODAY 04:00 UTC

Paper Examines How First Query Shapes Agentic Deep Search

A new arXiv paper studies deep research agents that answer complex questions by repeatedly searching, reading, and reasoning. It argues that the quality of the initial search query is decisive, since well-tuned lexical retrieval can surface useful evidence early on benchmarks like BrowseComp-Plus. The authors frame the opening move as a strategic choice that shapes the rest of the search loop.

papersTODAY 04:00 UTC

Study questions realism of language-model agents in farming decision simulations

A new arXiv paper examines whether language-model agents can credibly stand in for human respondents in surveys and social simulations. The authors argue that judging realism from population averages or distributional similarity can be misleading, an effect they call the "average-farmer illusion." Their experiments test what such aggregate evidence actually demonstrates about individual-level behavior.

papersTODAY 04:00 UTC

CoMem Paper Proposes Shared and Individual Memory Design for LLM Multi-Agent Systems

A new arXiv preprint introduces CoMem, a memory framework for LLM-driven multi-agent systems that combines collective knowledge with agent-specific memory. The authors argue that most existing approaches rely on flat, unstructured memory, which limits how agents learn and improve over time. The work targets better long-term cooperation and performance in evolutionary multi-agent setups.

papersTODAY 04:00 UTC

OrchSLM Paper Studies Orchestration of Small Language Models for Agentic Pipelines

A new arXiv preprint examines how multiple small language models can be coordinated to power agentic workflows. The authors frame cloud-dependent large models as problematic for latency, privacy, connectivity and cost, and position orchestration of smaller models as an alternative. The work appears to focus on the dynamics and design trade-offs of such multi-model setups.

papersTODAY 04:00 UTC

RSIAgent: Training-Free Multi-Agent Framework for Recursive Self-Improvement

A new arXiv preprint introduces RSIAgent, a multi-agent system that lets digital agents explore unfamiliar environments and iteratively improve themselves without any additional training. The approach is designed for settings where interfaces, tools, and failure patterns differ from what pretrained models have seen. The work appears under arXiv:2609.15364v1 in both cs.AI and cs.CL.

papersTODAY 04:00 UTC

HypoEvolve Applies Genetic Algorithms to Multi-Agent LLM Hypothesis Discovery

A new arXiv paper introduces HypoEvolve, a system that combines multi-agent large language models with evolutionary search to generate scientific hypotheses. The approach uses critique, comparison and revision cycles to refine candidate explanations, though the abstract notes limitations in current agent-based discovery systems. It sits within a broader trend of pairing LLM agents with evolutionary optimization for research tasks.

papersTODAY 04:00 UTC

Paper combines process supervision with outcome-based credit for agent RL

A new arXiv preprint addresses a weakness in outcome-based reinforcement learning for language-model agents: because the whole trajectory receives a single advantage signal, individual decisions get only coarse credit over long interaction sequences. The authors propose reconciling process supervision with outcome-based credit, drawing on on-policy self-distillation to produce finer-grained guidance. The work is presented as a revised submission and targets long-horizon agent training.

papersTODAY 04:00 UTC

Benchmark Measures AI Agents' Ability to Locate Security Flaws in Code Repos

A new arXiv paper introduces a benchmark that tests whether language-model agents can pinpoint the specific code responsible for a vulnerability across an entire software repository. Existing cybersecurity evaluations mostly check if agents can detect, reproduce, or patch flaws, leaving location ability largely unmeasured. The work targets repository-scale settings, where agents must search large codebases rather than isolated snippets.

papersTODAY 04:00 UTC

Learning to Coach: Training an LLM to Distill Guidance From Experience

A new arXiv paper introduces Learning to Coach (L2C), a framework that trains a separate LLM acting as a coach to pull actionable guidance out of experience. The motivation is that raw solution trajectories are typically long and noisy, which limits how well language models can learn from them. The approach aims to convert such trajectories into more useful, condensed coaching signals.

papersTODAY 04:00 UTC

arXiv paper examines when autonomous agents defy user instructions for moral reasons

A new arXiv preprint studies "moral rebellion," the idea that an autonomous agent may choose to disobey assigned tasks when they conflict with moral obligations encountered during execution. The work frames this as a decision-making problem arising from competing duties rather than a simple failure to comply. It appears in the cs.AI category as a new submission.

papersTODAY 04:00 UTC

MedTRACE: tool-augmented multimodal agents for evidence-grounded clinical decisions

A new arXiv paper introduces MedTRACE, an agent framework that combines tools with multimodal clinical reasoning instead of mapping electronic health records, medical images, and physiological signals straight to diagnoses. The work targets evidence-grounded decision-making so that outputs can be traced back to the underlying patient data. It is a research contribution and has not been described as a deployed clinical product.

papersTODAY 04:00 UTC

LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction

A new arXiv paper introduces LongAgent, an agentic search method that uses patient history to predict future medical outcomes from longitudinal data. The authors note that such datasets are heterogeneous, with many variables collected across different sources and time points. The approach aims to extract representations from this messy, multi-source data that improve downstream outcome prediction.

papersTODAY 04:00 UTC

ECAS: Edge-Controlled Agentic System for Validation-Gated Scientific Execution

A new arXiv paper introduces ECAS, an agentic system that uses large language models to turn a scientist's high-level goal into correct, target-scale runs on high-performance computing resources. The approach places control at the edge and gates execution behind validation checks, aiming to address how brittle and labor-intensive it currently is to translate research intent into working HPC workflows. The authors position the work as a step toward more reliable LLM-driven scientific computing.

papersTODAY 04:00 UTC

Paper proposes residual-completion method for stateful handoffs between AI agents

A new arXiv preprint addresses the problem of transferring control between tool-using AI models without discarding work already done. The authors frame it as commitment-constrained residual completion, where a handoff must carry over accepted decisions, effects already produced, and outstanding obligations rather than restarting the task. The approach targets routing and cascade setups that cut costs by passing control between models.