papersSEP 10 04:00 UTC
BRACE paper proposes anchored Bellman-residual correction for stale critics in asynchronous RL
A new arXiv preprint introduces BRACE, a method aimed at value-function staleness in asynchronous reinforcement learning. As training of language models increasingly relies on asynchronous setups, delays between acting and learning bias the critic toward outdated policies, while prior asynchronous-training fixes targeted only the actor. The proposed approach applies an anchored Bellman-residual correction to keep the critic aligned with the current policy.
arXivBRACEBellman-residual correctionactor-criticasynchronous reinforcement learningvalue-function staleness
COVERAGE · 2 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.AIBRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL ↗SEP 10 04:00 UTC
arXiv cs.LGBRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL ↗SEP 10 04:00 UTC