papersTODAY 04:00 UTC
Segment-Aware Listwise Alignment Targets Reasoning Safety in Large Reasoning Models
A new arXiv paper argues that safety alignment for large reasoning models must address two surfaces at once: the intermediate chain of thought and the final answer. The authors note that existing methods typically align whole responses, which can leave harmful reasoning steps intact even when the visible answer looks safe. Their proposed approach, segment-aware listwise alignment, treats reasoning traces and outputs as distinct segments to be optimized together.
arXivAI safety alignmentchain-of-thoughtlarge reasoning modelsreasoning safetysegment-aware listwise alignment
COVERAGE · 2 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.AIBeyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models ↗TODAY 04:00 UTC
arXiv cs.CLBeyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models ↗TODAY 04:00 UTC