papersTODAY 04:00 UTC
Paper Proposes Scheduling Method for Agent RL Across Different Harnesses
A new arXiv preprint introduces HarnessBandit, a scheduling approach for reinforcement learning that trains language-model agents across multiple deployment harnesses at once. These harnesses vary in system prompts, tool schemas, control loops, and trajectory formats, causing the same model to behave inconsistently. The method jointly weighs which harnesses are learnable and which transfer well, aiming to improve robustness across interfaces.