Paper Proposes Multi-Armed Bandit Approach to Compute Allocation in Self-Evolving LLMs
A new arXiv preprint examines how compute is allocated during LLM-guided evolutionary search, noting that prior work typically reports only the best result from many runs rather than the full distribution. The authors propose reframing the depth-versus-breadth tradeoff as a multi-armed bandit problem. The work is listed as a cross-list replacement across arXiv's AI and machine learning sections.