papersSEP 10 04:00 UTC
ESSA Paper Proposes Evolutionary Strategies for Scalable LLM Alignment
A new arXiv paper in machine learning introduces ESSA, which uses evolutionary strategies as an alternative to gradient-based RLHF methods like PPO and GRPO for aligning large language models. The authors argue that existing pipelines are costly because they require backpropagation through long rollouts, and their approach avoids this bottleneck. The work targets more scalable online alignment of LLMs.