papersSEP 12 04:00 UTC
COBRA-Skills: Bandit-Guided Evolution for LLM Agent Skill Optimization
A new arXiv preprint proposes COBRA-Skills, a method that applies contextual bandit guidance to evolve reusable skills for large language model agents. The approach aims to cut the reliance on expensive execution-based evaluation and large task datasets that limit existing skill optimization techniques. It targets agents that reuse skills distilled from earlier task experience.