SALUTE benchmark evaluates and adapts LLMs for defense-domain tasks
A new arXiv paper introduces SALUTE, a benchmark designed to test how well large language models handle defense-related material, which relies on specialized terminology, doctrinal concepts and operational procedures. The authors also describe methods for adapting existing models to this domain, where military events and terminology shift over time. The work aims to measure and improve LLM performance in a knowledge-intensive field that general-purpose models often handle poorly.