papersSEP 10 04:00 UTC
Era by Eon Benchmark provides ground-truth enterprise estate for evaluating LLM agents
The authors argue that LLM agents operating on enterprise systems of record are difficult to evaluate because production customer data cannot be used for testing and no existing substitute offers reliable ground truth. Their new benchmark addresses this by generating a synthetic enterprise environment paired with exact ground-truth labels, enabling systematic scoring of agents that use enterprise tools.