arXiv 2607.02032
PACE: A Proxy for Agentic Capability Evaluation
By Yueqi Song, Lintang Sutawika, et al.
Published 2026-07-02
Discussion
Read the public discussion and references gathered around this paper.
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benc…