arXiv 2607.02032
PACE: A Proxy for Agentic Capability Evaluation
By Yueqi Song, Lintang Sutawika, et al.
Published 2026-07-02
Wiki summary
Explore the paper's summary, context, and related research on Papiers.
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benc…