arXiv 2607.02032
PACE: A Proxy for Agentic Capability Evaluation
By Yueqi Song, Lintang Sutawika, et al.
Published 2026-07-02
Mindmap
Browse the paper's core ideas, clusters, and relationships in a structured outline.
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benc…