arXiv 2607.02032

PACE: A Proxy for Agentic Capability Evaluation

By Yueqi Song, Lintang Sutawika, et al.

Published 2026-07-02

Mindmap

Browse the paper's core ideas, clusters, and relationships in a structured outline.

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benc…

View the original paper on arXiv