AI Agents That Matter (Kapoor et al., 2024)¶
Reference
Citation: Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., Narayanan, A. "AI Agents That Matter." arXiv:2407.01502 (2024). Type: paper. Link: arxiv.org/abs/2407.01502.
What it is¶
An analysis of how agent benchmarks are reported. It shows that evaluating accuracy alone, without cost, rewards needlessly complex and expensive agents, and that cost-controlled evaluation lets much simpler baselines match state-of-the-art accuracy.
Role in the record¶
- Grounds BP01: once cost is measured, complex agents often do not beat simple baselines, so an agent is not the default choice for a task.
Atom-level for/against detail and quotes are in the provenance data
(assets/provenance.yml), keyed by practice atom.