Skip to content

AI Agents That Matter (Kapoor et al., 2024)

Reference

Citation: Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., Narayanan, A. "AI Agents That Matter." arXiv:2407.01502 (2024). Type: paper. Link: arxiv.org/abs/2407.01502.

What it is

An analysis of how agent benchmarks are reported. It shows that evaluating accuracy alone, without cost, rewards needlessly complex and expensive agents, and that cost-controlled evaluation lets much simpler baselines match state-of-the-art accuracy.

Role in the record

  • Grounds BP01: once cost is measured, complex agents often do not beat simple baselines, so an agent is not the default choice for a task.

Atom-level for/against detail and quotes are in the provenance data (assets/provenance.yml), keyed by practice atom.

Discussion