Agent cost-performance 6 products · 11 models · 3072 records
Dev-crew success-vs-cost across the scorecard ledgers (ARCH-107). Newest record: 2026-08-27. Pick a window or a custom range; hover any metric label for what it measures and what events feed it.
Cost vs quality — coding agents
Each row is an agent's cost-vs-quality ratio on code-related tasks — coder, refactorer, bug-fixer, plus Pavo PRs. These land as pull requests, so their quality is observable (merge %, 7-day survival, CI-first-pass); agents without that signal (reviewers, planners, the orchestrator) are out of scope. Autonomous dev-crew runs only — interactive co sessions (Opus) never enter here; for those see the Spending dashboard. Full definitions: the Cost vs quality metrics reference on this site.
→
This table covers autonomous dev-crew runs + Pavo PRs only — interactive co sessions (which run on Opus) never enter it. For total spend including interactive Opus, see the Spending dashboard. The "—" row is Pavo PRs, whose ledger records no model yet.
No records match these filters.
Cost per outcome — dispatch-chain-joined
Each row is a pull request, with the cost summed across the dispatch chain that produced it (the producing agent + anything it spawned). Complements the per-agent ratios above: this is the per-PR economics (ARCH-191/239). Note: reviewers/security-reviewers are dispatched as the orchestrator's siblings, so their cost sits in the session total (see the Activity log), not folded into a PR here; and merge status isn't joined yet (this counts PRs opened).
Cost by model — recomputes as you exclude runs below
Per-PR detail. Click ✕ to exclude a run from the totals above (e.g. a one-off runaway, to see the model's economics without it); click ↩ to restore. Click any column header to sort; click again to reverse.