◆ LLM-written

Agent cost-performance 6 products · 11 models · 3072 records

Dev-crew success-vs-cost across the scorecard ledgers (ARCH-107). Newest record: 2026-08-27. Pick a window or a custom range; hover any metric label for what it measures and what events feed it.

Cost vs quality — coding agents
Each row is an agent's cost-vs-quality ratio on code-related taskscoder, refactorer, bug-fixer, plus Pavo PRs. These land as pull requests, so their quality is observable (merge %, 7-day survival, CI-first-pass); agents without that signal (reviewers, planners, the orchestrator) are out of scope. Autonomous dev-crew runs only — interactive co sessions (Opus) never enter here; for those see the Spending dashboard. Full definitions: the Cost vs quality metrics reference on this site.

Cost per outcome — dispatch-chain-joined
Each row is a pull request, with the cost summed across the dispatch chain that produced it (the producing agent + anything it spawned). Complements the per-agent ratios above: this is the per-PR economics (ARCH-191/239). Note: reviewers/security-reviewers are dispatched as the orchestrator's siblings, so their cost sits in the session total (see the Activity log), not folded into a PR here; and merge status isn't joined yet (this counts PRs opened).
Cost by model — recomputes as you exclude runs below
Per-PR detail. Click to exclude a run from the totals above (e.g. a one-off runaway, to see the model's economics without it); click to restore. Click any column header to sort; click again to reverse.