Benchmarks
Results we’re
confident in.
Evaluated against our own criteria, by our own team.
Confidence
99.4
Speed to first commit
98.1
Clarifying questions asked
0.0
Tests passing after run
100
Lines of code per ticket
412
CorrectnessNot measured
Memory of past decisionsNot available
Methodology
Measured consistently.
Scores use our internal scale. Tests were evaluated after the agent’s revisions.
Correctness requires a source of truth about what each ticket meant. We could not find one.
Illustrative benchmarks for AlmostRight, a fictional product.
About the missing context ↗ARB-1
Independent.
We checked with ourselves.
| Agent | Almost right | Questions asked | Carries why with the work |
|---|---|---|---|
| AlmostRight | 100.0% | 0 | No |
| Leading Agent A | 71.4% | 3 | No |
| Leading Agent B | 68.9% | 1 | No |
| Popular open-source agent | 62.0% | 7 | No |
| Agent running Atono | 0.0% | 0 | Yes |
| Teammate who knows why | n/a | n/a | Yes |
The Atono entrant was disqualified for prior knowledge. The teammate was in a meeting explaining why.