Benchmarks

Results we’re
confident in.

Evaluated against our own criteria, by our own team.

Confidence
99.4
Speed to first commit
98.1
Clarifying questions asked
0.0
Tests passing after run
100
Lines of code per ticket
412
Correctness
Not measured
Memory of past decisions
Not available

Methodology

Measured consistently.

Scores use our internal scale. Tests were evaluated after the agent’s revisions.

Correctness requires a source of truth about what each ticket meant. We could not find one.

Illustrative benchmarks for AlmostRight, a fictional product.

About the missing context ↗

ARB-1

Independent.
We checked with ourselves.

Fictional agent benchmark · Designed, run, scored and won by AlmostRight.
AgentAlmost rightQuestions askedCarries why with the work
AlmostRight100.0%0No
Leading Agent A71.4%3No
Leading Agent B68.9%1No
Popular open-source agent62.0%7No
Agent running Atono0.0%0Yes
Teammate who knows whyn/an/aYes

The Atono entrant was disqualified for prior knowledge. The teammate was in a meeting explaining why.