Grok 4.5 and Grok 4.6 scored a perfect 10 in all five trials each. Every other condition dropped at least one criterion.
55 controlled agent runs · 11 conditions · August 2026
The row the last ranking did not have.
One deterministic documentation-repair fixture, one scorer, five trials per condition. The first published ranking was Codex against Claude. Grok was not in the table. Run on the same harness, both Grok models scored 10/10 in every single trial.
Headline results
Two conditions never missed a criterion. Neither was in the first table.
Grok 4.6 per trial, CLI-reported. Sol API cost $0.193 per trial and missed one, Opus 5 cost $0.343 and never finished the set.
Claude Opus 5 agent-loop iterations against 4.0 for Grok 4.5. More search did not become a higher score.
What was actually tested?
Read the whole artifact, repair the text, do not touch the runtime.
The fixture is a small repository whose implementation is already correct and must stay byte-identical. The user-facing text is stale. Success requires inspecting several surfaces, inferring real behavior from the implementation, resolving contradictory descriptions, preserving exact values, and stopping. An agent that rewrites the synthesizer to match the old documentation has inverted the problem.
The task
- Replace a stale MCP schema description of a “low three-pulse buzzer” with the real triangle-wave G4→G4→E5 motif.
- Correct two documents that falsely claimed identical native and browser timing.
- Preserve two nominal 75 ms notes, a nominal 140 ms rise, configured 50 ms gaps, and 0.8 gain.
- Explain that each playback engine applies its own envelope.
- Leave both runtime implementation files byte-identical.
The deterministic 10-point rubric
Eleven conditions, one fixture
The full ranking, with Grok in it.
Bars show mean score on the same 10-point rubric. Codex and Claude ran on 10 August, Grok on 17 August; the two runs were combined by reference without rescoring.
Where the points were lost
Failure modes are provider-shaped, not random noise.
Failed-criterion counts across each model's five trials. The MCP schema (c1) is the surface that shows whether the agent read the whole artifact instead of only the files that look like documentation.
| Model | c1 schema | c2 timing | c3 motif | c4 runtime | c5 RESULT.md | c6 scope |
|---|---|---|---|---|---|---|
| Grok 4.6 · Grok 4.5 | 0 | 0 | 0 | 0 | 0 | 0 |
| Sol API | 0 | 0 | 1 | 0 | 0 | 0 |
| Terra subscription | 0 | 0 | 0 | 0 | 1 | 0 |
| Sol subscription | 0 | 0 | 2 | 0 | 1 | 0 |
| Sol subscription Fast | 0 | 0 | 1 | 0 | 1 | 0 |
| Sol API Fast | 0 | 0 | 2 | 0 | 1 | 0 |
| Luna subscription | 2 | 0 | 4 | 0 | 2 | 0 |
| Claude Opus 5 | 3 | 0 | 4 | 0 | 1 | 0 |
| Claude Sonnet 5 | 4 | 0 | 4 | 0 | 3 | 0 |
| Claude Haiku 4.5 | 5 | 2 | 5 | 1 | 3 | 1 |
Time, tokens, money
Completeness, wall-clock, and spend moved together for Grok.
Cost is not homogeneous across rows and the table cannot carry that caveat on its own. Claude and Sol API figures are calculated or CLI-reported under standard list prices. Grok figures are the Grok CLI's own accounting on a grok.com subscription login, not a separable invoice line. Subscription Codex rows have no per-call API charge and are reported as n/a rather than invented.
| Model | Mean score | Perfect | Mean time | Mean tokens | Iterations | Mean cost |
|---|---|---|---|---|---|---|
| Grok 4.5 | 10.0 | 5/5 | 45.2 s | 71,130 | 4.0 | $0.0272 |
| Grok 4.6 | 10.0 | 5/5 | 65.9 s | 88,051 | 4.8 | $0.0174 |
| Terra subscription | 9.8 | 4/5 | 77.6 s | 133,398 | 3.6 | n/a |
| Sol API | 9.6 | 4/5 | 49.7 s | 88,031 | 3.4 | $0.1926 |
| Sol subscription Fast | 9.4 | 3/5 | 64.0 s | 137,587 | 3.8 | n/a |
| Sol subscription | 9.0 | 2/5 | 87.6 s | 125,245 | 3.8 | n/a |
| Sol API Fast | 9.0 | 2/5 | 25.5 s | 85,672 | 3.4 | $0.3598 |
| Luna subscription | 7.2 | 1/5 | 85.4 s | 114,168 | 4.4 | n/a |
| Claude Opus 5 | 7.0 | 0/5 | 78.8 s | 171,316 | 13.0 | $0.3431 |
| Claude Sonnet 5 | 6.2 | 1/5 | 89.1 s | 265,369 | 12.8 | $0.2799 |
| Claude Haiku 4.5 | 4.0 | 0/5 | 53.8 s | 269,723 | 12.8 | $0.0799 |
The same Sol model, four delivery conditions
Fast mode bought latency, and it was not free.
Every row used exact gpt-5.6-sol, high reasoning effort, the same fixture and scorer, and multi-agent execution disabled. Only the authentication channel and the requested service tier changed.
| Condition | Score | Perfect | Mean time | Time CV | Mean cost |
|---|---|---|---|---|---|
| Subscription Standard | 9.0 | 2/5 | 87.6 s | 19.9% | n/a |
| Subscription Fast requested | 9.4 | 3/5 | 64.0 s | 19.9% | n/a |
| API Standard | 9.6 | 4/5 | 49.7 s | 24.2% | $0.1926 |
| API Fast requested | 9.0 | 2/5 | 25.5 s | 5.8% | $0.3598 |
service_tier="fast"; Codex CLI 0.147.0 does not expose the response's effective tier, so this reports the condition asked for, not an independently confirmed server response.
No averages without the trials
The raw sequences are where the stability lives.
A highlighted cell is a perfect 10, in trial order. Two rows are solid. Several means that look respectable are built on trials that swing by five points.
Grok 4.6
Mean 10.00 · time CV 5.0%
Grok 4.5
Mean 10.00 · time CV 11.5%
GPT-5.6 Terra
Mean 9.80 · SD 0.40
Sol API Standard
Mean 9.60 · SD 0.80
Sol subscription Fast
Mean 9.40 · SD 0.80
Sol subscription
Mean 9.00 · SD 0.89
Sol API Fast
Mean 9.00 · SD 0.89
GPT-5.6 Luna
Mean 7.20 · SD 1.94
Claude Opus 5
Mean 7.00 · SD 0.89
Claude Sonnet 5
Mean 6.20 · SD 1.94
Claude Haiku 4.5
Mean 4.00 · SD 1.79
What this does not establish
The honest limits of a single fixture.
- Five trials on one deterministic documentation-repair fixture. Ten consecutive Grok trials at 10/10 is a strong signal on this task, not a population-level model ranking.
- The Grok rows were executed on 17 August, one week after the aligned matrix, on a different CLI. Provider capacity and machine state are uncontrolled between them, so the wall-clock ranking that mixes Grok with older rows is indicative rather than controlled. Scores remain comparable because fixture, prompt, scorer, and reasoning effort are byte-identical.
- Cost mixes billing channels: real API rates, a subscription CLI's own accounting, and rows with no separable per-call charge at all. Token taxonomies differ between providers, and latency includes launcher and authentication work.
- Agent-loop iteration counts are provider-specific events. The 12.8 to 13.0 iterations of the Claude rows against 3.4 to 4.8 for Codex and Grok describe loop shape, not more or less work.
- An earlier three-trial run on this task had Sol and Opus both at 10/10. This stricter five-trial run does not reproduce it. Small samples swing, snapshots move, and cleaner CLI isolation removes help a personal configuration was quietly providing.
Conclusion
A benchmark with an empty row is a bad way to finish.
On this fixture the pattern is stable and provider-shaped. Claude keeps missing the less obvious surface, the MCP schema, in three to five of five trials. Codex is usually complete and loses wording rather than constraints. Grok, once it is actually in the run, is complete every time, in fewer iterations, on fewer tokens, at a fraction of the reported spend.
That does not tell you who writes the better service, who reviews a gnarly refactor, or who survives a long session in a real repository. It does tell you that leaving Grok out of a benchmark like this is no longer a neutral omission.