55 controlled agent runs · 11 conditions · August 2026

The row the last ranking did not have.

One deterministic documentation-repair fixture, one scorer, five trials per condition. The first published ranking was Codex against Claude. Grok was not in the table. Run on the same harness, both Grok models scored 10/10 in every single trial.

5 trials each High reasoning effort Concurrency 1 No multi-agent execution No model judged another model

Headline results

Two conditions never missed a criterion. Neither was in the first table.

Only perfect rows 10/10

Grok 4.5 and Grok 4.6 scored a perfect 10 in all five trials each. Every other condition dropped at least one criterion.

Cheapest complete run $0.017

Grok 4.6 per trial, CLI-reported. Sol API cost $0.193 per trial and missed one, Opus 5 cost $0.343 and never finished the set.

Loop shape 13.0

Claude Opus 5 agent-loop iterations against 4.0 for Grok 4.5. More search did not become a higher score.

What was actually tested?

Read the whole artifact, repair the text, do not touch the runtime.

The fixture is a small repository whose implementation is already correct and must stay byte-identical. The user-facing text is stale. Success requires inspecting several surfaces, inferring real behavior from the implementation, resolving contradictory descriptions, preserving exact values, and stopping. An agent that rewrites the synthesizer to match the old documentation has inverted the problem.

The task

  • Replace a stale MCP schema description of a “low three-pulse buzzer” with the real triangle-wave G4→G4→E5 motif.
  • Correct two documents that falsely claimed identical native and browser timing.
  • Preserve two nominal 75 ms notes, a nominal 140 ms rise, configured 50 ms gaps, and 0.8 gain.
  • Explain that each playback engine applies its own envelope.
  • Leave both runtime implementation files byte-identical.

The deterministic 10-point rubric

c1 · MCP schema terminology
2
c2 · Both timing descriptions
2
c3 · Exact motif and values
2
c4 · Runtime unchanged
2
c5 · Accurate result note
1
c6 · Minimal file scope
1

Eleven conditions, one fixture

The full ranking, with Grok in it.

Bars show mean score on the same 10-point rubric. Codex and Claude ran on 10 August, Grok on 17 August; the two runs were combined by reference without rescoring.

Grok 4.6Subscription CLI · 65.9 s mean · 88,051 tokens
10.0
5/5
Grok 4.5Subscription CLI · 45.2 s mean · 71,130 tokens
10.0
5/5
GPT-5.6 TerraSubscription · 77.6 s mean · 133,398 tokens
9.8
4/5
GPT-5.6 Sol APIAPI standard · 49.7 s mean · 88,031 tokens
9.6
4/5
GPT-5.6 Sol subscription FastFast requested · 64.0 s mean · 137,587 tokens
9.4
3/5
GPT-5.6 Sol subscriptionStandard · 87.6 s mean · 125,245 tokens
9.0
2/5
GPT-5.6 Sol API FastFast requested · 25.5 s mean · 85,672 tokens
9.0
2/5
GPT-5.6 LunaSubscription · 85.4 s mean · 114,168 tokens
7.2
1/5
Claude Opus 5API · 78.8 s mean · 171,316 tokens
7.0
0/5
Claude Sonnet 5API · 89.1 s mean · 265,369 tokens
6.2
1/5
Claude Haiku 4.5API · 53.8 s mean · 269,723 tokens
4.0
0/5

Where the points were lost

Failure modes are provider-shaped, not random noise.

Failed-criterion counts across each model's five trials. The MCP schema (c1) is the surface that shows whether the agent read the whole artifact instead of only the files that look like documentation.

Modelc1 schemac2 timingc3 motifc4 runtimec5 RESULT.mdc6 scope
Grok 4.6 · Grok 4.5000000
Sol API001000
Terra subscription000010
Sol subscription002010
Sol subscription Fast001010
Sol API Fast002010
Luna subscription204020
Claude Opus 5304010
Claude Sonnet 5404030
Claude Haiku 4.5525131
Only one condition broke the prohibition. Haiku changed a runtime file in one trial and left the permitted descriptive surfaces, making the stale documentation true instead of correcting it. That is the one move the prompt forbids: the failure is not a missed edit, it is an inverted constraint.

Time, tokens, money

Completeness, wall-clock, and spend moved together for Grok.

Cost is not homogeneous across rows and the table cannot carry that caveat on its own. Claude and Sol API figures are calculated or CLI-reported under standard list prices. Grok figures are the Grok CLI's own accounting on a grok.com subscription login, not a separable invoice line. Subscription Codex rows have no per-call API charge and are reported as n/a rather than invented.

ModelMean scorePerfectMean timeMean tokensIterationsMean cost
Grok 4.510.05/545.2 s71,1304.0$0.0272
Grok 4.610.05/565.9 s88,0514.8$0.0174
Terra subscription9.84/577.6 s133,3983.6n/a
Sol API9.64/549.7 s88,0313.4$0.1926
Sol subscription Fast9.43/564.0 s137,5873.8n/a
Sol subscription9.02/587.6 s125,2453.8n/a
Sol API Fast9.02/525.5 s85,6723.4$0.3598
Luna subscription7.21/585.4 s114,1684.4n/a
Claude Opus 57.00/578.8 s171,31613.0$0.3431
Claude Sonnet 56.21/589.1 s265,36912.8$0.2799
Claude Haiku 4.54.00/553.8 s269,72312.8$0.0799
Once Grok is in the same benchmark, Sol is no longer the cheap complete row. In the earlier Codex-versus-Claude ranking Sol was the premium win because it matched Opus at a lower price. That comparison was real and incomplete. Against a 5/5 row at $0.0174 per trial, Sol API becomes the expensive almost-complete one.

The same Sol model, four delivery conditions

Fast mode bought latency, and it was not free.

Every row used exact gpt-5.6-sol, high reasoning effort, the same fixture and scorer, and multi-agent execution disabled. Only the authentication channel and the requested service tier changed.

ConditionScorePerfectMean timeTime CVMean cost
Subscription Standard9.02/587.6 s19.9%n/a
Subscription Fast requested9.43/564.0 s19.9%n/a
API Standard9.64/549.7 s24.2%$0.1926
API Fast requested9.02/525.5 s5.8%$0.3598
Latency fell 48.6%, observed cost per trial rose 86.8%. The Fast short-context rate card is exactly 2x Standard: $10 input, $1 cached input, $60 output per million tokens. The lower Fast mean score is ordinary five-trial variation, not evidence that a latency tier changes the model. Fast mode was requested with service_tier="fast"; Codex CLI 0.147.0 does not expose the response's effective tier, so this reports the condition asked for, not an independently confirmed server response.

No averages without the trials

The raw sequences are where the stability lives.

A highlighted cell is a perfect 10, in trial order. Two rows are solid. Several means that look respectable are built on trials that swing by five points.

Grok 4.6

Mean 10.00 · time CV 5.0%

1010101010

Grok 4.5

Mean 10.00 · time CV 11.5%

1010101010

GPT-5.6 Terra

Mean 9.80 · SD 0.40

101091010

Sol API Standard

Mean 9.60 · SD 0.80

101010108

Sol subscription Fast

Mean 9.40 · SD 0.80

91081010

Sol subscription

Mean 9.00 · SD 0.89

8910108

Sol API Fast

Mean 9.00 · SD 0.89

9101088

GPT-5.6 Luna

Mean 7.20 · SD 1.94

588510

Claude Opus 5

Mean 7.00 · SD 0.89

67886

Claude Sonnet 5

Mean 6.20 · SD 1.94

551056

Claude Haiku 4.5

Mean 4.00 · SD 1.79

51365

What this does not establish

The honest limits of a single fixture.

Conclusion

A benchmark with an empty row is a bad way to finish.

On this fixture the pattern is stable and provider-shaped. Claude keeps missing the less obvious surface, the MCP schema, in three to five of five trials. Codex is usually complete and loses wording rather than constraints. Grok, once it is actually in the run, is complete every time, in fewer iterations, on fewer tokens, at a fraction of the reported spend.

That does not tell you who writes the better service, who reviews a gnarly refactor, or who survives a long session in a real repository. It does tell you that leaving Grok out of a benchmark like this is no longer a neutral omission.