Claude Code vs Codex vs MiniMax Code: Who Actually Wins Terminal-Bench 4.0?

Posted by Reda Fornera on 2026-10-11
Estimated Reading Time 22 Minutes
Words 3.7k In Total

A dark code-editor screen filled with colorful programming source code — generic developer stock photo illustrating the Claude Code vs Codex comparison below, not the official Terminal-Bench 4.0 leaderboard

Claude Code vs Codex is the sharpest AI-coding-agent matchup of late 2026 — and all figures below were fetched on 2026-10-11 and cited per source. Where two boards disagree, both are reported and neither is averaged. Where a claim could not be verified against a fetched source, we say so.

Three agents, three different leaderboards

MiniMax Code is now trying to crash the party. Open the official Terminal-Bench 4.0 leaderboard as of October 11, 2026, and the podium looks like this: Claude Code driving Claude Opus 5.5 (max) at 64.8% (± 3.1%) with a full-board cost of $4.7k, Claude Code on Sonnet 5.5 (max) at 61.8% (± 2.9%) at $7.3k, and OpenAI’s Codex harness in a dead heat — GPT-6 Astra (max) and GPT-6.1 Sol (max) both at 58.2%, the Astra run costing $3.3k and the Sol run $634.24 (official Terminal-Bench 4.0 leaderboard, run dates Sep 22, Sep 28, Sep 3, and Sep 29 respectively).

Now look for the third player. MiniMax’s M2.5 has been headline news in coding-agent circles, with a Terminal Bench number circulating widely. On the official TB 4.0 board, MiniMax does not appear at all — not as a harness entry, not as a model behind anyone’s scaffold.

That absence is the entire story of this comparison. The loudest MiniMax coding-agent claim is not on a comparable board — different benchmark version, different scaffold, different sandbox, different timeout, different evidence format. Let’s walk through exactly why, then grade what actually can be graded on the official data alone.

What each agent actually is

  • Claude Code (Anthropic) is “an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools,” available in the terminal, IDE, desktop, and browser (Claude Code docs). On the board it scaffolds Anthropic’s Opus 5.5, Sonnet 5.5 and Fable 5.1 — and third-party models like Z.ai’s GLM-5.3.
  • Codex (OpenAI) is the coding agent behind the Codex CLI — “a coding agent from OpenAI that runs locally on your computer” (openai/codex README, fetched 2026-10-11) — the scaffold under every GPT-6-family entry — the OpenAI side of the Claude Code vs Codex matchup.
  • MiniMax Code is the newest entrant: its CLI was open-sourced September 18, 2026 as v0.4.12 under the MIT license, “with MiniMax or your own model” (MiniMax-AI/minimax-code README). The forensic wrinkle: the M2.5 benchmark numbers you’ve seen were not produced by MiniMax Code — they came from borrowed scaffolding, below.

The versioning trap, pinned forensically

Terminal-Bench is now a continuously maintained benchmark with semantic versioning, so a major-version bump is a rerun trigger, not a sequel: “incrementing x in (x, y, z) means rerun: agent-environment changes like new tasks, prompt fixes, data and tool modifications, or agent resource adjustments” (tbench.ai announcement, fetched 2026-10-11). TB 4.0 removed 8 tasks (2 saturated, 2 refusals, 2 public solutions, 2 unresolved quality/platform issues), fixed 19 tasks, calibrated task resources, and set “a flat agent timeout of 8 hours” on every task (same announcement).

So what was the MiniMax M2.5 number actually run on? The vendor’s model card states, verbatim:

“Terminal Bench 2: We tested Terminal Bench 2 using Claude Code 2.0.64 as the evaluation scaffolding. We modified the Dockerfiles of some problems to ensure the correctness of the problems themselves, uniformly expanded sandbox specifications to 8-core CPU and 16 GB memory, set the timeout uniformly to 7,200 seconds, and equipped each problem with a basic toolset (ps, curl, git, etc.). While not retrying on timeouts, we added a detection mechanism for empty scaffolding responses, retrying tasks whose final response was empty to handle various abnormal interruption scenarios. Final results are averaged over 4 runs.”

Read that against the official TB 4.0 methodology and the mismatch is total:

Dimension Official TB 4.0 board (Claude Code / Codex rows) MiniMax M2.5 model-card setup
Benchmark version Terminal-Bench 4.0 (66 tasks per the version history published by the Synergised analysis site; the 330-trial-per-entry figure per board transcriptions, not the HTML render) Terminal Bench 2-era
Scaffolding Claude Code (current) and Codex — each vendor’s own harness Borrowed Claude Code 2.0.64, an older release, not MiniMax Code’s own CLI
Environment Task resources calibrated via the Anthropic infrastructure-noise methodology Modified Dockerfiles for “some problems”; sandbox uniformly expanded to 8-core CPU / 16 GB memory
Timeout Flat 8 hours per task 7,200 seconds = 2 hours, per the vendor README
Run count Quoted resolution ± 95% CI from ~330 trials Averaged over 4 runs (vendor README), plus an empty-response retry mechanism
Extra retries Not described for these rows Empty-scaffolding-response retry mechanism added
Evidence form Numeric resolution, CI band, cost, token count, run date per row Terminal-Bench results presented as chart images in the README; the README’s numeric appendix table contains AIME25, GPQA-D and similar text scores but no Terminal-Bench row

Two of these are decision-grade alone: the timeout (TB 4.0 gives agents eight hours; the M2.5 run gave them two) and the sandbox (TB 4.0 calibrated resources down to realistic settings; the M2.5 run uniformly expanded to 8 cores/16 GB). Either would move an agentic score; together, plus modified tasks, they make the two numbers different measurements, not different results. That is also why any Claude Code vs Codex verdict can only be drawn one board at a time.

And the cross-version collapse is not theoretical. The Synergised analysis of the TB 4.0 reset (fetched 2026-10-11, crediting its version-history table to Capital and Compute) reports the top score at ~83% on 2.1 (89 tasks, May 2026), ~34% on 3.0 (74 tasks, July), and ~52–58% on 4.0 (66 tasks) — “three separate measurements that happen to share a name, not a rise and fall in machine capability.” A TB 2-era score of 80%+ and a TB 4.0 score of 58% overlap only accidentally.

The independent-reproduction angle is documented too. In modelscope/evalscope issue #1387 (fetched 2026-10-11), after a user re-running M2.5 on terminal_bench_v2 got 1.16%, the maintainers noted the vendor figure — 51.7% — was “只在模型卡公布,官方 leaderboard 无对应提交条目,harness 未公开”: published only on the model card, with no official-leaderboard entry and an undisclosed harness — and that “terminal-bench 对 harness 极敏感(同模型可差 10+ 个点)” (the same model can differ by 10+ points across harnesses; translation ours). Fairness requires noting the issue also traces the 1.16% to a misconfigured local setup (Docker failures, not the model), so it is not evidence against M2.5’s capability — but it directly documents that the official 51.7 has no leaderboard entry behind it. The criticism is attributed to those maintainers; the methodology facts come from MiniMax’s own README. No like-for-like TB 4.0 run of M2.5 exists, so nothing here rescales or implies an equivalent MiniMax score.

The official board: Claude Code vs Codex, head to head

Agent (scaffold) Model Reasoning effort Resolution ± CI Full-board cost Run date Tokens
Claude Code Opus 5.5 max 64.8% ± 3.1% $4.7k Sep 22, 2026 8.0B
Claude Code Sonnet 5.5 max 61.8% ± 2.9% $7.3k Sep 28, 2026 19.4B
Codex GPT-6 Astra max 58.2% ± 2.8% $3.3k Sep 3, 2026 1.5B
Codex GPT-6.1 Sol max 58.2% ± 3.1% $634.24 Sep 29, 2026 1.5B

(Source: official Terminal-Bench 4.0 leaderboard via Harbor Hub, fetched 2026-10-11. The raw board renders also list both entries at 58.18% with n_trials = 330 — a caveat on how we know that: the n_trials column and full-precision scores are not exposed in the board’s HTML render, so these exact values come from two independent transcriptions of the board (the Synergised analysis and Capital and Compute, fetched 2026-10-11), not from the render itself. The 58.2% figure quoted above is render-verified for both entries.)

On the official run, Claude Code holds the top of the board: Opus 5.5 (max) at 64.8% is three points clear of Sonnet 5.5 (max) and 6.6 clear of the best Codex entry, with CIs (±3.1 vs ±2.8) putting the Opus–Astra gap comfortably outside noise. But check the cost column: Anthropic’s entries run $4.7k and $7.3k for the full 330-trial sweep, against $3.3k for Astra and $634 for Sol.

Our derived arithmetic on the board: roughly $22 per solved task for Opus 5.5, $36 for Sonnet 5.5, $17 for Astra, $3.30 for GPT-6.1 Sol. The Synergised analysis, crediting Capital and Compute, puts the full board’s spread at $6.08–$234.24 per solved task — that site’s arithmetic, but the shape is the same: score-only rankings hide an order-of-magnitude value spread. (Provenance caveat: that spread was derived from the Sep 22 board snapshot, which predates the Sol, Opus 5.5, and Sonnet 5.5 entries — GPT-6.1 Sol’s derived $3.30 above extends the efficiency frontier below that spread’s $6.08 floor.) For anyone weighing Claude Code vs Codex on real budgets, that spread is the headline the score column hides.

A developer's terminal window with scrolling command-line code — generic tech stock photo illustrating the head-to-head scores and costs discussed in this section, not an actual render of Table A or the leaderboard

The mirror check: Artificial Analysis runs its own TB 4.0

The BenchLeader TB 4.0 mirror (data as of 11 Oct 2026), which reruns nothing and republishes Artificial Analysis’s own run, reads differently in the middle of the board:

Rank Model (effort) AA score
1 Claude Sonnet 5.5 (max) 63.6%
2 Claude Opus 5.5 (max) 59.6%
3 Claude Opus 5.5 (xhigh) 59.6%
4 GPT-6 Astra (xhigh) 59.6%
5 GPT-6 Astra (max) 59.1%
9 GPT-6.1 Sol (max) 56.1%

BenchLeader flags it explicitly: “Run by Artificial Analysis; the Terminal-Bench team’s own run is the one in the index,” and classifies the AA board as “Reference only” — it does not feed BenchLeader’s composite index.

The divergence is the lesson, not a bug. On AA’s run, Sonnet 5.5 leads Opus 5.5 — the official board’s top two in reverse, Sonnet up 1.8 points, Opus down 5.2. Two competent organizations ran the same benchmark and the ordering flipped; that is exactly why we cite each board separately and never merge or average them. So treat any single Claude Code vs Codex score as one organization’s measurement, not the matchup’s truth. The honest comparison is a range, board by board.

Same harness (Codex), same 58.2% resolution, both using 1.5B tokens across the board — but the official cost column reads $3,300 for Astra and $634.24 for Sol, a 5.2× derived gap. The CIs (±2.8 and ±3.1) overlap almost completely: on this benchmark, the two are statistically indistinguishable.

That’s the vendor’s own design. OpenAI’s pricing page (fetched 2026-10-11) lists standard short-context rates at exactly one-fifth intervals: Astra $10/M input, $50/M output; Sol $2/M input, $10/M output. The launch post adds that Sol’s cached input is $0.10/M — “95% less than standard input pricing” — and that Sol “nearly matches GPT-6 Astra’s intelligence on agentic coding” (OpenAI, “Introducing GPT-6.1 Sol”). That 5× pricing split is the core of our GPT-6.1 Sol launch analysis. The board’s economics bear the claim out: five times less per token, zero measured resolution lost. When is Astra worth it? When you need headroom above “nearly matches” — the Opus row shows another 6.6 points exist above the Astra/Sol band — or for workloads the benchmark can’t see. What the board can see: for budget agentic coding, Astra’s premium bought nothing measurable on TB 4.0.

For the model-level view behind these two boards, see our Sonnet 5.5 vs GPT-6 Sol comparison.

The pricing context: Anthropic’s menu runs from $0.10 to $50

Anthropic’s own lineup now spans the whole cost menu. Claude Haiku 5.5 (released October 7, 2026) is the first Claude priced by prompt length: $0.10/M input, $0.50/M output up to 100k prompt tokens, jumping to $0.50/$2.50 above — the platform docs confirm “a request whose prompt is over 100,000 tokens pays higher prices” (Anthropic pricing page and Haiku 5.5 docs, fetched 2026-10-11); Anthropic’s own pricing table shows the comparison directly — Haiku 4.5 at $1/$5 vs Haiku 5.5 at $0.10/$0.50 low-tier, a 90% input cut (third-party trackers such as LLMCost and aicost.tools make the same point, per the draft’s source file; we did not independently re-fetch them). Anthropic positions it for subagent work — “pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work” — while Opus 5.5 lists at $4/$20. Against OpenAI’s $2/$10 Sol row, the point is sharp: Anthropic’s cheap-model story aims at token logistics inside Claude Code workflows, not at the benchmark itself — Haiku 5.5 has no TB 4.0 entry on either fetched board. One provenance note that a careful reader will need: Anthropic’s own Haiku 5.5 announcement does publish Terminal-Bench 4.0 figures in its benchmark chart — 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5 (as reference) — but these are vendor-run, chart-only numbers with no stated run protocol in the post and no leaderboard entry behind them; they are not the official board’s independent runs (on which Sonnet 5.5 scored 61.8%), and the 70.6% vendor figure exceeds every row on both fetched boards. We treat them exactly like MiniMax’s chart figures: the same caveat stack applies to vendor-published numbers whichever vendor publishes them. Never quote an agent’s cost without knowing which model, effort setting, and cache behavior the run used.

The MiniMax Code side: what can and can’t be said

What can be said, per fetched sources:

  • MiniMax M2.5 exists; its model card claims 80.2% on SWE-Bench Verified and “It costs just $1 to run the model continuously for an hour at a rate of 100 tokens per second” (MiniMax-M2.5 README) — vendor figures on vendor-chosen setups.
  • The circulating Terminal Bench number (51.7% on TB 2, per the evalscope maintainers’ characterization) carries the full caveat stack above: TB-2-era tasks, borrowed Claude Code 2.0.64 scaffolding, modified Dockerfiles, 8-core/16 GB sandbox, 2-hour timeout, 4-run averaging — and the numeric results live only in chart images, with no Terminal-Bench row in the README’s text appendix.
  • No official-leaderboard entry backs the figure, and harness sensitivity alone moves scores by 10+ points (evalscope #1387).

What can’t be said: anything about how MiniMax Code running M2.5 on its own MIT-licensed harness would score on TB 4.0. No such entry exists on the official board as of October 11, 2026. The methodology pin is the finding; a score is unavailable. The methodological contrast with the Claude Code vs Codex rows above could not be sharper — which is exactly why no MiniMax number appears in either table.

Caveated verdict

  • Best raw score: Opus 5.5 (max) on Claude Code — 64.8% ± 3.1 at $4.7k (official board; the AA mirror tells a different story at 59.6% — hence naming the board in every verdict line).
  • Best value: GPT-6.1 Sol (max) on Codex — 58.2% ± 3.1 at $634.24, statistically tied with Astra at 5.2× less derived cost (official board).
  • Best cost-adjusted Claude option: Sonnet 5.5 — first on the AA mirror (63.6%), second on the official board (61.8%), but the priciest of the four headline rows ($7.3k, 19.4B tokens); check whether that premium survives on your cost axis.
  • MiniMax Code: unrated. No like-for-like TB 4.0 data exists. Treating its TB-2-era vendor numbers as comparable to the board would be the exact error this post exists to prevent.

Caveat box: Never merge TB 2.x, 3.x and 4.0 numbers. A Claude Code vs Codex reading is only valid within a single board. Official-board and AA-mirror figures are cited per source and never averaged. All figures as fetched 2026-10-11. Derived per-solved-task costs are arithmetic on published figures, flagged as ours. Benchmark rates are worst-case units of work — nobody pays benchmark rates for production coding.

A computer monitor displaying a business analytics dashboard — generic stock photo marking the verdict section below, not the Terminal-Bench verdict table or any live benchmark dashboard

Decision guide

Four picks, one per persona — three drawn from the official Claude Code vs Codex board, one guardrail before you trust a vendor chart:

  • Need the max ceiling, cost no object: Claude Code + Opus 5.5 (max) — 64.8% ± 3.1 on the official board.
  • Need frontier-band results on a budget: Codex + GPT-6.1 Sol (max) — 58.2% at ~$3.30 per derived solved task.
  • Claude-flavored workflows at value: Claude Code + Sonnet 5.5 — first on the AA mirror (63.6%), second officially (61.8%), but watch the board-leading token burn (19.4B).
  • Considering MiniMax M2.5 on benchmark claims alone: before trusting the README chart, require independently verifiable answers — an official TB 4.0 entry, a public harness config, numeric scores with CIs, unmodified environments, per-run (not 4-run-averaged) protocol.

Two different angles on the same ecosystem: Codex as a Python tooling play and AI-native development environments compared.

The reusable checklist for any vendor benchmark claim: (1) version, and is it comparable to the board you use? (2) harness — whose scaffold, which release? (3) environment — any Dockerfile, sandbox, or timeout deviations? (4) run count — averaging and retry policies? (5) evidence format — numeric rows with CIs, or chart images? Fail any of the five — as the 51.7% circulating for MiniMax M2.5 fails several at once — and the number isn’t evidence. It’s marketing wearing a leaderboard’s clothes.


Sources (all fetched 2026-10-11)

  • Official Terminal-Bench 4.0 leaderboard (via Harbor Hub and tbench.ai render, fetched 2026-10-11): all Table A figures (render-verified); n_trials = 330 and full-precision 58.18% per board transcriptions (Synergised; Capital and Compute), not visible in the HTML render.
  • tbench.ai, “Terminal-Bench 4.0” announcement: versioning semantics, 8-task removals, 19 fixes, flat 8-hour timeout, resource calibration, Sonnet 5 token figures.
  • BenchLeader, “Terminal-Bench 4.0 (AA) leaderboard” (data as of 11 Oct 2026; republishing Artificial Analysis’s run): all Table B figures.
  • MiniMax-AI/MiniMax-M2.5 GitHub README: Terminal Bench 2 methodology (Claude Code 2.0.64, modified Dockerfiles, 8-core/16 GB, 7,200 s timeout, 4-run average), SWE-Bench/BrowseComp claims, $1/hour cost claim, no Terminal-Bench numeric row in the text appendix.
  • modelscope/evalscope issue #1387: independent-reproduction thread; maintainer statements on the 51.7% figure, missing leaderboard entry, undisclosed harness, and 10+ point harness sensitivity (Chinese original, translation ours).
  • Synergised, “Why Terminal-Bench 4.0 Reset Every AI Agent Score”: TB version history (2.1: 89 tasks / ~83%; 3.0: 74 tasks / ~34%; 4.0: 66 tasks / ~52%+), crediting Capital and Compute for the version table and the $6.08–$234.24 per-solved-task derivation.
  • anthropics/claude-code docs and the openai/codex GitHub README (current Codex CLI tagline: “a coding agent from OpenAI that runs locally on your computer”); MiniMax-AI/minimax-code GitHub repo (README, LICENSE, open-source-status.md: MIT license, v0.4.12, Sep 18, 2026 open-source release).
  • Anthropic pricing docs and Haiku 5.5 announcement: Haiku 5.5 two-tier pricing, cache rates, model lineup pricing, and the announcement’s vendor-run TB 4.0 chart figures (Haiku 5.5 39.2%; Sonnet 5.5 reference 70.6%), cited as vendor-run numbers with the caveats above.
  • OpenAI API pricing page and “Introducing GPT-6.1 Sol”: Astra $10/$50, Sol $2/$10, cached-input claims.

References and further reading


Please let us know if you enjoyed this blog post. Share it with others to spread the knowledge! If you believe any images in this post infringe your copyright, please contact us promptly so we can remove them.



// adding consent banner