Choose the coding tool for its workflow, controls, and billing model. Choose the underlying model separately. Tool subscriptions, API token prices, and public benchmark scores are not interchangeable.

Quick Decision

PriorityStart withWhyVerify before committing
OpenAI-native cloud agents and parallel tasksCodexOpenAI account integration and isolated task workflowsPlan limits, current model picker, workspace controls, and data terms
Claude-native terminal work and premium reviewClaude CodeSonnet 5 daily lane and Opus 5 premium escalationSubscription/API boundary, model availability, and retention requirements
Kimi-native codingKimi Code / K3 / K2.7 CodeK3 for newest flagship tests; K2.7 Code for cheaper routine Kimi codingMembership routing, quota, exact model ID, K3 entitlement, and checkout offer

Current recommendation: use the tool that fits your repository controls, then run the same real task through the model lanes you can actually access. Do not select a tool from an old SWE-bench row alone.

Current Model Map

Tool or providerActive model contextGuarded or pending contextStatus rule
OpenAI CodexGPT-5.6 Sol, Terra, and LunaTrusted Access for less-restricted cyber workGPT-5.6 is generally available; product access and effort choices vary by plan
Claude CodeSonnet 5 for daily work; Opus 5 for premium reviewFable 5 and Mythos 5Fable is restored but guarded and high-cost; Mythos remains trusted-access only
Kimi Code / APIKimi K3, Kimi K2.7 Code, and K2.7 Code HighSpeedK3 weights promised by July 27K3 is the newest Kimi flagship; K2.7 is the cheaper routine coding lane; K2.5/K2.6 are historical or compatibility context

The model exposed by a subscription or tool-facing alias can differ from the public API model discussed in a benchmark. Confirm the exact model ID or account UI instead of inferring it from the product name.

Model Evidence

benchmark artifact

Coding Tool Model Lanes

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-5.5OpenAIactive
Generally available prior-generation OpenAI model retained for existing integrations and comparisons.
1.05M API; 400K Codex$5.00 / 1M$30.00 / 1Mnot verifiednot verifiednot verifiednot verifiedPrimary coding seat while ChatGPT/Codex limits fit the workload.OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset2026-06-28
GPT-5.6 SolOpenAIactive
Generally available through ChatGPT paid plans, Codex paid plans, and the OpenAI API; plan and effort options vary.
1.05M$5.00 / 1M$30.00 / 1MArtificial Analysis reports 80 on its Coding Agent Index at max effort; OpenAI reports 64.6% on SWE-bench Pro.Generally available in API and paid Codex plans; max and ultra modes are vendor-documented.
  • Artificial Analysis Intelligence Index: 59 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 80 at max effort (independent)
  • SWE-bench Pro: 64.6% (vendor)
  • AIHackers repo eval: not-run (site-owned)
OpenAI announced a selected-customer Cerebras preview for July; production latency is not verified.Generally available flagship; test on real tasks and apply stronger controls for agentic or cyber work.OpenAI GPT-5.6 general availability [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
Claude Sonnet 5Anthropicactive
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5; escalate only when the premium pass changes the accepted result.Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-07-25
Claude Opus 5Anthropicactive
Current premium Claude baseline; default on Max, strongest Pro model, and available through the Claude API and supported cloud platforms.
1M$5.00 / 1M$25.00 / 1MAnthropic reports major agentic-coding gains; Artificial Analysis reports joint first on its Coding Agent Index at xhigh.Thinking is on by default; five effort settings materially change cost, latency, and task performance.
  • Artificial Analysis Intelligence Index: 61 at max effort; $2.03 per task (independent)
  • AA-Briefcase: 1720 Elo / $17.79 max; 1606 Elo / $10.41 high (independent)
  • Anthropic launch evaluations: vendor-reported; configuration varies by evaluation (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports high/xhigh/max AA-Briefcase runtimes of 25.7/34.3/36.2 minutes per task; Fast mode is a separate API research preview.Premium escalation for consequential agentic and knowledge work; compare complete-task cost against Sonnet 5 before default routing.Anthropic Claude Opus 5 launch [archive], What's new in Claude Opus 5 [archive], Claude Opus 5 system card [archive], Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Claude Opus 4.8Anthropichistorical
Still available, but superseded by Opus 5 for current premium comparisons.
1M$5.00 / 1M$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25
Claude Fable 5Anthropicactive
Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits.
1M$10.00 / 1M$50.00 / 1MAnthropic reports frontier launch results; independent reproducible ranking is pending.Guarded-domain requests can refuse or fall back; verify account behavior before routing.
  • Artificial Analysis Intelligence Index: 60 at max (independent)
  • AA-Briefcase: 1574 Elo / $22.30 per task (independent)
  • AIHackers repo eval: not verified (site-owned)
Task latency varies; compare complete-task runtime before escalation.High-cost guarded escalation only; use Opus 5 as the practical Claude premium baseline.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Kimi K3Moonshot AIactive
Active Kimi API/product flagship; official weights and serving materials released under the Kimi K3 License.
1M$3.00 / 1M$15.00 / 1MMoonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified.Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching.
  • Artificial Analysis Intelligence Index v4.1: 57 (independent)
  • Artificial Analysis output speed: 62 tokens/s (independent)
  • Moonshot launch benchmark suite: vendor-reported max-reasoning table (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task.Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests.Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive]2026-08-01
Kimi K2.7 CodeMoonshot AIactive
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
GLM-5.2Z.AIactive
Current Z.AI flagship coding model and supported-tool value lane.
1M$1.40 / 1M$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass.Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

This table separates model status and evidence provenance. It does not rank the surrounding coding tools or claim that one benchmark predicts repository productivity.

Useful evidence has three levels:

  1. Independent: a named third party publishes a methodology and model-specific result.
  2. Vendor: the provider publishes a result or relative improvement; useful for deciding what to test, not for adopting the claim as an AIHackers ranking.
  3. Site-owned: the same repository task, harness, completion rules, and artifacts are available for review.

AIHackers does not yet have a controlled cross-tool repository evaluation for current Codex, Claude Code, and Kimi Code. Claims such as “95% of another model’s capability,” “best quality,” or “8x better value” are therefore not supported here.

Pricing: Keep Three Lanes Separate

Tool subscriptions

Codex, Claude Code, and Kimi Code expose plan limits, quotas, credits, or checkout-controlled offers. Those entitlements can change independently of API list prices. Verify the live plan page and account before purchase.

API token prices

ModelInputCache readOutputStatus
Kimi K3$3.00$0.30$15.00Active API; weights released under the Kimi K3 License
Kimi K2.7 Code$0.95$0.19$4.00Active
GLM-5.2$1.40$0.26$4.40Active alternative
Claude Sonnet 5$2 intro; $3 standardverify current docs$10 intro; $15 standardActive; intro ends Aug 31
Claude Opus 5$5.00$0.50$25.00Active premium
GPT-5.5$5.00check current OpenAI docs$30.00Existing integrations
GPT-5.6 Sol$5.00$0.50$30.00Generally available

Prices are per 1 million tokens. They do not include tool subscription fees, retries, cache writes, failed patches, or human review.

Cost per successful task

This is the metric that matters. Record:

  • input, cache writes/reads, and output tokens;
  • wall-clock time and retries;
  • whether tests pass;
  • whether the patch is accepted without repair;
  • human review and cleanup time.

A cheaper token price can lose if a model produces longer outputs, retries more often, or requires expensive review. A premium model can be rational when it prevents a costly mistake.

Workflow Comparison

Workflow questionCodexClaude CodeKimi Code
Parallel cloud task workflowStrong product fitVerify current Claude workflowVerify current Kimi workflow
Terminal-first local orchestrationSupported product paths varyCore Claude Code workflowKimi CLI and compatible tools
Premium second-pass reviewUse active OpenAI modelOpus 5 is the current Claude baselineRoute to another provider if needed
Low API token priceCompare current OpenAI APISonnet/Opus cost moreK2.7 Code is the low-cost Kimi lane
Guarded frontier workGPT-5.6 Trusted Access for eligible cyber defendersFable/Mythos only after access/compliance checksNot applicable in this comparison

No tool makes an autonomous agent safe by default. Apply least-privilege credentials, scoped worktrees, confirmation gates for destructive actions, and independent test/diff review.

Evaluation Protocol

Run the same four tasks:

TaskPass condition
Repository mapCorrect modules, contracts, and commands; no invented files
Real bug fixMinimal correct patch with relevant tests
Cross-file refactorPreserves behavior and repository conventions
Patch reviewFile-grounded findings with no generic filler

Pin the tool version, model ID, reasoning setting, repository commit, prompt, and acceptance rules. If a tool hides the exact model, record that as an evaluation limitation.

Current Verdict

  • Choose Codex for OpenAI-native cloud-agent workflows after confirming the current model and plan limits.
  • Choose Claude Code when Claude-native terminal work and an Opus 5 premium review lane matter.
  • Choose Kimi Code/K3 when you want the newest Kimi flagship and can measure CAR; choose K2.7 Code when low Kimi API pricing fits.
  • Route GPT-5.6 by workload and plan; use Trusted Access only for eligible cyber work. Keep Mythos 5 behind its access checks and test restored Fable only as a measured escalation after Opus.
  • Keep Opus 4.5, Kimi K2.5/K2.6, and older GPT rows only as historical or product-specific context.

Sources


Last verified: July 26, 2026. Tool plans, model pickers, access, quotas, open-weight status, and API prices can change independently.