Choose the coding tool for its workflow, controls, and billing model. Choose the underlying model separately. Tool subscriptions, API token prices, and public benchmark scores are not interchangeable.
Quick Decision
| Priority | Start with | Why | Verify before committing |
|---|---|---|---|
| OpenAI-native cloud agents and parallel tasks | Codex | OpenAI account integration and isolated task workflows | Plan limits, current model picker, workspace controls, and data terms |
| Claude-native terminal work and premium review | Claude Code | Sonnet 5 daily lane and Opus 5 premium escalation | Subscription/API boundary, model availability, and retention requirements |
| Kimi-native coding | Kimi Code / K3 / K2.7 Code | K3 for newest flagship tests; K2.7 Code for cheaper routine Kimi coding | Membership routing, quota, exact model ID, K3 entitlement, and checkout offer |
Current recommendation: use the tool that fits your repository controls, then run the same real task through the model lanes you can actually access. Do not select a tool from an old SWE-bench row alone.
Current Model Map
| Tool or provider | Active model context | Guarded or pending context | Status rule |
|---|---|---|---|
| OpenAI Codex | GPT-5.6 Sol, Terra, and Luna | Trusted Access for less-restricted cyber work | GPT-5.6 is generally available; product access and effort choices vary by plan |
| Claude Code | Sonnet 5 for daily work; Opus 5 for premium review | Fable 5 and Mythos 5 | Fable is restored but guarded and high-cost; Mythos remains trusted-access only |
| Kimi Code / API | Kimi K3, Kimi K2.7 Code, and K2.7 Code HighSpeed | K3 weights promised by July 27 | K3 is the newest Kimi flagship; K2.7 is the cheaper routine coding lane; K2.5/K2.6 are historical or compatibility context |
The model exposed by a subscription or tool-facing alias can differ from the public API model discussed in a benchmark. Confirm the exact model ID or account UI instead of inferring it from the product name.
Model Evidence
benchmark artifact
Coding Tool Model Lanes
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.5 | OpenAI | active Generally available prior-generation OpenAI model retained for existing integrations and comparisons. | 1.05M API; 400K Codex | $5.00 / 1M | $30.00 / 1M | not verified | not verified | not verified | not verified | Primary coding seat while ChatGPT/Codex limits fit the workload. | OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset | 2026-06-28 |
| GPT-5.6 Sol | OpenAI | active Generally available through ChatGPT paid plans, Codex paid plans, and the OpenAI API; plan and effort options vary. | 1.05M | $5.00 / 1M | $30.00 / 1M | Artificial Analysis reports 80 on its Coding Agent Index at max effort; OpenAI reports 64.6% on SWE-bench Pro. | Generally available in API and paid Codex plans; max and ultra modes are vendor-documented. |
| OpenAI announced a selected-customer Cerebras preview for July; production latency is not verified. | Generally available flagship; test on real tasks and apply stronger controls for agentic or cyber work. | OpenAI GPT-5.6 general availability [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card | 2026-08-01 |
| Claude Sonnet 5 | Anthropic | active Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths. | 1M | $2.00 / 1M | $10.00 / 1M | Anthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending. | Available in Claude Code and the Claude API; adaptive thinking is on by default. |
| No site-owned normalized latency result is verified. | First Claude cost/performance test before Opus 5; escalate only when the premium pass changes the accepted result. | Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive] | 2026-07-25 |
| Claude Opus 5 | Anthropic | active Current premium Claude baseline; default on Max, strongest Pro model, and available through the Claude API and supported cloud platforms. | 1M | $5.00 / 1M | $25.00 / 1M | Anthropic reports major agentic-coding gains; Artificial Analysis reports joint first on its Coding Agent Index at xhigh. | Thinking is on by default; five effort settings materially change cost, latency, and task performance. |
| Artificial Analysis reports high/xhigh/max AA-Briefcase runtimes of 25.7/34.3/36.2 minutes per task; Fast mode is a separate API research preview. | Premium escalation for consequential agentic and knowledge work; compare complete-task cost against Sonnet 5 before default routing. | Anthropic Claude Opus 5 launch [archive], What's new in Claude Opus 5 [archive], Claude Opus 5 system card [archive], Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive] | 2026-07-25 |
| Claude Opus 4.8 | Anthropic | historical Still available, but superseded by Opus 5 for current premium comparisons. | 1M | $5.00 / 1M | $25.00 / 1M | Historical premium Claude baseline; use Opus 5 for new task-level comparisons. | Still available for pinned integrations; new Claude premium routing should test Opus 5. |
| Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary. | Historical premium baseline. Use Claude Opus 5 for current Claude premium routing. | Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard | 2026-07-25 |
| Claude Fable 5 | Anthropic | active Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits. | 1M | $10.00 / 1M | $50.00 / 1M | Anthropic reports frontier launch results; independent reproducible ranking is pending. | Guarded-domain requests can refuse or fall back; verify account behavior before routing. |
| Task latency varies; compare complete-task runtime before escalation. | High-cost guarded escalation only; use Opus 5 as the practical Claude premium baseline. | Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive] | 2026-07-25 |
| Kimi K3 | Moonshot AI | active Active Kimi API/product flagship; official weights and serving materials released under the Kimi K3 License. | 1M | $3.00 / 1M | $15.00 / 1M | Moonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified. | Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching. |
| Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task. | Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests. | Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive] | 2026-08-01 |
| Kimi K2.7 Code | Moonshot AI | active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| GLM-5.2 | Z.AI | active Current Z.AI flagship coding model and supported-tool value lane. | 1M | $1.40 / 1M | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass. | Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
This table separates model status and evidence provenance. It does not rank the surrounding coding tools or claim that one benchmark predicts repository productivity.
Useful evidence has three levels:
- Independent: a named third party publishes a methodology and model-specific result.
- Vendor: the provider publishes a result or relative improvement; useful for deciding what to test, not for adopting the claim as an AIHackers ranking.
- Site-owned: the same repository task, harness, completion rules, and artifacts are available for review.
AIHackers does not yet have a controlled cross-tool repository evaluation for current Codex, Claude Code, and Kimi Code. Claims such as “95% of another model’s capability,” “best quality,” or “8x better value” are therefore not supported here.
Pricing: Keep Three Lanes Separate
Tool subscriptions
Codex, Claude Code, and Kimi Code expose plan limits, quotas, credits, or checkout-controlled offers. Those entitlements can change independently of API list prices. Verify the live plan page and account before purchase.
API token prices
| Model | Input | Cache read | Output | Status |
|---|---|---|---|---|
| Kimi K3 | $3.00 | $0.30 | $15.00 | Active API; weights released under the Kimi K3 License |
| Kimi K2.7 Code | $0.95 | $0.19 | $4.00 | Active |
| GLM-5.2 | $1.40 | $0.26 | $4.40 | Active alternative |
| Claude Sonnet 5 | $2 intro; $3 standard | verify current docs | $10 intro; $15 standard | Active; intro ends Aug 31 |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | Active premium |
| GPT-5.5 | $5.00 | check current OpenAI docs | $30.00 | Existing integrations |
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | Generally available |
Prices are per 1 million tokens. They do not include tool subscription fees, retries, cache writes, failed patches, or human review.
Cost per successful task
This is the metric that matters. Record:
- input, cache writes/reads, and output tokens;
- wall-clock time and retries;
- whether tests pass;
- whether the patch is accepted without repair;
- human review and cleanup time.
A cheaper token price can lose if a model produces longer outputs, retries more often, or requires expensive review. A premium model can be rational when it prevents a costly mistake.
Workflow Comparison
| Workflow question | Codex | Claude Code | Kimi Code |
|---|---|---|---|
| Parallel cloud task workflow | Strong product fit | Verify current Claude workflow | Verify current Kimi workflow |
| Terminal-first local orchestration | Supported product paths vary | Core Claude Code workflow | Kimi CLI and compatible tools |
| Premium second-pass review | Use active OpenAI model | Opus 5 is the current Claude baseline | Route to another provider if needed |
| Low API token price | Compare current OpenAI API | Sonnet/Opus cost more | K2.7 Code is the low-cost Kimi lane |
| Guarded frontier work | GPT-5.6 Trusted Access for eligible cyber defenders | Fable/Mythos only after access/compliance checks | Not applicable in this comparison |
No tool makes an autonomous agent safe by default. Apply least-privilege credentials, scoped worktrees, confirmation gates for destructive actions, and independent test/diff review.
Evaluation Protocol
Run the same four tasks:
| Task | Pass condition |
|---|---|
| Repository map | Correct modules, contracts, and commands; no invented files |
| Real bug fix | Minimal correct patch with relevant tests |
| Cross-file refactor | Preserves behavior and repository conventions |
| Patch review | File-grounded findings with no generic filler |
Pin the tool version, model ID, reasoning setting, repository commit, prompt, and acceptance rules. If a tool hides the exact model, record that as an evaluation limitation.
Current Verdict
- Choose Codex for OpenAI-native cloud-agent workflows after confirming the current model and plan limits.
- Choose Claude Code when Claude-native terminal work and an Opus 5 premium review lane matter.
- Choose Kimi Code/K3 when you want the newest Kimi flagship and can measure CAR; choose K2.7 Code when low Kimi API pricing fits.
- Route GPT-5.6 by workload and plan; use Trusted Access only for eligible cyber work. Keep Mythos 5 behind its access checks and test restored Fable only as a measured escalation after Opus.
- Keep Opus 4.5, Kimi K2.5/K2.6, and older GPT rows only as historical or product-specific context.
Sources
- OpenAI: GPT-5.6 general availability and system card
- OpenAI: Codex
- Anthropic: Claude model overview and pricing
- Kimi: K3 launch blog, K3 quickstart, K2.7 Code quickstart, and K2.7 pricing
- Artificial Analysis: Claude Opus 5, Claude Opus 4.8, and GLM-5.2
Related links
- /models/gpt-5-6/
- /models/claude-opus-5/
- /models/claude-opus-4-8/
- /models/kimi-k3/
- /models/kimi-k2.7-code/
- /tools/kimi-code/
- /value/smart-spend/
- /risks/codex/cloud-dependency-risks/
Last verified: July 26, 2026. Tool plans, model pickers, access, quotas, open-weight status, and API prices can change independently.