Use this section to answer two separate questions: what is newest, and what can be used normally today. An announced model can be newer without being generally available.

Active Models

These are the current generally usable lanes in this comparison set.

Claude Sonnet 5

The first Claude cost/performance test before Opus API pricing: 1M context, 128K maximum output, $2/$10 introductory pricing through August 31, then $3/$15. Launch comparisons are vendor evidence; retain Opus for highest-accuracy arbitration.

Claude Opus 5

The new premium Claude baseline for architecture, consequential agentic coding, knowledge work, and final arbitration. It keeps $5/$25 pricing, adds five effort levels, and scored 61 at max effort on the Artificial Analysis Intelligence Index.

GLM-5.2

Z.AI’s current flagship coding model: 1M context, 128K maximum output, $1.40/$4.40 input/output pricing, and Artificial Analysis Index 51. Treat it as a value lane to test, not a proven Opus replacement.

Kimi K3

Moonshot’s newest Kimi flagship: 2.8T vendor-stated parameters, 1M context, native vision, $0.30/$3/$15 cache-hit/input/output pricing, and Artificial Analysis Index 57. Official weights and serving materials are released under the Kimi K3 License; self-hosting cost and AIHackers CAR remain unverified.

Kimi K2.7 Code

Moonshot’s cheaper routine Kimi coding API lane, with base and HighSpeed model IDs, 256K context, multimodal input, required thinking mode, and $0.95/$4.00 cache-miss input/output pricing.

GPT-5.6 Sol, Terra, and Luna

OpenAI’s generally available family across ChatGPT, Codex, and API: Sol is the flagship at $5/$30, Terra is the balanced tier at $2/$12, and Luna is the high-volume tier at $0.20/$1.20. These are standard short-context rates; requests above 272K input use higher long-context rates. See GPT-5.6 Luna Is OpenAI’s Value Default for subscription ranges, API economics, and the CAR decision.

DeepSeek V4 Flash 0731

DeepSeek’s direct-API and MIT-weight value candidate: 1M context, $0.14 cache-miss input, $0.0028 cache-hit input, and $0.28 output. Artificial Analysis scores 0731 at 50 and places it on its intelligence-versus-cost Pareto frontier; high output use, an 84% hallucination rate, and site-owned not-run status keep it in the evaluate-first lane.

Active supporting lanes

Guarded and restricted: Fable and Mythos

Claude Fable 5 and Mythos 5

Keep these models out of default routing for different reasons. Anthropic restored Fable globally on native Claude surfaces July 1, but included usage is plan-specific, cloud rollout is incomplete, and safeguards and retention still matter. Mythos remains approved-organization access.

Current Evidence Table

benchmark artifact

Current Model Status, Cost, and Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GPT-5.6 SolOpenAIactive
Generally available through ChatGPT paid plans, Codex paid plans, and the OpenAI API; plan and effort options vary.
1.05M$5.00 / 1M$30.00 / 1MArtificial Analysis reports 80 on its Coding Agent Index at max effort; OpenAI reports 64.6% on SWE-bench Pro.Generally available in API and paid Codex plans; max and ultra modes are vendor-documented.
  • Artificial Analysis Intelligence Index: 59 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 80 at max effort (independent)
  • SWE-bench Pro: 64.6% (vendor)
  • AIHackers repo eval: not-run (site-owned)
OpenAI announced a selected-customer Cerebras preview for July; production latency is not verified.Generally available flagship; test on real tasks and apply stronger controls for agentic or cyber work.OpenAI GPT-5.6 general availability [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
GPT-5.6 TerraOpenAIactive
Generally available through ChatGPT Work, Codex including Free and Go, and the OpenAI API.
1.05M$2.00 / 1M$12.00 / 1MArtificial Analysis reports 77 on its Coding Agent Index at max effort; OpenAI reports 63.4% on SWE-bench Pro.Generally available in API and Codex, including the Free and Go Codex lane.
  • Artificial Analysis Intelligence Index: 55 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 77 at max effort (independent)
  • SWE-bench Pro: 63.4% (vendor)
  • AIHackers repo eval: not-run (site-owned)
not verifiedGenerally available balanced lane; compare accepted-task cost with Sol, Luna, and other providers.OpenAI GPT-5.6 general availability [archive], OpenAI GPT-5.6 price update [archive], OpenAI API pricing [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
GPT-5.6 LunaOpenAIactive
Generally available through paid ChatGPT Work/Codex plans and the OpenAI API.
1.05M$0.20 / 1M$1.20 / 1MArtificial Analysis reports 75 on its Coding Agent Index at max effort; BenchLM ranks Luna #6/129 in coding with an Estimated overall position.Generally available in API and paid Codex plans.
  • Artificial Analysis Intelligence Index: 51 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 75 at max effort (independent)
  • BenchLM coding rank: #6/129; overall Estimated (independent aggregator)
  • SWE-bench Pro: 62.7% (vendor)
  • AIHackers repo eval: not-run (site-owned)
Vendor-positioned as fastest; measured production latency is not verified.Generally available lowest-cost GPT-5.6 tier; verify quality and cost per accepted task.OpenAI GPT-5.6 general availability [archive], OpenAI GPT-5.6 price update [archive], OpenAI API pricing [archive], OpenAI GPT-5.6 Luna model page [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, BenchLM GPT-5.6 Luna profile [archive], OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
DeepSeek V4 Flash 0731DeepSeekactive
Official 0731 release available through DeepSeek's direct API; MIT-licensed weights are published separately. OpenCode's promotional free route has unverified backend revision identity.
1M$0.14 / 1M$0.28 / 1MArtificial Analysis reports Intelligence Index 50, one point behind GPT-5.6 Luna at max effort; this is shortlist evidence, not a universal coding-quality claim.DeepSeek documents tool calls, Responses API, and Anthropic-compatible access; real harness behavior remains workload-specific.
  • Artificial Analysis Intelligence Index: 50 (independent)
  • AA-Omniscience hallucination rate: 84% (independent)
  • Historical security field test: 0/10; pre-0731 preview only (site-owned historical)
  • AIHackers 0731 repo eval: not-run (site-owned)
Independent write-up reports about 206M total output tokens for the Intelligence Index; production speed and accepted-task efficiency are not verified.Direct-API and open-weight value candidate; test accepted-task cost, token use, hallucinations, and tool behavior before production routing.DeepSeek V4 Flash 0731 update [archive], DeepSeek V4 Flash 0731 model card [archive], DeepSeek API pricing [archive], Artificial Analysis DeepSeek V4 Flash 0731, OpenCode Zen pricing [archive], AIHackers app-hacking field test2026-08-01
GPT-5.5OpenAIactive
Generally available prior-generation OpenAI model retained for existing integrations and comparisons.
1.05M API; 400K Codex$5.00 / 1M$30.00 / 1Mnot verifiednot verifiednot verifiednot verifiedPrimary coding seat while ChatGPT/Codex limits fit the workload.OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset2026-06-28
Claude Sonnet 5Anthropicactive
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5; escalate only when the premium pass changes the accepted result.Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-07-25
Claude Opus 5Anthropicactive
Current premium Claude baseline; default on Max, strongest Pro model, and available through the Claude API and supported cloud platforms.
1M$5.00 / 1M$25.00 / 1MAnthropic reports major agentic-coding gains; Artificial Analysis reports joint first on its Coding Agent Index at xhigh.Thinking is on by default; five effort settings materially change cost, latency, and task performance.
  • Artificial Analysis Intelligence Index: 61 at max effort; $2.03 per task (independent)
  • AA-Briefcase: 1720 Elo / $17.79 max; 1606 Elo / $10.41 high (independent)
  • Anthropic launch evaluations: vendor-reported; configuration varies by evaluation (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports high/xhigh/max AA-Briefcase runtimes of 25.7/34.3/36.2 minutes per task; Fast mode is a separate API research preview.Premium escalation for consequential agentic and knowledge work; compare complete-task cost against Sonnet 5 before default routing.Anthropic Claude Opus 5 launch [archive], What's new in Claude Opus 5 [archive], Claude Opus 5 system card [archive], Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Claude Opus 4.8Anthropichistorical
Still available, but superseded by Opus 5 for current premium comparisons.
1M$5.00 / 1M$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25
Claude Fable 5Anthropicactive
Generally available; temporary subscription allowances ended July 7 and current subscription use is through usage credits.
1M$10.00 / 1M$50.00 / 1MAnthropic reports frontier launch results; independent reproducible ranking is pending.Guarded-domain requests can refuse or fall back; verify account behavior before routing.
  • Artificial Analysis Intelligence Index: 60 at max (independent)
  • AA-Briefcase: 1574 Elo / $22.30 per task (independent)
  • AIHackers repo eval: not verified (site-owned)
Task latency varies; compare complete-task runtime before escalation.High-cost guarded escalation only; use Opus 5 as the practical Claude premium baseline.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive], Artificial Analysis: Claude Opus 5 [archive], Artificial Analysis: Claude Opus 5 on AA-Briefcase [archive]2026-07-25
Claude Mythos 5Anthropicrestricted
Restored to a set of approved US organizations; broader Glasswing access remains restricted.
1M$10.00 / 1M$50.00 / 1MGeneral coding quality is not independently verified for an accessible production route.Invitation-only research access; account and compliance approval required.
  • Independent cross-model evaluation: not verified (independent)
  • AIHackers repo eval: not verified (site-owned)
not verifiedRestricted research context, not a normal production or buying recommendation.Claude models overview [archive], Claude API pricing [archive], Anthropic Fable 5 and Mythos 5 [archive], Anthropic Fable/Mythos access statement [archive], Anthropic Fable 5 redeployment [archive]2026-07-01
Kimi K3Moonshot AIactive
Active Kimi API/product flagship; official weights and serving materials released under the Kimi K3 License.
1M$3.00 / 1M$15.00 / 1MMoonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified.Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching.
  • Artificial Analysis Intelligence Index v4.1: 57 (independent)
  • Artificial Analysis output speed: 62 tokens/s (independent)
  • Moonshot launch benchmark suite: vendor-reported max-reasoning table (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task.Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests.Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive]2026-08-01
GLM-5.2Z.AIactive
Current Z.AI flagship coding model and supported-tool value lane.
1M$1.40 / 1M$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass.Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Kimi K2.7 CodeMoonshot AIactive
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

Independent, vendor, historical-preview, and site-owned evidence are labeled separately. Model availability, a weight release, or a vendor benchmark does not establish cost per successful task.

Do not compare scores without matching the benchmark variant and harness. SWE-bench Verified, SWE-Bench Pro, Terminal-Bench, Artificial Analysis, vendor preference tests, and site-owned repository tests answer different questions. Use How to Read AI Benchmarks to turn those results into a shortlist before running your own tasks.

Historical Guides

These pages remain indexable for older integrations, pricing history, and search intent. They are not current recommendations.

Older model names can still be correct when a tool genuinely exposes that version or a dated article describes the period. They should not be relabeled as current leaders.

Selection Shortcuts

NeedStart withEscalate or compare with
Daily Claude productionSonnet 5Opus 5
Premium Claude reviewOpus 5Fable only after access, cost, and compliance checks
Low-cost long-context codingGLM-5.2Opus 5 or GPT-5.6 for arbitration
Latest Kimi frontier testKimi K3Opus 5 or GPT-5.6 for arbitration
Cheaper Kimi coding APIKimi K2.7 CodeGLM-5.2 when 1M context matters
OpenAI productionGPT-5.6 Terra or LunaGPT-5.6 Sol for harder work
Lowest-cost direct API or MIT weightsDeepSeek V4 Flash 0731Luna, GLM-5.2, or a stronger reviewed lane when acceptance fails
Local or private inferenceGemma 4Compare hardware and quantization fit

See Budget Tier, Mid-Range, Premium, and Smart Spend for workload-specific decisions.


Last verified: August 1, 2026. Model access, prices, limits, safeguards, open-weight status, and benchmark positions can change independently.

Models

Claude Sonnet 5 Guide

Claude Sonnet 5 pricing, 1M context, API migration changes, tokenizer cost implications, availability, and source-labeled launch evidence.

Models

GPT-5.6 Sol, Terra, and Luna Guide

Current GPT-5.6 Sol, Terra, and Luna access, short- and long-context pricing, independent benchmark signals, and deployment controls.

Models

Claude Opus 4.8 Historical Guide

Claude Opus 4.8 historical pricing and benchmark context, retained for older integrations after Opus 5 became the premium Claude baseline.

Models

GLM-5.2: July Value Coding Pick

GLM-5.2 is the July best value coding model to test: 1M context, AA Index 51, $1.40/$4.40 API pricing, and Opus 4.8 comparison math.

Models

Kimi K2.7 Code: Coding Model Guide

Kimi K2.7 Code is Moonshot's cheaper routine coding lane after K3: 256K context, base/HighSpeed pricing, source-labeled benchmark deltas, and eval caveats.

Models

MiniMax M3: Value Coding Model Guide

MiniMax M3 is a value coding model candidate with 1M context, multimodal input, Token Plan economics, and vendor-reported benchmark strength. Test it before replacing premium lanes.

Models

Kimi K2.5 Historical Guide

Historical Kimi K2.5 capabilities, benchmarks, and older free-access routes. Use Kimi K3 for newest-Kimi decisions and K2.7 Code for cheaper coding decisions.

Models

GLM 4.7: Legacy Z.AI Model Guide

Legacy guide for Z.AI GLM 4.7. GLM-5.2 is now the current Z.AI coding-plan model; use this page for historical context and GLM-4.7 fallback routing.

Models

Claude Opus 4.5 Historical Guide

Historical Claude Opus 4.5 guide with current routing notes. Compare Sonnet 5, Opus 5, restored Fable 5, and GPT-5.6 before buying premium capacity.