Kimi now has a new flagship route, a cheaper coding route, and several older access paths that should not be mixed together.

1
2
3
4
5
Need the newest Kimi flagship? → Kimi K3
Need cheaper routine coding? → Kimi K2.7 Code
Need lower K2.7 latency? → Kimi K2.7 Code HighSpeed
Need Kimi's membership-backed CLI? → Kimi Code with k3 / kimi-for-coding
Need an NVIDIA-hosted experiment? → Verify its current K2.6 trial and account terms

Current Kimi Routes

RouteModel identityBillingBest for
Kimi K3 APIkimi-k3Pay as you goLatest Kimi flagship, 1M context, multimodal/agentic evals
Kimi K2.7 Code APIkimi-k2.7-codePay as you goReproducible coding API evaluation
Kimi K2.7 Code HighSpeedkimi-k2.7-code-highspeedPay as you go at 2x base token pricesLatency-sensitive coding
Kimi Code membershipk3, kimi-for-coding, provider-managed backend IDsSubscription credits and quotasOfficial CLI and supported coding agents
NVIDIA NIM trialmoonshotai/kimi-k2.6Trial terms; production requires a separate subscriptionBounded internal compatibility evaluation

K3 is the newest Kimi flagship and the right route for latest-Kimi intent. K2.7 Code remains the cheaper routine coding model. K2.6 remains a compatibility, NVIDIA NIM, and prior-comparison lane; K2.5 remains historical context.

API Pricing

Prices are per 1 million tokens.

ModelCache-hit inputCache-miss inputOutputContext
K3$0.30$3.00$15.001M
K2.7 Code$0.19$0.95$4.00256K
K2.7 Code HighSpeed$0.38$1.90$8.00256K
K2.6$0.16$0.95$4.00256K

K3 is priced as a flagship escalation lane, not a budget lane. Moonshot has released the model weights and serving materials under the Kimi K3 License. Treat them as open weights, not as a generic open-source grant; review that license and your serving requirements before self-hosting.

K2.7 Code requires thinking mode in the documented API path. HighSpeed is the same model at higher token prices, with Kimi advertising approximately 180 tokens/s and up to 260 tokens/s for short contexts while warning that constrained resources can make performance fluctuate.

Kimi Code Membership

The current Kimi Code model configuration shows K3 and K2.7 Code across plan-dependent model IDs:

1
2
curl -fsSL https://code.kimi.com/kimi-code/install.sh | bash
kimi

Membership-backed third-party tools can use k3 or kimi-for-coding, depending on plan and context entitlement. Membership k3 supports low/high/max reasoning; Moderato exposes up to 256K context, while Allegretto and higher tiers expose up to 1M. That is convenient but less reproducible than pinning kimi-k3 or kimi-k2.7-code in the API, where K3 is currently always-on and max-only.

Kimi documents weekly quota refresh, a rolling five-hour frequency window, shared device/API-key quota, and a unified monthly credit pool across membership features. Check the signed-in console for the current balance and exact plan terms.

The VS Code path is transitional: current docs say new extension installs are limited to legacy Python CLI users while the TypeScript CLI integration is adapted. Other editors can connect through ACP.

Free and Promotional Routes

The safe wording is narrow:

  • NVIDIA’s live catalog currently exposes Kimi K2.6 as a trial service, not K3.
  • The public docs do not establish a universal payment requirement, quota, or expiry; verify the signed-in account terms.
  • NVIDIA’s trial terms prohibit production use without a separate NVIDIA or service-provider subscription.
  • Do not configure the trial as an autonomous fallback.
  • That does not make K2.7 Code free.
  • A no-payment-required trial can be labeled Free only while the live signup requires no payment.
  • A discounted first month is Low-cost, not free.
  • Checkout screenshots and old campaign rules are not evergreen entitlements.
  • Kimi’s current Moonshot Together rules describe a referral draw with published odds. The 30- and 365-day credit tiers require a paid-subscription task chance, awarded credits expire December 31, and AIHackers has not run redemption or published a referral CTA. Use the neutral draw guide.

NVIDIA NIM K2.6 setup covers the current official NVIDIA model ID and non-production restriction. NVIDIA did not list K3 when checked August 1, so use Moonshot’s direct API or released weights for latest-Kimi intent. Recheck the provider catalog, signup requirements, quota, and model ID before recommending the trial.

For a free coding IDE rather than Kimi specifically, use the Free Stack guide and verify its current hosted model roster.

Evidence and Quality

benchmark artifact

Kimi K3, K2.7, and Current Coding Alternatives

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
Kimi K3Moonshot AIactive
Active Kimi API/product flagship; official weights and serving materials released under the Kimi K3 License.
1M$3.00 / 1M$15.00 / 1MMoonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified.Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching.
  • Artificial Analysis Intelligence Index v4.1: 57 (independent)
  • Artificial Analysis output speed: 62 tokens/s (independent)
  • Moonshot launch benchmark suite: vendor-reported max-reasoning table (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task.Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests.Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive]2026-08-01
Kimi K2.7 CodeMoonshot AIactive
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
GLM-5.2Z.AIactive
Current Z.AI flagship coding model and supported-tool value lane.
1M$1.40 / 1M$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass.Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Gemini 3 FlashGoogleactive
Current Gemini value lane where Gemini API or Vertex AI fits.
1.05M input$0.50 / 1M$3.00 / 1Mnot verifiedFunction calling and code execution supported.not verifiedPreview model positioned for lower latency; independent value not imported.High-context value lane when Gemini API or Vertex AI fits.Gemini API models, Gemini API pricing, Artificial Analysis: Gemini 3 Flash, LMArena leaderboard dataset2026-05-26
Claude Sonnet 5Anthropicactive
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5; escalate only when the premium pass changes the accepted result.Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-07-25
Claude Opus 4.8Anthropichistorical
Still available, but superseded by Opus 5 for current premium comparisons.
1M$5.00 / 1M$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25

K3 has early independent evidence but no AIHackers-owned repository result. K2.7 improvements remain vendor-reported.

Kimi K3 has an Artificial Analysis Index 57 row and Moonshot launch-suite evidence. Kimi’s earlier K2.7 release notes report improvements over K2.6 of 10.4% on Program-Bench, 11.4% on MCP Mark Verified, and 76.2% on SWE Marathon, plus 30% lower reasoning-token use. Keep the source labels attached.

Do not reuse K2.5’s 76.8% SWE-bench Verified result as a K2.7 score.

Cost per Successful Task

Run the same repository task and record:

  • exact model ID or backend alias;
  • input, output, cache hits, and cache writes;
  • retries and wall-clock time;
  • tests and accepted completion;
  • human review and repair.

The HighSpeed premium is justified only if lower latency changes the workflow. The membership is justified only if its credits and tool convenience beat measured API use for the same accepted tasks.

Which Route to Choose

RequirementRecommended route
Newest Kimi flagship and 1M contextK3 API or K3-enabled Kimi Code plan
Explicit current model and repeatable API testK2.7 Code API
Lowest K2.7 cache-hit costBase K2.7 Code
Lowest latencyHighSpeed after measuring the premium
Official Kimi CLI and third-party coding agentsKimi Code membership
One-off zero-dollar testA currently verified trial, with the actual model labeled
1M context at lower token priceCompare GLM-5.2, Gemini, MiniMax, or Xiaomi instead
Premium final reviewClaude Opus 4.8 or another approved premium lane

Sources


Last verified: August 20, 2026. API prices, HighSpeed capacity, membership aliases, referral odds, credits, checkout terms, open-weight status, and third-party trial routes can change independently.