Kimi K2.7 Code is now the cheaper routine Kimi coding lane, not the newest Kimi model. For latest-Kimi, 1M-context, and K3 launch intent, use Kimi K3. Use this page when the search is specifically about K2.7 Code pricing, HighSpeed, 256K coding workflows, or a lower-cost Kimi API route.

Moonshot’s Kimi API docs still list kimi-k2.7-code and kimi-k2.7-code-highspeed IDs, 256K-class context, multimodal input, thinking required, automatic context caching, ToolCalls, JSON Mode, and Partial Mode.

Use K2.7 Code when you want cheaper Kimi coding economics. Use K3 when 1M context, native flagship behavior, or the newest Kimi benchmark evidence is the reason you are testing Kimi at all. Keep K2.6 in the comparison set because it is still a relevant multimodal/API lane and exact-match search term.

Quick Facts

SpecKimi K2.7 Code
ProviderMoonshot AI / Kimi
Model IDskimi-k2.7-code, kimi-k2.7-code-highspeed
Context window262,144 tokens, commonly described as 256K
ThinkingRequired for K2.7 Code; disabling thinking throws an error
Base API pricing$0.19 cache-hit input / $0.95 cache-miss input / $4.00 output per 1M tokens
HighSpeed pricing$0.38 cache-hit input / $1.90 cache-miss input / $8.00 output per 1M tokens
Useful forCoding agents, multimodal coding tasks, Kimi-native API workflows
CaveatK3 is the newest Kimi flagship; vendor benchmark claims and speed claims need workload-level checks

K2.7 Code vs K2.6

QuestionK2.7 CodeK2.6
Current roleCheaper routine Kimi coding lanePrior Kimi release and still relevant comparison lane
Model IDskimi-k2.7-code, kimi-k2.7-code-highspeedkimi-k2.6
Context256K-class256K-class
Thinking modeRequiredCan be enabled or disabled
Cache-hit input$0.19 / 1M$0.16 / 1M
Cache-miss input$0.95 / 1M$0.95 / 1M
Output$4.00 / 1M$4.00 / 1M
Fast variantHighSpeed at $0.38 / $1.90 / $8.00No separate HighSpeed price page in the current docs

The short version: choose K2.7 Code for lower-cost Kimi coding tests and K2.6 when you need non-thinking mode, have existing integrations, or are comparing exact K2.6 pricing/search intent.

Benchmark Evidence

Kimi’s June release notes report these changes relative to K2.6:

Benchmark or measureVendor-reported delta
Program-Bench+10.4%
MCP Mark Verified+11.4%
SWE Marathon+76.2%
Reasoning-token use30% lower

These are relative vendor-reported deltas, not absolute scores. They cannot be compared directly with SWE-bench Verified, SWE-Bench Pro, Terminal-Bench, or Artificial Analysis. AIHackers has not verified a normalized independent K2.7 score or run a controlled repository evaluation.

benchmark artifact

Kimi K2.7 and Current Alternatives

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
Kimi K2.7 CodeMoonshot AIactive
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
GLM-5.2Z.AIactive
Current Z.AI flagship coding model and supported-tool value lane.
1M$1.40 / 1M$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass.Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Claude Sonnet 5Anthropicactive
Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths.
1M$2.00 / 1M$10.00 / 1MAnthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending.Available in Claude Code and the Claude API; adaptive thinking is on by default.
  • Cross-model benchmark evidence: vendor-reported; updated chart and system card preferred (vendor)
  • Artificial Analysis task cost: $1.53 per Intelligence Index task at max (independent)
  • AIHackers repo eval: not verified (site-owned)
No site-owned normalized latency result is verified.First Claude cost/performance test before Opus 5; escalate only when the premium pass changes the accepted result.Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive]2026-07-25
Claude Opus 4.8Anthropichistorical
Still available, but superseded by Opus 5 for current premium comparisons.
1M$5.00 / 1M$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25
GPT-5.5OpenAIactive
Generally available prior-generation OpenAI model retained for existing integrations and comparisons.
1.05M API; 400K Codex$5.00 / 1M$30.00 / 1Mnot verifiednot verifiednot verifiednot verifiedPrimary coding seat while ChatGPT/Codex limits fit the workload.OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset2026-06-28

K2.7 metrics are vendor-reported improvements over K2.6. Missing independent and AIHackers-owned results remain not verified.

Kimi Code And API Routing

Kimi’s official coding surfaces have moved beyond the old “Kimi k2.5 only” framing. Current Kimi docs describe Kimi Code and API paths around K3 plus K2.7 Code, with k3, kimi-for-coding, and kimi-for-coding-highspeed appearing in coding-tool contexts and K2.7 Code remaining available through the API.

That means older Kimi k2.5 guides still matter for historical/free-access intent, while newest-Kimi searches should land on K3 and lower-cost Kimi coding searches should land here or on the Kimi access guide:

  • Access Kimi for membership, promo-check, and access routing.
  • Kimi Code for the official coding membership and tool path.
  • Kimi K3 for the newest flagship, 1M context, and K3 benchmark evidence.
  • Kimi K2.5 for older K2.5 context and legacy free-access paths.

GLM-5.2 vs Kimi K2.7 Code

Decision pointKimi K2.7 CodeGLM-5.2
Context256K-class1M
Base input pricing$0.19 cache hit / $0.95 cache miss$1.40 input / $0.26 cached input
Output pricing$4.00$4.40
Speed laneHighSpeed at 2x token priceGLM Coding Plan or API path
Tool fitKimi API, Kimi Code, Claude Code-compatible Kimi paths, OpenAI-compatible APIZ.AI-supported coding tools and OpenAI/Anthropic-compatible endpoints
Best first testKimi-native coding, multimodal tool calls, cheaper cache-hit workloadsWhole-repo context and long-horizon refactors

For the full decision table, use GLM-5.2 vs Kimi K2.6/K2.7.

Eval Plan

Before switching your daily coding lane:

TestWhat to checkPass signal
Coding bugOne real failing testCorrect small patch, no broad churn
Long contextMedium repo or feature areaMaintains constraints through follow-up turns
Tool callsMulti-step agent task with tool resultsPreserves reasoning/tool context and recovers from failures
MultimodalScreenshot, UI clip, or video taskUses visual input concretely rather than guessing
CostCache-hit and cache-miss workloadTotal cost per accepted patch beats your fallback

Sources


Last verified: July 18, 2026. Kimi model IDs, pricing, membership aliases, benchmark claims, K3 weight status, and HighSpeed availability can change quickly.