Kimi K2.7 Code is now the cheaper routine Kimi coding lane, not the newest Kimi model. For latest-Kimi, 1M-context, and K3 launch intent, use Kimi K3. Use this page when the search is specifically about K2.7 Code pricing, HighSpeed, 256K coding workflows, or a lower-cost Kimi API route.
Moonshot’s Kimi API docs still list kimi-k2.7-code and kimi-k2.7-code-highspeed IDs, 256K-class context, multimodal input, thinking required, automatic context caching, ToolCalls, JSON Mode, and Partial Mode.
Use K2.7 Code when you want cheaper Kimi coding economics. Use K3 when 1M context, native flagship behavior, or the newest Kimi benchmark evidence is the reason you are testing Kimi at all. Keep K2.6 in the comparison set because it is still a relevant multimodal/API lane and exact-match search term.
Quick Facts
| Spec | Kimi K2.7 Code |
|---|---|
| Provider | Moonshot AI / Kimi |
| Model IDs | kimi-k2.7-code, kimi-k2.7-code-highspeed |
| Context window | 262,144 tokens, commonly described as 256K |
| Thinking | Required for K2.7 Code; disabling thinking throws an error |
| Base API pricing | $0.19 cache-hit input / $0.95 cache-miss input / $4.00 output per 1M tokens |
| HighSpeed pricing | $0.38 cache-hit input / $1.90 cache-miss input / $8.00 output per 1M tokens |
| Useful for | Coding agents, multimodal coding tasks, Kimi-native API workflows |
| Caveat | K3 is the newest Kimi flagship; vendor benchmark claims and speed claims need workload-level checks |
K2.7 Code vs K2.6
| Question | K2.7 Code | K2.6 |
|---|---|---|
| Current role | Cheaper routine Kimi coding lane | Prior Kimi release and still relevant comparison lane |
| Model IDs | kimi-k2.7-code, kimi-k2.7-code-highspeed | kimi-k2.6 |
| Context | 256K-class | 256K-class |
| Thinking mode | Required | Can be enabled or disabled |
| Cache-hit input | $0.19 / 1M | $0.16 / 1M |
| Cache-miss input | $0.95 / 1M | $0.95 / 1M |
| Output | $4.00 / 1M | $4.00 / 1M |
| Fast variant | HighSpeed at $0.38 / $1.90 / $8.00 | No separate HighSpeed price page in the current docs |
The short version: choose K2.7 Code for lower-cost Kimi coding tests and K2.6 when you need non-thinking mode, have existing integrations, or are comparing exact K2.6 pricing/search intent.
Benchmark Evidence
Kimi’s June release notes report these changes relative to K2.6:
| Benchmark or measure | Vendor-reported delta |
|---|---|
| Program-Bench | +10.4% |
| MCP Mark Verified | +11.4% |
| SWE Marathon | +76.2% |
| Reasoning-token use | 30% lower |
These are relative vendor-reported deltas, not absolute scores. They cannot be compared directly with SWE-bench Verified, SWE-Bench Pro, Terminal-Bench, or Artificial Analysis. AIHackers has not verified a normalized independent K2.7 score or run a controlled repository evaluation.
benchmark artifact
Kimi K2.7 and Current Alternatives
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi K2.7 Code | Moonshot AI | active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| GLM-5.2 | Z.AI | active Current Z.AI flagship coding model and supported-tool value lane. | 1M | $1.40 / 1M | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass. | Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| Claude Sonnet 5 | Anthropic | active Generally available across Claude plans, Claude Code, the Claude API, GitHub Copilot, and supported AWS paths. | 1M | $2.00 / 1M | $10.00 / 1M | Anthropic reports substantial coding and agentic gains over Sonnet 4.6; independent normalized results are pending. | Available in Claude Code and the Claude API; adaptive thinking is on by default. |
| No site-owned normalized latency result is verified. | First Claude cost/performance test before Opus 5; escalate only when the premium pass changes the accepted result. | Anthropic Claude Sonnet 5 launch [archive], Claude Sonnet 5 migration guide [archive], GitHub Copilot Claude Sonnet 5 launch [archive], Claude Sonnet 5 on AWS [archive], Artificial Analysis: Claude Opus 5 [archive] | 2026-07-25 |
| Claude Opus 4.8 | Anthropic | historical Still available, but superseded by Opus 5 for current premium comparisons. | 1M | $5.00 / 1M | $25.00 / 1M | Historical premium Claude baseline; use Opus 5 for new task-level comparisons. | Still available for pinned integrations; new Claude premium routing should test Opus 5. |
| Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary. | Historical premium baseline. Use Claude Opus 5 for current Claude premium routing. | Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard | 2026-07-25 |
| GPT-5.5 | OpenAI | active Generally available prior-generation OpenAI model retained for existing integrations and comparisons. | 1.05M API; 400K Codex | $5.00 / 1M | $30.00 / 1M | not verified | not verified | not verified | not verified | Primary coding seat while ChatGPT/Codex limits fit the workload. | OpenAI GPT-5.5 API model page, OpenAI GPT-5.5 ChatGPT limits, Artificial Analysis: GPT-5.5, LMArena leaderboard dataset | 2026-06-28 |
K2.7 metrics are vendor-reported improvements over K2.6. Missing independent and AIHackers-owned results remain not verified.
Kimi Code And API Routing
Kimi’s official coding surfaces have moved beyond the old “Kimi k2.5 only” framing. Current Kimi docs describe Kimi Code and API paths around K3 plus K2.7 Code, with k3, kimi-for-coding, and kimi-for-coding-highspeed appearing in coding-tool contexts and K2.7 Code remaining available through the API.
That means older Kimi k2.5 guides still matter for historical/free-access intent, while newest-Kimi searches should land on K3 and lower-cost Kimi coding searches should land here or on the Kimi access guide:
- Access Kimi for membership, promo-check, and access routing.
- Kimi Code for the official coding membership and tool path.
- Kimi K3 for the newest flagship, 1M context, and K3 benchmark evidence.
- Kimi K2.5 for older K2.5 context and legacy free-access paths.
GLM-5.2 vs Kimi K2.7 Code
| Decision point | Kimi K2.7 Code | GLM-5.2 |
|---|---|---|
| Context | 256K-class | 1M |
| Base input pricing | $0.19 cache hit / $0.95 cache miss | $1.40 input / $0.26 cached input |
| Output pricing | $4.00 | $4.40 |
| Speed lane | HighSpeed at 2x token price | GLM Coding Plan or API path |
| Tool fit | Kimi API, Kimi Code, Claude Code-compatible Kimi paths, OpenAI-compatible API | Z.AI-supported coding tools and OpenAI/Anthropic-compatible endpoints |
| Best first test | Kimi-native coding, multimodal tool calls, cheaper cache-hit workloads | Whole-repo context and long-horizon refactors |
For the full decision table, use GLM-5.2 vs Kimi K2.6/K2.7.
Eval Plan
Before switching your daily coding lane:
| Test | What to check | Pass signal |
|---|---|---|
| Coding bug | One real failing test | Correct small patch, no broad churn |
| Long context | Medium repo or feature area | Maintains constraints through follow-up turns |
| Tool calls | Multi-step agent task with tool results | Preserves reasoning/tool context and recovers from failures |
| Multimodal | Screenshot, UI clip, or video task | Uses visual input concretely rather than guessing |
| Cost | Cache-hit and cache-miss workload | Total cost per accepted patch beats your fallback |
Related links
- /compare/models/glm-5.2-vs-kimi-k2.6/ - GLM vs Kimi coding-model decision guide
- /models/glm-5.2/ - Z.AI’s current GLM coding model
- /tools/kimi-code/ - Official Kimi coding membership and tool path
- /value/kimi-access/ - Kimi access, membership, and promo-check guidance
- /value/smart-spend/ - Paid-stack routing strategy
- /compare/models/mid-range/ - Production spend-band comparison
Sources
- Kimi K2.7 Code quickstart (Archive)
- Kimi K2.7 Code pricing (Archive)
- Kimi K2.7 Code release notes (Archive)
- Kimi model list
- Kimi K3 quickstart (Archive)
- Kimi K2.6 quickstart
- Kimi K2.6 pricing
- Z.AI GLM-5.2 docs
Last verified: July 18, 2026. Kimi model IDs, pricing, membership aliases, benchmark claims, K3 weight status, and HighSpeed availability can change quickly.