Budget models are useful when their lower token price survives retries, long outputs, and human review. The relevant metric is cost per successful task, not the cheapest input row.
Current Shortlist
| Model | Input / output per 1M | Context | Current role |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | $0.14 cache miss, $0.0028 cache hit / $0.28 | 1M | Direct-API and MIT-weight value candidate; verify accepted-task behavior |
| Kimi K2.7 Code | $0.95 cache miss, $0.19 cache hit / $4.00 | 256K | Cheaper routine Kimi coding API lane |
| Gemini 3 Flash | $0.50 / $3.00 | ~1M | Gemini high-context value lane |
| MiniMax M3 | Standard PAYG starts at $0.60 / $2.40 | 1M | Coding-agent value lane to test |
| Xiaomi MiMo-V2.5 | Overseas list starts at $0.14 / $0.28 | 1M | Low-cost long-context and open-weight evaluation lane |
| GPT-5.6 Luna | $0.20 / $1.20 | 1.05M | Generally available high-volume GPT-5.6 API tier; long-context rates apply above 272K input |
Subscription and Token Plan quotas are not API prices. Compare them separately and verify checkout before purchase.
Best Starting Points
Kimi K2.7 Code
Use K2.7 Code for lower-cost Kimi coding searches. Moonshot documents base and HighSpeed IDs, 256K context, required thinking mode, multimodal input, and automatic context caching. Use Kimi K3 for newest-Kimi, 1M-context, or K3 benchmark intent.
Kimi’s published improvements over K2.6 are vendor evidence. AIHackers has not verified a normalized independent K2.7 coding score or run a controlled repository comparison.
Gemini 3 Flash
Use Gemini when its API or Vertex AI route, context size, and tool support fit. Verify the current model stage, regional availability, and exact free/paid rates; preview and free-tier limits can change faster than this page.
MiniMax M3 and Xiaomi MiMo
Both are newer low-cost, long-context lanes worth testing. Their strongest coding claims are vendor-reported, so require the same real repository tasks and pass conditions used for Kimi, GLM, or Claude.
GPT-5.6 Luna
Luna is the OpenAI value default for bounded Codex work and high-volume API tasks. Plus guidance lists approximately 250–2,000 Luna messages per five hours, versus 25–200 Terra and 10–100 Sol, with possible weekly limits. OpenAI cut Luna’s standard short-context API rate by 80% on July 30; the value-default analysis keeps subscription allowance and API billing separate.
DeepSeek V4 Flash 0731
DeepSeek’s direct API maps deepseek-v4-flash to the official 0731 release. Artificial Analysis scores it 50, one point behind Luna max, and places it on the evaluator’s Intelligence-versus-Cost frontier. The 0731 analysis keeps that independent finding separate from vendor benchmarks, the older preview’s 0/10 security result, and AIHackers’ site-owned not-run state.
Evidence Table
benchmark artifact
Budget and Value Model Evidence
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | DeepSeek | active Official 0731 release available through DeepSeek's direct API; MIT-licensed weights are published separately. OpenCode's promotional free route has unverified backend revision identity. | 1M | $0.14 / 1M | $0.28 / 1M | Artificial Analysis reports Intelligence Index 50, one point behind GPT-5.6 Luna at max effort; this is shortlist evidence, not a universal coding-quality claim. | DeepSeek documents tool calls, Responses API, and Anthropic-compatible access; real harness behavior remains workload-specific. |
| Independent write-up reports about 206M total output tokens for the Intelligence Index; production speed and accepted-task efficiency are not verified. | Direct-API and open-weight value candidate; test accepted-task cost, token use, hallucinations, and tool behavior before production routing. | DeepSeek V4 Flash 0731 update [archive], DeepSeek V4 Flash 0731 model card [archive], DeepSeek API pricing [archive], Artificial Analysis DeepSeek V4 Flash 0731, OpenCode Zen pricing [archive], AIHackers app-hacking field test | 2026-08-01 |
| GPT-5.6 Luna | OpenAI | active Generally available through paid ChatGPT Work/Codex plans and the OpenAI API. | 1.05M | $0.20 / 1M | $1.20 / 1M | Artificial Analysis reports 75 on its Coding Agent Index at max effort; BenchLM ranks Luna #6/129 in coding with an Estimated overall position. | Generally available in API and paid Codex plans. |
| Vendor-positioned as fastest; measured production latency is not verified. | Generally available lowest-cost GPT-5.6 tier; verify quality and cost per accepted task. | OpenAI GPT-5.6 general availability [archive], OpenAI GPT-5.6 price update [archive], OpenAI API pricing [archive], OpenAI GPT-5.6 Luna model page [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, BenchLM GPT-5.6 Luna profile [archive], OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card | 2026-08-01 |
| Kimi K2.7 Code | Moonshot AI | active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| Gemini 3 Flash | active Current Gemini value lane where Gemini API or Vertex AI fits. | 1.05M input | $0.50 / 1M | $3.00 / 1M | not verified | Function calling and code execution supported. | not verified | Preview model positioned for lower latency; independent value not imported. | High-context value lane when Gemini API or Vertex AI fits. | Gemini API models, Gemini API pricing, Artificial Analysis: Gemini 3 Flash, LMArena leaderboard dataset | 2026-05-26 | |
| GLM-5.2 | Z.AI | active Current Z.AI flagship coding model and supported-tool value lane. | 1M | $1.40 / 1M | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass. | Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
DeepSeek 0731 and GPT-5.6 Luna are shortlist candidates; local quality, latency, and cost-per-successful-task results remain not-run or not verified.
GLM-5.2 sits above the strict $1 input threshold at $1.40, but it is the relevant value comparison because it adds 1M context and independent Artificial Analysis evidence.
What Not to Compare Directly
- SWE-bench Verified and SWE-Bench Pro.
- A vendor’s internal preference test and an independent aggregate index.
- Base API price and a monthly coding-plan quota.
- Cache-hit price and uncached first-run price.
- A model’s context limit and its ability to use that context accurately.
Do not claim one model delivers a percentage of another model’s total capability from a single benchmark.
Repository Test
Run each accessible model on:
| Task | Pass condition |
|---|---|
| Repository map | Correct modules and commands; no invented files |
| Bug fix | Minimal patch and relevant passing tests |
| Refactor | Preserved behavior and repository conventions |
| Review | Concrete file-grounded findings |
Record exact model ID, settings, prompt, tokens, cache behavior, retries, latency, accepted patch, and review time. Models that are inaccessible or provider-managed should be labeled accordingly rather than assigned estimated results.
Historical Budget Context
Kimi K2.5 remains useful for older free-hosting and integration searches. Kimi K2.6 remains relevant for compatible non-thinking and multimodal workflows. Neither should replace K3 for newest-Kimi intent or K2.7 Code for cheaper Kimi coding.
Older access pages can retain K2.5 when that exact hosted model is still offered. They must not describe it as Kimi’s latest coding model.
Verdict
- Start with Kimi K2.7 Code when Kimi’s API and 256K context fit.
- Start with DeepSeek V4 Flash 0731 when direct-API price or MIT weights are the main constraint.
- Test GLM-5.2 when 1M context and independent shortlist evidence justify slightly higher input pricing.
- Evaluate Gemini 3 Flash, MiniMax M3, or Xiaomi MiMo when their provider/tool route fits.
- Test GPT-5.6 Luna when the OpenAI API or paid Codex route fits the workload.
- Escalate to Opus 4.8 only when a premium second pass materially improves the outcome.
Sources
- Kimi: K2.7 Code quickstart and pricing
- Z.AI: GLM-5.2 and pricing
- Google: Gemini models and pricing
- OpenAI: GPT-5.6 general availability
- OpenAI: July 30 price update and API pricing
- Artificial Analysis: GLM-5.2
- DeepSeek: 0731 model card and direct pricing
- Artificial Analysis: DeepSeek V4 Flash 0731
Related links
- /models/kimi-k2.7-code/
- /models/kimi-k3/
- /models/glm-5.2/
- /models/gpt-5-6/
- /posts/gpt-5-6-luna-price-cut/
- /posts/deepseek-v4-flash-0731-value-frontier/
- /compare/models/mid-range/
- /value/free-stack/
Last verified: August 1, 2026. Prices, model stages, cache terms, benchmark positions, and free routes can change independently.