Budget models are useful when their lower token price survives retries, long outputs, and human review. The relevant metric is cost per successful task, not the cheapest input row.

Current Shortlist

ModelInput / output per 1MContextCurrent role
DeepSeek V4 Flash 0731$0.14 cache miss, $0.0028 cache hit / $0.281MDirect-API and MIT-weight value candidate; verify accepted-task behavior
Kimi K2.7 Code$0.95 cache miss, $0.19 cache hit / $4.00256KCheaper routine Kimi coding API lane
Gemini 3 Flash$0.50 / $3.00~1MGemini high-context value lane
MiniMax M3Standard PAYG starts at $0.60 / $2.401MCoding-agent value lane to test
Xiaomi MiMo-V2.5Overseas list starts at $0.14 / $0.281MLow-cost long-context and open-weight evaluation lane
GPT-5.6 Luna$0.20 / $1.201.05MGenerally available high-volume GPT-5.6 API tier; long-context rates apply above 272K input

Subscription and Token Plan quotas are not API prices. Compare them separately and verify checkout before purchase.

Best Starting Points

Kimi K2.7 Code

Use K2.7 Code for lower-cost Kimi coding searches. Moonshot documents base and HighSpeed IDs, 256K context, required thinking mode, multimodal input, and automatic context caching. Use Kimi K3 for newest-Kimi, 1M-context, or K3 benchmark intent.

Kimi’s published improvements over K2.6 are vendor evidence. AIHackers has not verified a normalized independent K2.7 coding score or run a controlled repository comparison.

Gemini 3 Flash

Use Gemini when its API or Vertex AI route, context size, and tool support fit. Verify the current model stage, regional availability, and exact free/paid rates; preview and free-tier limits can change faster than this page.

MiniMax M3 and Xiaomi MiMo

Both are newer low-cost, long-context lanes worth testing. Their strongest coding claims are vendor-reported, so require the same real repository tasks and pass conditions used for Kimi, GLM, or Claude.

GPT-5.6 Luna

Luna is the OpenAI value default for bounded Codex work and high-volume API tasks. Plus guidance lists approximately 250–2,000 Luna messages per five hours, versus 25–200 Terra and 10–100 Sol, with possible weekly limits. OpenAI cut Luna’s standard short-context API rate by 80% on July 30; the value-default analysis keeps subscription allowance and API billing separate.

DeepSeek V4 Flash 0731

DeepSeek’s direct API maps deepseek-v4-flash to the official 0731 release. Artificial Analysis scores it 50, one point behind Luna max, and places it on the evaluator’s Intelligence-versus-Cost frontier. The 0731 analysis keeps that independent finding separate from vendor benchmarks, the older preview’s 0/10 security result, and AIHackers’ site-owned not-run state.

Evidence Table

benchmark artifact

Budget and Value Model Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
DeepSeek V4 Flash 0731DeepSeekactive
Official 0731 release available through DeepSeek's direct API; MIT-licensed weights are published separately. OpenCode's promotional free route has unverified backend revision identity.
1M$0.14 / 1M$0.28 / 1MArtificial Analysis reports Intelligence Index 50, one point behind GPT-5.6 Luna at max effort; this is shortlist evidence, not a universal coding-quality claim.DeepSeek documents tool calls, Responses API, and Anthropic-compatible access; real harness behavior remains workload-specific.
  • Artificial Analysis Intelligence Index: 50 (independent)
  • AA-Omniscience hallucination rate: 84% (independent)
  • Historical security field test: 0/10; pre-0731 preview only (site-owned historical)
  • AIHackers 0731 repo eval: not-run (site-owned)
Independent write-up reports about 206M total output tokens for the Intelligence Index; production speed and accepted-task efficiency are not verified.Direct-API and open-weight value candidate; test accepted-task cost, token use, hallucinations, and tool behavior before production routing.DeepSeek V4 Flash 0731 update [archive], DeepSeek V4 Flash 0731 model card [archive], DeepSeek API pricing [archive], Artificial Analysis DeepSeek V4 Flash 0731, OpenCode Zen pricing [archive], AIHackers app-hacking field test2026-08-01
GPT-5.6 LunaOpenAIactive
Generally available through paid ChatGPT Work/Codex plans and the OpenAI API.
1.05M$0.20 / 1M$1.20 / 1MArtificial Analysis reports 75 on its Coding Agent Index at max effort; BenchLM ranks Luna #6/129 in coding with an Estimated overall position.Generally available in API and paid Codex plans.
  • Artificial Analysis Intelligence Index: 51 at max effort (independent)
  • Artificial Analysis Coding Agent Index: 75 at max effort (independent)
  • BenchLM coding rank: #6/129; overall Estimated (independent aggregator)
  • SWE-bench Pro: 62.7% (vendor)
  • AIHackers repo eval: not-run (site-owned)
Vendor-positioned as fastest; measured production latency is not verified.Generally available lowest-cost GPT-5.6 tier; verify quality and cost per accepted task.OpenAI GPT-5.6 general availability [archive], OpenAI GPT-5.6 price update [archive], OpenAI API pricing [archive], OpenAI GPT-5.6 Luna model page [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, BenchLM GPT-5.6 Luna profile [archive], OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card2026-08-01
Kimi K2.7 CodeMoonshot AIactive
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Gemini 3 FlashGoogleactive
Current Gemini value lane where Gemini API or Vertex AI fits.
1.05M input$0.50 / 1M$3.00 / 1Mnot verifiedFunction calling and code execution supported.not verifiedPreview model positioned for lower latency; independent value not imported.High-context value lane when Gemini API or Vertex AI fits.Gemini API models, Gemini API pricing, Artificial Analysis: Gemini 3 Flash, LMArena leaderboard dataset2026-05-26
GLM-5.2Z.AIactive
Current Z.AI flagship coding model and supported-tool value lane.
1M$1.40 / 1M$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass.Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28

DeepSeek 0731 and GPT-5.6 Luna are shortlist candidates; local quality, latency, and cost-per-successful-task results remain not-run or not verified.

GLM-5.2 sits above the strict $1 input threshold at $1.40, but it is the relevant value comparison because it adds 1M context and independent Artificial Analysis evidence.

What Not to Compare Directly

  • SWE-bench Verified and SWE-Bench Pro.
  • A vendor’s internal preference test and an independent aggregate index.
  • Base API price and a monthly coding-plan quota.
  • Cache-hit price and uncached first-run price.
  • A model’s context limit and its ability to use that context accurately.

Do not claim one model delivers a percentage of another model’s total capability from a single benchmark.

Repository Test

Run each accessible model on:

TaskPass condition
Repository mapCorrect modules and commands; no invented files
Bug fixMinimal patch and relevant passing tests
RefactorPreserved behavior and repository conventions
ReviewConcrete file-grounded findings

Record exact model ID, settings, prompt, tokens, cache behavior, retries, latency, accepted patch, and review time. Models that are inaccessible or provider-managed should be labeled accordingly rather than assigned estimated results.

Historical Budget Context

Kimi K2.5 remains useful for older free-hosting and integration searches. Kimi K2.6 remains relevant for compatible non-thinking and multimodal workflows. Neither should replace K3 for newest-Kimi intent or K2.7 Code for cheaper Kimi coding.

Older access pages can retain K2.5 when that exact hosted model is still offered. They must not describe it as Kimi’s latest coding model.

Verdict

  • Start with Kimi K2.7 Code when Kimi’s API and 256K context fit.
  • Start with DeepSeek V4 Flash 0731 when direct-API price or MIT weights are the main constraint.
  • Test GLM-5.2 when 1M context and independent shortlist evidence justify slightly higher input pricing.
  • Evaluate Gemini 3 Flash, MiniMax M3, or Xiaomi MiMo when their provider/tool route fits.
  • Test GPT-5.6 Luna when the OpenAI API or paid Codex route fits the workload.
  • Escalate to Opus 4.8 only when a premium second pass materially improves the outcome.

Sources


Last verified: August 1, 2026. Prices, model stages, cache terms, benchmark positions, and free routes can change independently.