DeepSeek V4 Flash 0731 is the direct-API and open-weight value model to test first when OpenAI is not a requirement. It combines a 1M-token context window, MIT-licensed weights, and a first-party API at $0.14 cache-miss input / $0.0028 cache-hit input / $0.28 output per million tokens.
That price-performance position is credible shortlist evidence. It is not proof that 0731 is the cheapest accepted result for every coding, agent, or knowledge task.
What the 0731 release is
DeepSeek’s model card calls 0731 the official release that supersedes the earlier V4 Flash preview. The repository publishes the weights under the MIT License. DeepSeek’s direct pricing page maps the API name deepseek-v4-flash to model version DeepSeek-V4-Flash-0731, with 1M context and a maximum 384K output.
Keep three identities separate:
- DeepSeek V4 Flash 0731 through DeepSeek’s API: the current first-party hosted model and price authority.
- DeepSeek V4 Flash 0731 weights: the MIT-licensed release for self-hosting, with infrastructure cost and operational behavior outside the API price.
- DeepSeek V4 Flash (Preview): the earlier model used by older evaluations, including AIHackers’ June security field-test record.
The model card’s agent benchmarks are vendor evidence. DeepSeek used the not-yet-released minimal mode of its own harness, max reasoning effort, temperature 1.0, and top_p=0.95 for public code-agent tasks. Those results should not be projected onto another harness or lower effort setting.
Direct API economics
| Direct DeepSeek API item | Price per 1M tokens |
|---|---|
| Cache-miss input | $0.14 |
| Cache-hit input | $0.0028 |
| Output | $0.28 |
DeepSeek says a future peak-hours policy will charge 2× the regular rates, but the effective date is subject to a later announcement. Until that date exists, the current list price is the defensible comparison anchor; production budgets should still include a recheck trigger.
The cache-hit discount is unusually large, but it only applies when the provider records a hit. First-run context, unstable prefixes, tool output, reasoning, retries, and cache misses can dominate real bills. Self-hosting the weights also replaces token list price with hardware, throughput, reliability, and operations costs.
Independent evidence: frontier value, not universal quality
Artificial Analysis scored DeepSeek V4 Flash 0731 at 50 on Intelligence Index v4.1, one point behind GPT-5.6 Luna at max effort (51). The evaluator explicitly places 0731 on its Intelligence-versus-Cost-per-Task Pareto frontier and reports its first-party cost per task at roughly 60% below post-cut Luna max.
That finding belongs to Artificial Analysis. It is not an AIHackers benchmark, CAR result, or claim that the two models are interchangeable. One composite index can justify a trial; it cannot establish tool reliability, instruction fit, security behavior, latency, or review cost on your repository.
The caveats are material. Artificial Analysis reports that 0731 used about 206 million output tokens across the Intelligence Index run—12% fewer than the preview, but still a high absolute total. Its AA-Omniscience hallucination rate was 84%, while accuracy remained 37%. A low token price can coexist with verbose reasoning and unreliable factual answers.
benchmark artifact
DeepSeek 0731 Value-Frontier Evidence
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | DeepSeek | active Official 0731 release available through DeepSeek's direct API; MIT-licensed weights are published separately. OpenCode's promotional free route has unverified backend revision identity. | 1M | $0.14 / 1M | $0.28 / 1M | Artificial Analysis reports Intelligence Index 50, one point behind GPT-5.6 Luna at max effort; this is shortlist evidence, not a universal coding-quality claim. | DeepSeek documents tool calls, Responses API, and Anthropic-compatible access; real harness behavior remains workload-specific. |
| Independent write-up reports about 206M total output tokens for the Intelligence Index; production speed and accepted-task efficiency are not verified. | Direct-API and open-weight value candidate; test accepted-task cost, token use, hallucinations, and tool behavior before production routing. | DeepSeek V4 Flash 0731 update [archive], DeepSeek V4 Flash 0731 model card [archive], DeepSeek API pricing [archive], Artificial Analysis DeepSeek V4 Flash 0731, OpenCode Zen pricing [archive], AIHackers app-hacking field test | 2026-08-01 |
| GPT-5.6 Luna | OpenAI | active Generally available through paid ChatGPT Work/Codex plans and the OpenAI API. | 1.05M | $0.20 / 1M | $1.20 / 1M | Artificial Analysis reports 75 on its Coding Agent Index at max effort; BenchLM ranks Luna #6/129 in coding with an Estimated overall position. | Generally available in API and paid Codex plans. |
| Vendor-positioned as fastest; measured production latency is not verified. | Generally available lowest-cost GPT-5.6 tier; verify quality and cost per accepted task. | OpenAI GPT-5.6 general availability [archive], OpenAI GPT-5.6 price update [archive], OpenAI API pricing [archive], OpenAI GPT-5.6 Luna model page [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, BenchLM GPT-5.6 Luna profile [archive], OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card | 2026-08-01 |
| GLM-5.2 | Z.AI | active Current Z.AI flagship coding model and supported-tool value lane. | 1M | $1.40 / 1M | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | July value pick to test for supported coding-tool workflows; keep Opus/GPT for final arbitration until local evals pass. | Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
Independent, vendor, historical-preview, and site-owned evidence remain separate. Scores do not establish universal quality or cost per accepted result.
The old 0/10 result does not evaluate 0731
AIHackers’ June app-hacking field test records 0 solves in 10 runs for DeepSeek V4 Flash, with an average $0.08 per run. That result is preserved as a historical security-task record for the pre-0731 preview version. It does not evaluate the July 31 release and cannot be relabeled as a 0731 failure.
The boundary cuts both ways: the newer independent score does not erase the old security result, and the old result does not predict the new model’s security performance. A current comparison requires the same vulnerable app, harness, budget, effort, acceptance criteria, and exact model version.
Direct API is not OpenCode’s free promotion
OpenCode Zen lists DeepSeek V4 Flash Free as a limited-time promotional route. OpenCode does not explicitly identify that route’s underlying revision as 0731 in the checked source. Treat it as a separate provider product with its own billing, limits, routing, and data terms—not as evidence of free access to DeepSeek’s direct 0731 API.
If exact version identity matters, use DeepSeek’s direct deepseek-v4-flash endpoint or self-host the named 0731 weights. Do not infer backend identity from a display label.
Recommendation
Start with DeepSeek V4 Flash 0731 for high-volume direct-API work or when MIT weights and self-hosting flexibility matter. Start with GPT-5.6 Luna when the OpenAI API, Codex subscription workflow, multimodal input, or OpenAI tooling is the constraint. Escalate to stronger models when retries, hallucinations, tool failures, or human review erase the price advantage.
AIHackers has not run a 0731 repository evaluation or calculated site-owned Cost per Accepted Result. That state is not-run. Record exact version, provider, effort, cache hits, output and reasoning tokens, retries, latency, receipts, accepted tasks, and review time before declaring a production winner. The Smart Spend guide and benchmark mini-eval provide the routing and evaluation method.
Sources and evidence state
- DeepSeek: 0731 model card and weights (Archive) — official release identity, MIT license, and vendor benchmark settings;
validated - DeepSeek: models and pricing (Archive) — direct version mapping, context, and prices;
validated - DeepSeek: updates (Archive) — July 31 release record;
validated - Artificial Analysis: DeepSeek V4 Flash 0731 — independent score, Pareto finding, output-token total, and hallucination evidence;
validated, archivearchive-pending - OpenCode: Zen pricing (Archive) — limited-time named free route without explicit 0731 backend confirmation;
validated
Related links
- /value/smart-spend/
- /compare/models/budget-tier/
- /models/
- /compare/
- /lab/investigations/llm-app-hacking-field-test/
- /posts/gpt-5-6-luna-price-cut/
Checked August 1, 2026. Pricing, peak-hour policy, provider routing, benchmark positions, and model availability can change independently.