DeepSeek V4.1 Flash is a strong API value candidate to evaluate for coding, research, and high-volume work. Native vision broadens its use, and MIT weights provide a separate deployment option. Its low token prices justify a trial; they do not establish lower cost per accepted result than Luna or a stronger model.

Use the September stack to choose an access lane. The recommendation here concerns model capability and metered economics. Temporary subscription boosts do not determine its ranking.

What changed

The official model card describes a multimodal mixture-of-experts model with 552B backbone parameters, native image/text input, text output, and a one-million-token context. The weights carry the MIT license. Native image support makes screenshots and visual documents evaluation targets; it does not prove dependable interpretation of every diagram or interface.

DeepSeek’s release announcement identifies the new Flash release. Keep its vendor evaluations separate from independent results: a model-card score reflects the named setup, reasoning configuration, and harness. It is not an AIHackers repository test. Self-hosting also requires hardware, serving capacity, and operational work; MIT weights do not mean free inference.

Access and direct prices

Use deepseek-flash for DeepSeek’s direct API. The current pricing page maps it to V4.1 Flash. Legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names remain accepted but route to V4.1 Flash and use Flash billing. Treat that compatibility routing as transitional, not a way to pin the retired backend; no guaranteed alias lifetime is established here.

Prices below are US dollars per million tokens, checked September 14:

Direct API rateOff-peakPeak
Input, cache miss$0.15$0.30
Input, cache hit$0.003$0.006
Output$0.60$1.20

Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. All other hours, including weekends, are off-peak. For one million uncached input tokens and one million output tokens, that is $0.75 off-peak or $1.50 peak, before any separately applicable charges. This arithmetic is a token bill example, not a task forecast.

Cache savings require recorded hits. Repeated prefixes may help, while changing context, first requests, and output-heavy reasoning alter the mix. Track misses and hits separately rather than applying the cache-read price to every input token. Check provider identity too: OpenCode and Command Code are distinct services with their own accounting.

Independent evidence

Artificial Analysis’s V4.1 Flash Max profile, checked September 14 under the current v4.3 index, reports 40, approximately $0.27 per Intelligence Index task, and 250M output tokens across its evaluation. It reports roughly 214 output tokens per second. That supports testing price/performance while highlighting substantial output use; generation speed is not end-to-end task latency.

Those figures describe the evaluator’s reasoning/Max run and price assumptions, not our receipts. The v4.3 methodology differs from August’s v4.1. Never subtract the old 0731 score of 50 from today’s 40 and call it a ten-point model regression. A defensible change comparison needs both revisions evaluated with the same current methodology.

benchmark artifact

DeepSeek V4.1 Flash evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
DeepSeek V4.1 FlashDeepSeekactive
Direct deepseek-flash API; native vision; MIT weights. Legacy Flash names now route here.
1M$0.15 / 1M$0.60 / 1Mnot verifiednot verified
  • AA Intelligence Index v4.3 (Max): 40 (independent)
  • AIHackers CAR: not-run (site-owned)
not verifiedStrong API value candidate; evaluate accepted results, tool behavior, and vision. Historical 0731/preview tests do not evaluate V4.1.DeepSeek V4.1 Flash model card, DeepSeek September pricing and aliases, Artificial Analysis V4.1 Flash Max profile2026-09-14

September v4.3 independent evidence; AIHackers CAR remains not-run.

Historical results and limitations

Our 0731 article preserves that release’s sources and scores. The June app-hacking field test preserves the older preview’s 0/10 security result. Neither record evaluates V4.1 Flash. Newer results do not erase old failures, and old failures cannot be relabeled as a new-model test.

Before routing production work, test tool arguments, instruction following, factual grounding, image interpretation, and recovery after a failed step. Count retries and human correction. A low token bill can lose its advantage when the output requires substantial repair, and a long context window does not establish reliable recall of everything inside it.

Recommendation and sources

Try V4.1 Flash where direct metering or self-hosting flexibility matters. Retain Luna where your OpenAI workflow and accepted results justify it; use Astra Low as a separate planning candidate. Command Code and OpenCode offer temporary V4.1 allowances ending September 20, detailed in the subscription comparison; these are limited coding-access offers, not unlimited usage or general API credit.

AIHackers’ V4.1 evaluation and CAR remain not-run. Sources linked above are vendor documentation and independent measurements, checked September 14. The release log timed out during the bounded fetch; release identity is supported by the accessible announcement and model card. Archives remain archive-pending until replay validation.