The August value stack is a router, not a universal leaderboard. Buy one strong subscription lane, test one low-cost API lane, and escalate only when the cheaper route misses a prewritten acceptance rule. Token price, subscription capacity, benchmark cost, and accepted-result cost are different measures.
For the provider-spanning method behind that router, use the LLM cost-saving playbook to work through deterministic steps, access paths, caching, batching, promotions, and Cost per Accepted Result in order.
The 30-second stack
| Workload | Start here | Escalate when |
|---|---|---|
| Well-defined coding, research, coverage, transformation, and review groundwork | GPT-5.6 Luna High/XHigh (Extra High in the product UI) | The task remains ambiguous, the acceptance check fails, or errors are consequential |
| Planning, architecture, delegation, synthesis, and consequential review | GPT-5.6 Sol High/XHigh (Extra High in the product UI) | Use Max for one exceptional hard problem; Ultra only when independent parts benefit from subagents |
| API research, batch work, coverage, or coding-agent evaluation | DeepSeek V4 Flash 0731 | Test GLM-5.2 for supported-tool fit; test Kimi K3/Kimi Code for Kimi capability, weights, or membership fit |
| Sustained premium-seat usage | OpenAI Pro or Claude Max, after measuring the base seat | Choose only when avoided separate spend justifies the higher seat price; included use is still not API credit |
This is AIHackers’ recommendation, not OpenAI’s product taxonomy. OpenAI describes Luna as the clear, repeatable, high-volume lane; Terra as the everyday workhorse; and Sol as the complex, open-ended lane. Our stronger Luna-first rule applies only when “done” is explicit and the output is checked.
OpenAI may show account-specific referral promotions or a purchased full-reset control. Neither changes this August recommendation, price comparisons, benchmark interpretation, or CAR. Community-observed $8 Plus and $80 Pro 20x reset prices are not universal terms, and redemption is reported to start a new seven-day window. See the extra Codex usage referral guide and capacity guide for the separate mechanisms; AIHackers publishes no OpenAI referral link.
Subscription router: Luna first, Sol for uncertainty
For August, ChatGPT/Codex is the primary frontier subscription choice. Start a defined task on GPT-5.6 Luna High. Move to XHigh when the same bounded task needs more checking. Suitable work includes implementing an approved plan, extracting evidence, transforming content, adding coverage, running a focused review, or researching against a fixed question set.
Escalate to GPT-5.6 Sol High or XHigh for unclear requirements, architecture, strategy, multi-source synthesis, delegation design, security-sensitive judgment, or a Luna attempt that fails the acceptance rule. OpenAI’s own example uses Sol to resolve uncertainty and plan, then Luna to implement specified changes and tests. That is vendor guidance; whether it works economically for your repository still needs your evidence.
Reasoning level is another cost control. Max gives one model more time on a single exceptionally difficult problem. Ultra uses subagents for separable work and therefore consumes more total model work. Most jobs need neither. The rule is simple: first tighten the task and acceptance test, then increase effort, then change model.
OpenAI says Luna became 80% cheaper on July 30: standard short-context API pricing is now $0.20 input / $1.20 output per 1M tokens. Terra fell 20% to $2/$12. Subscription prices and quota budgets did not change, but Luna and Terra now consume fewer credits in Codex and ChatGPT Work. Do not turn that into a universal “token subsidy”: plan usage varies with context, reasoning, tools, caching, and the task.
OpenAI now also publishes this token-based Codex credit rate card:
| Model | Input / 1M tokens | Cached input / 1M tokens | Output / 1M tokens |
|---|---|---|---|
| GPT-5.6 Sol | 125 credits | 12.5 credits | 750 credits |
| GPT-5.6 Terra | 50 credits | 5 credits | 300 credits |
| GPT-5.6 Luna | 5 credits | 0.5 credits | 30 credits |
Keep that rate card separate from three other ledgers. A plan’s included capacity is not a purchased-credit balance; additional ChatGPT credits extend eligible use after included limits; and an API-key session is billed at standard API rates rather than this Codex credit table. The rates make relative credit burn legible, but they do not predict how many tasks a seat will finish.
OpenAI Plus is the normal $20 seat. Current guidance lists approximately 250–2,000 Luna messages per five hours, versus 25–200 Terra and 10–100 Sol, with possible weekly limits still applying. Pro starts at $100 with higher limits. Anthropic’s Claude Max likewise starts at $100 and offers 5× or 20× Pro usage. These can be valuable for sustained use, but neither provider’s included subscription capacity is purchased API credit, cash, or a transferable token balance.
API router: DeepSeek first, GLM and Kimi remain contenders
DeepSeek V4 Flash 0731 is the first API value candidate, not a proven winner. DeepSeek identifies it as the July 31 official release served through deepseek-v4-flash. First-party list price is $0.14 cache-miss input / $0.0028 cache-hit input / $0.28 output per 1M tokens. DeepSeek has announced future 2× peak-hours pricing but has not published the effective date.
The vendor’s agent results used max effort, top_p=0.95, temperature 1.0, and an unreleased DeepSeek Harness minimal mode. Artificial Analysis now labels its current independent result as 0731, an important change from the older April evidence, but AIHackers still has no accepted-task result for this model. Third-party endpoints may list lower prices; provider identity, quantization, uptime, tool behavior, retention, and routing make them separate products.
GLM-5.2 remains a credible API and supported-tool contender. Its current international API anchor is $1.40 input / $4.40 output, with a separate GLM Coding Plan for supported tools. The subscription quota is not general API credit. GLM’s best case is workflow fit: long context, supported coding tools, and a predictable overflow lane.
Kimi K3 and Kimi Code remain serious capability, open-weight, and subscription contenders. Moonshot has released K3 weights under the Kimi K3 License—say open weights, not generically open source. The global Kimi API surface lists $0.30 cache-hit input / $3 cache-miss input / $15 output per 1M tokens, while Kimi Code uses membership credits and provider-managed model IDs. At that managed API price, K3 is not the default value route unless its accepted-task economics beat Luna, DeepSeek, or GLM for your work. No primary source checked on August 1 established a stable free Kimi K3 API with an exact quota and expiry.
Cost per task versus accepted-result cost
Keep five ledgers separate: API token list price; subscription capacity and credit burn; an evaluator’s benchmark cost per task; a third-party field test’s cost per solve; and your own Cost per Accepted Result (CAR).
| Evidence checked Aug. 1 | Published setup and result | Published cost, tokens, or runtime | Boundary |
|---|---|---|---|
| GPT-5.6 Luna, max | AA Intelligence Index v4.1: 51; Coding Agent Index in Codex: 75 | $0.21/Intelligence task | Pre-dates the 80% cut and used the old $1/$6 price; do not silently reprice |
| GPT-5.6 Sol, max | AA Intelligence Index v4.1: 59; Coding Agent Index in Codex: 80 | $1.04/Intelligence task; about 15K output tokens/task | Benchmark cost, not repository acceptance or subscription economics |
| DeepSeek V4 Flash 0731, effort not disclosed in AA’s public write-up | AA Intelligence Index v4.1: 50 | Exact Index cost/task not exposed in the article | Current 0731 evidence; DeepSeek’s separate vendor benchmarks used max effort and its unreleased harness |
| GLM-5.2, max | AA Intelligence Index v4.1: 51 | $0.46/task; 43K output, including 37K reasoning tokens | June 16 snapshot; no site-owned result |
| Kimi K3, effort not disclosed in AA’s public write-up | AA-Briefcase Elo 1543 on the private agentic knowledge-work set | $10.57/task; 120K output; 83 turns; 56.4 minutes | A different benchmark, not token list price or CAR |
The June app-hacking field test is another separate measure. An older, pre-0731 DeepSeek V4 Flash recorded 0 solves in 10 runs; V4 Pro recorded 3/10 and a task-specific $0.62 per solve. The harness, vulnerable app, budget, version, and acceptance rule constrain those numbers. They neither condemn 0731 nor prove V4 Pro is a general bargain.
Small paid and trial ways to test before committing
CAR gate, recheck triggers, sources, and related links
AIHackers defines CAR as:
CAR = (model/tool cost + human review and cleanup hours × loaded hourly rate) / accepted tasks
There is no August AIHackers-owned CAR result for this stack. Run the same five to ten representative tasks with a fixed harness, model version, effort, budget, and acceptance rule. Record retries, token use, runtime, receipts, review time, cleanup time, and accepted count. Choose the cheapest lane that clears the rule, not the cheapest successful demo.
Recheck this guide when OpenAI changes plan credit rates, DeepSeek activates peak pricing or releases its harness, an evaluator refreshes post-cut costs, a provider changes a free route, GLM changes supported tools, or Kimi changes API or membership terms. Do not recompute an evaluator’s historical cost without labeling the arithmetic and formula as an AIHackers scenario.
Primary sources: OpenAI price update, Codex models, plans, credit rates, and usage, DeepSeek 0731 update, DeepSeek pricing, OpenCode Zen, GLM-5.2, Kimi API pricing, Kimi K3 weights, NVIDIA model catalog, NVIDIA Kimi K2.6, and NVIDIA API Trial Terms (exact PDF capture, 2026-07-16 19:06:06 UTC). The current Codex credit-rate revision is archive-pending; the latest replayable capture checked still contains older Terra and Luna rates.
Independent evidence: Artificial Analysis GPT-5.6, DeepSeek V4 Flash 0731, GLM-5.2, and Kimi K3 AA-Briefcase.
Related: LLM cost-saving playbook, GPT-5.6 family, Luna value-default analysis, perishable subscription capacity, DeepSeek 0731 value frontier, GLM-5.2, Kimi K3, Kimi access, benchmark mini-eval, and the app-hacking field test.