GPT-5.6 Luna is OpenAI’s value default for bounded, routine Codex work and high-volume API tasks—not a universal frontier replacement. On July 30, OpenAI cut Luna’s standard short-context API price by 80%, from $1 input / $6 output to $0.20 / $1.20 per million tokens. Terra fell 20% to $2/$12. Sol stayed at $5/$30.
That changes Luna’s shortlist position. It does not prove that cheaper tokens produce cheaper accepted work.
What changed on July 30
| GPT-5.6 tier | Launch input / output | Current input / output | Change |
|---|---|---|---|
| Luna | $1 / $6 | $0.20 / $1.20 | 80% lower |
| Terra | $2.50 / $15 | $2 / $12 | 20% lower |
| Sol | $5 / $30 | $5 / $30 | Unchanged |
These are OpenAI’s standard short-context rates per million tokens. Luna cache reads cost $0.02 and cache writes $0.25. The complete API pricing table lists separate Batch, Flex, Fast, regional, and long-context rates.
The context footnote matters. Luna supports 1.05 million tokens, with 922,000 maximum input and 128,000 maximum output. Once a request exceeds 272K input tokens, OpenAI prices the full request at twice the short-context input rate and 1.5 times the output rate. For Luna, that is $0.40 input, $0.04 cached input, $0.50 cache writes, and $1.80 output.
OpenAI attributes the cut to serving-efficiency work. That is a vendor explanation; it is not independent evidence about Luna’s quality. The durable buying question is whether the same accepted result now costs less.
Subscription value is not API value
OpenAI’s current Codex pricing guidance gives Plus users approximately 250–2,000 Luna local messages per five hours, compared with 25–200 for Terra and 10–100 for Sol. Local messages and cloud chats share the five-hour window, and additional weekly limits may still apply. These are approximate plan allowances, not purchased tokens, API credit, or a promise that every task consumes the meter equally.
The direct API is a separate buying lane. Luna’s $0.20/$1.20 short-context rates make it the first OpenAI API route to test for high-volume extraction, transformation, routing, and bounded coding. A Plus allowance does not offset an API invoice, and cheap API tokens do not guarantee more accepted work after retries and review.
What the benchmarks actually say
No single board supports “Luna is frontier quality.” The current evidence supports a narrower claim: Luna can be close enough on some coding tests to deserve a workload-specific trial, while harder agentic work still separates the tiers.
| Source, checked Aug. 1 | Luna | Terra | Sol | What it means |
|---|---|---|---|---|
| Artificial Analysis Intelligence Index at max effort | 51 | 55 | 59 | Independent aggregate; effort, latency, and token use travel with the score |
| Artificial Analysis Coding Agent Index at max effort | 75 | 77 | 80 | Independent harness result, not an AIHackers repository test |
| OpenAI-published SWE-bench Pro | 62.7% | 63.4% | 64.6% | Vendor result on one named coding benchmark; Sonnet 5 reports 63.2% in its own launch material |
| Agent Arena, xHigh | #18 | #17 | #4 | Live agentic sessions still favor Sol; ranks move as the board updates |
BenchLM’s Luna profile adds a useful but limited signal: Luna was #25 of 214 overall and #6 of 129 in coding, with a $0.70 blended input/output price. The overall position is Estimated, not a fully supported rank, and the blended figure is a token-price average—not task cost.
There is also a live data-quality trap. Artificial Analysis’s July 9 article and model pages still show Luna’s old $1/$6 launch price, even though their benchmark scores remain useful. BenchLM has the new price, and OpenAI’s pricing table is authoritative for direct API billing. A recent crawl timestamp does not prove that every field reflects the latest event.
The CAR verdict: token savings are not accepted-work savings
AIHackers uses Cost per Accepted Result (CAR):
CAR = (model/tool cost + human review and cleanup hours × loaded hourly rate) ÷ accepted tasks
For a short-context request with 100K fresh input and 20K output, the model-token subtotal is about $0.044 on Luna, $0.44 on Terra, $1.10 on Sol, and $0.40 on Sonnet 5 at its introductory rate. Luna is 25 times cheaper than Sol on that fixed token mix.
That is not CAR. It excludes reasoning expansion, cache behavior, tools, retries, failed attempts, latency, and human review. AIHackers has not run an accepted-task Luna evaluation, so Luna’s site-owned CAR is not-run, not a number to infer from a leaderboard.
Use the mini-eval and CAR method on the same repository tasks before moving production traffic. The Smart Spend guide explains why list price, subscription allowance, and successful-task cost must stay separate.
Who should switch—and who should not
| Workload | Starting lane | Production gate |
|---|---|---|
| Routing, classification, extraction, and structured transformation | Luna | Schema validity, factual error rate, and retry budget |
| Bounded tool calls and background automation | Luna | Tool success, recovery behavior, latency, and CAR |
| PR triage, first-pass review, and well-specified test implementation | Luna, then Terra if needed | Accepted patches on the same repo and harness |
| Ambiguous multi-file planning or architecture | Sol or a measured premium peer | Whether the stronger tier changes the accepted result |
| Security-sensitive, destructive, or high-stakes final review | Stronger reviewed lane plus human approval | Explicit controls, evidence, and independent verification |
| Long-horizon autonomous coding | Do not switch on price alone | Same-task completion, recovery, review time, and failure cost |
The practical ladder is Luna first for bounded routine Codex work and high-volume OpenAI API tasks, Terra when Luna’s retries or review erase the advantage, and Sol for hard or consequential work where the stronger tier changes the outcome. The GPT-5.6 family guide keeps the specifications and current prices together; the Sonnet 5 guide is the relevant Claude comparison, while the Sonnet 5 launch analysis preserves its dated evidence.
The price cut makes Luna easier to justify as a first route. It does not remove the need to measure what reaches your acceptance bar.
Sources and evidence state
- OpenAI: July 30 price-cut announcement (Archive) — current prices and vendor efficiency explanation;
validated - OpenAI API docs: pricing (Archive) and Luna model specification (Archive) — short/long rates and limits;
validated - OpenAI Codex: plans and usage (Archive) — current Plus model ranges, shared five-hour window, and possible weekly limits; live fields
validated - Artificial Analysis: GPT-5.6 evaluation (Archive) — independent benchmark scores with pre-cut pricing still visible;
validated - Arena: Agent Arena — live Aug. 1 snapshot;
validated,archive-pending - BenchLM: GPT-5.6 Luna profile (Historical archive) — estimated aggregate, category ranks, current price, and source coverage; live fields
validated, current archivearchive-pending
Archive note: exact CDX and save attempts were run August 1, 2026. Any source without a replayable capture remains archive-pending; a wildcard lookup URL is not treated as evidence.
Checked August 1, 2026. Prices, rankings, benchmark coverage, model limits, and provider terms can change independently.