GPT-5.6 Luna is OpenAI’s value default for bounded, routine Codex work and high-volume API tasks—not a universal frontier replacement. On July 30, OpenAI cut Luna’s standard short-context API price by 80%, from $1 input / $6 output to $0.20 / $1.20 per million tokens. Terra fell 20% to $2/$12. Sol stayed at $5/$30.

That changes Luna’s shortlist position. It does not prove that cheaper tokens produce cheaper accepted work.

What changed on July 30

GPT-5.6 tierLaunch input / outputCurrent input / outputChange
Luna$1 / $6$0.20 / $1.2080% lower
Terra$2.50 / $15$2 / $1220% lower
Sol$5 / $30$5 / $30Unchanged

These are OpenAI’s standard short-context rates per million tokens. Luna cache reads cost $0.02 and cache writes $0.25. The complete API pricing table lists separate Batch, Flex, Fast, regional, and long-context rates.

The context footnote matters. Luna supports 1.05 million tokens, with 922,000 maximum input and 128,000 maximum output. Once a request exceeds 272K input tokens, OpenAI prices the full request at twice the short-context input rate and 1.5 times the output rate. For Luna, that is $0.40 input, $0.04 cached input, $0.50 cache writes, and $1.80 output.

OpenAI attributes the cut to serving-efficiency work. That is a vendor explanation; it is not independent evidence about Luna’s quality. The durable buying question is whether the same accepted result now costs less.

Subscription value is not API value

OpenAI’s current Codex pricing guidance gives Plus users approximately 250–2,000 Luna local messages per five hours, compared with 25–200 for Terra and 10–100 for Sol. Local messages and cloud chats share the five-hour window, and additional weekly limits may still apply. These are approximate plan allowances, not purchased tokens, API credit, or a promise that every task consumes the meter equally.

The direct API is a separate buying lane. Luna’s $0.20/$1.20 short-context rates make it the first OpenAI API route to test for high-volume extraction, transformation, routing, and bounded coding. A Plus allowance does not offset an API invoice, and cheap API tokens do not guarantee more accepted work after retries and review.

What the benchmarks actually say

No single board supports “Luna is frontier quality.” The current evidence supports a narrower claim: Luna can be close enough on some coding tests to deserve a workload-specific trial, while harder agentic work still separates the tiers.

Source, checked Aug. 1LunaTerraSolWhat it means
Artificial Analysis Intelligence Index at max effort515559Independent aggregate; effort, latency, and token use travel with the score
Artificial Analysis Coding Agent Index at max effort757780Independent harness result, not an AIHackers repository test
OpenAI-published SWE-bench Pro62.7%63.4%64.6%Vendor result on one named coding benchmark; Sonnet 5 reports 63.2% in its own launch material
Agent Arena, xHigh#18#17#4Live agentic sessions still favor Sol; ranks move as the board updates

BenchLM’s Luna profile adds a useful but limited signal: Luna was #25 of 214 overall and #6 of 129 in coding, with a $0.70 blended input/output price. The overall position is Estimated, not a fully supported rank, and the blended figure is a token-price average—not task cost.

There is also a live data-quality trap. Artificial Analysis’s July 9 article and model pages still show Luna’s old $1/$6 launch price, even though their benchmark scores remain useful. BenchLM has the new price, and OpenAI’s pricing table is authoritative for direct API billing. A recent crawl timestamp does not prove that every field reflects the latest event.

The CAR verdict: token savings are not accepted-work savings

AIHackers uses Cost per Accepted Result (CAR):

CAR = (model/tool cost + human review and cleanup hours × loaded hourly rate) ÷ accepted tasks

For a short-context request with 100K fresh input and 20K output, the model-token subtotal is about $0.044 on Luna, $0.44 on Terra, $1.10 on Sol, and $0.40 on Sonnet 5 at its introductory rate. Luna is 25 times cheaper than Sol on that fixed token mix.

That is not CAR. It excludes reasoning expansion, cache behavior, tools, retries, failed attempts, latency, and human review. AIHackers has not run an accepted-task Luna evaluation, so Luna’s site-owned CAR is not-run, not a number to infer from a leaderboard.

Use the mini-eval and CAR method on the same repository tasks before moving production traffic. The Smart Spend guide explains why list price, subscription allowance, and successful-task cost must stay separate.

Who should switch—and who should not

WorkloadStarting laneProduction gate
Routing, classification, extraction, and structured transformationLunaSchema validity, factual error rate, and retry budget
Bounded tool calls and background automationLunaTool success, recovery behavior, latency, and CAR
PR triage, first-pass review, and well-specified test implementationLuna, then Terra if neededAccepted patches on the same repo and harness
Ambiguous multi-file planning or architectureSol or a measured premium peerWhether the stronger tier changes the accepted result
Security-sensitive, destructive, or high-stakes final reviewStronger reviewed lane plus human approvalExplicit controls, evidence, and independent verification
Long-horizon autonomous codingDo not switch on price aloneSame-task completion, recovery, review time, and failure cost

The practical ladder is Luna first for bounded routine Codex work and high-volume OpenAI API tasks, Terra when Luna’s retries or review erase the advantage, and Sol for hard or consequential work where the stronger tier changes the outcome. The GPT-5.6 family guide keeps the specifications and current prices together; the Sonnet 5 guide is the relevant Claude comparison, while the Sonnet 5 launch analysis preserves its dated evidence.

The price cut makes Luna easier to justify as a first route. It does not remove the need to measure what reaches your acceptance bar.

Sources and evidence state

Archive note: exact CDX and save attempts were run August 1, 2026. Any source without a replayable capture remains archive-pending; a wildcard lookup URL is not treated as evidence.


Checked August 1, 2026. Prices, rankings, benchmark coverage, model limits, and provider terms can change independently.