Evaluate GPT-6 Astra Low first for planning and orchestration, then raise effort when acceptance checks fail. Keep Luna for well-defined implementation and retain Sol where your existing results justify it. This is a starting hypothesis for your workload, not a measured AIHackers winner.

The September value guide puts that recommendation beside DeepSeek and subscription choices. Our August router remains historical: a new planning candidate does not invalidate every successful Luna or Sol workflow.

What changed

OpenAI’s model documentation positions Astra for difficult end-to-end reasoning, coding, research, computer use, and document work. It accepts text and images, produces text, and exposes Low through Max API reasoning effort. More capable planning is useful when requirements cross files, tools, or stages; it still needs a concrete definition of done.

On September 6, Tibo Sottiaux said Astra Low performed better than Sol High and suggested Low or Medium to people satisfied with Sol High. That is a broader performance comparison, not a planning-only benchmark. The original post was inaccessible during our check; the September 7 forum reproduction preserves the attribution. The forum author’s orchestration interpretation is community commentary, not official documentation or an independent evaluation.

Our narrower recommendation is to test that starting point on planning. Give Astra the constraints, relevant files, permitted actions, and acceptance checks. Increase effort only after identifying what failed. If the plan is already explicit, handing bounded implementation to Luna can avoid paying for another round of open-ended analysis.

Access and pricing

Standard API rates are $10 input / $50 output per million tokens, with $1 cache reads and $12.50 cache writes. Above 272K input tokens, the model page specifies double input/cache rates and 1.5× output rates for the full request. API billing is separate from included Codex use.

Codex pricing documents Astra access and workload-dependent allowance estimates. Existing Pro 20x is our preferred premium Codex lane when sustained utilization justifies $200/month. New subscriptions and upgrades are unavailable during the current pause; see the verified access note. Do not cancel with an assumption that immediate repurchase is possible.

A shorter answer can save output tokens without saving total money. Input, cached context, reasoning, tools, retries, and the chosen billing route all matter. Measure elapsed time separately: fewer tokens do not guarantee lower latency if a request waits or runs more tools. Subscription consumption is another ledger; API price cannot predict how many projects your seat will finish.

Independent evidence and regressions

Artificial Analysis’s September 9 evaluation, using Intelligence Index v4.3, reports Astra Max at 53, six points above Sol Max. Every Astra effort is on its intelligence-versus-cost frontier; Low costs $0.82 per index task and Max $3.26. Those are evaluator costs, not accepted-result costs for your work.

The same report finds regressions: Astra trails Sol by about 45 Elo on GDPval-AA v2 and loses presentation quality on AA-Briefcase. In the Coding Agent Index, Codex/Astra scores 62 overall but DeepSWE falls to 68% from Sol’s 72%. Strong aggregate results therefore support evaluation without establishing universal superiority.

Do not compare these numbers directly with August’s v4.1 scores. The v4.3 methodology update changes the evaluation mix. Vendor benchmarks, independent indices, and site-owned tests remain separate evidence classes.

Limits and a practical decision

Dated community experience is mixed. In a September 5 discussion, the author described clever problem-solving alongside scope drift and unnecessary implementation. Replies also report solving problems Sol missed while struggling with overengineering. These self-selected observations do not establish prevalence, a universal effort setting, or a reproducible regression.

Keep the experiment bounded: compare the same task, review the plan before consequential execution, and record correction time. Accept Astra only when its result passes your checks at an acceptable combined cost. Keep Sol if presentation or collaboration is better on your tasks. AIHackers has run no paid Astra comparison or Cost per Accepted Result study: CAR is not-run.

Sources above were checked September 14, 2026. Archive status is archive-pending unless a replay has been validated. Continue with the cost-saving playbook and DeepSeek V4.1 Flash evaluation guide.