You can overpay for AI before sending a single prompt. Checked September 27, 2026: Apple’s Cambodia storefront lists Claude Max 20x at US$299.99/month. Anthropic advertises the same named tier at a US$200 web base price, while warning that mobile prices vary. That is a $99.99 listed-price gap, not a verified saving after tax and payment fees. Anthropic’s Max pricing.

Check your own bill before changing renewal settings.

Then come the costs that are harder to see: an API agent session that loses its cache discount, a reset that moves your next refill, and a cheap model that needs an hour of cleanup. These are different problems. Fix the one affecting your bill or your work first.

This is AIHackers’ main guide to reducing subscription, coding-agent and API costs. Keep the result useful, then reduce what it takes to get there. The priority below is editorial judgment about likely impact and effort, not a measured savings ranking.

Start with your bill · Understand cold caches · Time a reset · Copy the monthly audit

Where to look first

PriorityWhen it mattersFirst actionWhat improves
1. Billing route and duplicate seatsYou pay for subscriptionsMatch final web/app renewal totals; review unused seatsRecurring cash spend
2. Context and cache reuseAPI sessions with cache billingInspect cache reads/writes before changing the sessionMetered input cost and latency
3. Reset timing and plan limitsA usage cap blocks useful workCheck both reset clocks, expiry and the meter you exhaustedUsable capacity before your deadline
4. Instructions, model choice and retriesAgents repeat work or need cleanupRemove obsolete instructions; test one routing changeModel work and human time
5. Offers and an alternative providerRenewal or a material product changeVerify eligibility and the cost of switchingFuture spend and continuity

An expiring reset or blocked deadline can move to the top. For a high-volume API service, caching or routing may dwarf subscription savings. Skip checks that do not apply to you.

1. Check your bill before your prompts

The opening gap is nearly 50% above the web base; reversing $299.99 to $200 would be 33.3% below the mobile listing, before tax or fees. Those percentages use different denominators. The US Apple listing instead shows Max 20x at $249.99: the gap varies by market and plan.

Compare the same tier, billing period, legitimate billing location, taxes, currency conversion and renewal benefits. Do not change country details to chase a listing. Use the Claude price-check checklist, including its observed Thai checkout example and steps to avoid overlapping subscriptions. A store listing is neither your checkout quote nor a settled charge.

For every provider, record who bills you and what renews next. If a second seat does no useful work, assess cancellation before optimizing pennies per request. Confirm any loss of access, data or grandfathered terms first. Already-paid allowance can avoid another charge; using it merely to empty a meter creates no saving.

2. The conversation you paid for can become expensive again

A saved conversation and a live prompt cache are different things. The conversation stores history; a cache can reuse computation for an unchanged beginning of a request. Anthropic documents a default five-minute cache lifetime, refreshed by reuse, and an optional one-hour duration at a higher write price. After expiry, retained history can need processing again. Claude cache mechanics.

Here is the scale of the difference in Anthropic’s published API example, checked September 27: on Fable 5.1, a 100,000-token prefix costs $1.25 for a five-minute cache write versus $0.025 for a cache read (about $0.03 rounded). The exact rates imply 50× for that prefix; dividing the rounded $0.03 does not. This excludes output, new input and tools, and is not a 50× whole-session bill or a subscription-quota formula. Vendor example · Exact model rates.

What triggers it? An expired entry, a different model, or changes ahead of the reusable prefix can cause a miss. Editing instructions or tool definitions and changing top-level effort/thinking settings can alter that prefix. The scope depends on model and API; it is not true that every setting change always invalidates everything. Anthropic documents supported per-message changes, and OpenAI documents a GPT-6 configuration-update route that preserves earlier context. OpenAI cache guidance.

What to do: keep a useful active thread intact. At a task boundary, compare continuing it with a compact handoff containing the objective, current revision, relevant files, decisions, failed attempts and next check. Shortening context can lose evidence and itself require a cache write; do not compact reflexively. Switch model or effort when the quality gain warrants it. Check actual usage fields afterward. The caching guide separates API billing, context occupancy and subscription limits.

3. Check when the weekly clock starts

Weekly windows do not accumulate like a credit balance: two idle weeks are not a budgetable 2× allowance. Check whether your clock is already running or starts with use; do not assume an old refill is still one day away. For a usage-triggered window, a small useful task starts the clock before a larger session later. The weekly-clock guide explains the scheduling value, with OpenAI’s paid-reset rule as a documented example. Confirm the signed-in meter instead of assuming rollover.

OpenAI’s current banked-reset instructions say a full reset refreshes both Codex windows and changes the weekly reset date. You do not also get the old scheduled weekly refill. Availability and expiry depend on the account and offer.

Illustrative timeline: suppose your account displays a natural refill on Friday and an available banked reset expiring the following Monday, and today is Tuesday.

ChoiceWhat followsWhen it makes sense
Redeem and resume TuesdayNext weekly refill moves to around the following Tuesday; no refill this FridayValuable work cannot wait
Wait for Friday’s natural refillYou retain the banked reset until its expiry; redeem later only if neededRemaining capacity or another route covers the wait
Wait past Monday’s expiryThe unused banked reset is lostThere was no useful workload needing it

Waiting can preserve access to the imminent natural refill. It does not guarantee double useful compute: demand, remaining allowance, eligibility and expiry decide the benefit. Read the updated account clock after redemption.

Purchased weekly resets are another product: OpenAI says they pull the next allowance forward, with the next automatic weekly reset seven days after the first subsequent request. They are not top-up credits. Use the capacity guide’s current reset rules for the distinction and the chronology for historical reports.

4. Ask “20x of which limit?”

Anthropic defines Max 5x and 20x against Pro’s per-session allowance, with five-hour resets and a separate weekly cap. That does not establish a 20× weekly allowance. Official Max limits.

OpenAI’s Pro tier FAQ describes 5×/20× usage versus Plus, while also noting model-specific allowances. That wording alone cannot prove an exact multiplier for every weekly meter. Check current purchase and return eligibility before cancelling a plan.

Before upgrading, identify the exhausted meter: session, week, model, tool or workspace. Check whether the upgrade changes that meter, whether products share it, and what happens at the cap. The Claude limits ledger owns the detailed evidence. One frustrated account report cannot establish everybody’s effective capacity.

5. Maintain instructions like dependencies

Old AGENTS.md, CLAUDE.md, skills and task prompts can make an agent reread irrelevant documents or repeat checks. OpenAI’s Astra guidance recommends revisiting inherited instructions and loading documents when relevant to the task.

Audit duplicated rules, retired model names, obsolete commands and mandatory reads. Keep security boundaries, acceptance criteria and required release checks. Move task-specific detail behind a short pointer. Compare a representative task before and after; shorter instructions are only better if they preserve correct behavior.

6. Stop paying twice for the same failure

Use ordinary code for known rules: sorting, arithmetic, parsing and validation. For work needing judgment, try the least expensive candidate that passes your checks. Reserve stronger reasoning for ambiguity or failures it can actually resolve. More agents add coordination and duplicated context; compare their total cost with a single-agent baseline.

After failure, name the cause before paying for another attempt. Missing evidence needs retrieval; a broken tool needs repair; unclear scope needs clarification; a capability gap may need another model. A transient error may justify a bounded retry. Set a maximum number of attempts and a deadline. The three workflows below show concrete stopping rules.

7. Check offers without turning model shopping into a job

A qualifying promotion can matter more than a small token discount, but enrollment, region, expiry, cashback delay and ordinary renewal determine its value. Use the Deals ledger; unresolved eligibility is not an available saving.

Keep one primary route and a tested fallback if continuity matters. Reconsider them after a material price, limit or capability change. Include migration, privacy requirements, integration work and learning time before switching. A second recurring subscription is not required merely to have an alternative.

The ten-minute cost audit

Start with your renewal totals if you only need a subscription decision. For a repeated workflow, use this more detailed comparison:

  1. Pick one repeated workflow and a representative set of tasks. Write its acceptance rule, deadline, and retry budget before choosing a model.
  2. Record the access route, submitted tasks, attempts, accepted results, model/tool spend, and review and cleanup minutes. Mark missing evidence as unknown.
  3. Identify deterministic operations and irrelevant context you can remove.
  4. Choose one intervention from the stack below. Rerun the same submitted tasks under the same acceptance rule.
  5. Compare accepted-result cost and completion rate. Keep the change only if it meets your quality and timing needs.

Keep separate spending ledgers

LedgerWhat it buysWhat it is not
Included subscription allowanceAccount-specific use inside a product and its supported surfacesCash, portable tokens, or general API credit
Purchased product credits or usage bundlesAdditional use under that product’s rulesA refill of included allowance unless the provider says so
API billingMetered model, tool, cache, service-tier, and data-processing useSubscription capacity or evidence that the output was accepted
Promotional accessTemporary or eligibility-limited useA durable price or entitlement
Local inferenceCompute on hardware you controlZero cost; hardware, power, setup, maintenance, and review remain
Human and retry costTime spent rerunning, checking, correcting, and integratingUsually visible on a provider receipt

Track cash and capacity separately. Already-paid allowance may avoid an additional API charge, but unused allowance is not money saved. Check the active billing route before starting: subscription access and API-key authentication can lead to different bills. The capacity guide owns product-specific allowance rules.

The savings stack, in order

  1. Define acceptance first. Specify outputs, checks, source requirements, latency, rejection conditions, and a stopping budget. Keep this contract fixed during a comparison.
  2. Remove deterministic work. Parse, filter, sort, diff, calculate, and validate with ordinary code when the rule is known. Send only work that still needs model judgment.
  3. Choose access. Compare already-paid capacity, metered API use, purchased credits, and local inference. Confirm permitted use and account access; do not treat a listing or promotion as an entitlement.
  4. Bound context. Search filenames, symbols, or source sections first. Retrieve the relevant ranges and cap noisy tool output. Preserve the evidence needed to verify the answer.
  5. Route by task. Start with the least expensive candidate that has passed representative acceptance checks. Escalate unresolved ambiguity or rejected work within the agreed budget.
  6. Cache and batch where suitable. Keep reusable instructions, schemas, and examples stable; measure actual cache writes and hits. Batch independent deferred tasks only when their completion window fits. A discount cannot rescue a missed deadline.
  7. Check timing and offers. Verify reset state, off-peak rules, eligibility, expiry, and the final charge at execution time. Compare the ordinary renewal terms as well as the temporary benefit.
  8. Compare accepted-result cost. Include all attempts and review time across equal submitted task coverage. Stop testing when the evidence supports the decision or the budget is exhausted.

OpenAI’s cache documentation and Anthropic’s cache documentation describe provider-specific eligibility and billing. Stable text alone does not guarantee a hit. OpenAI’s Batch guide illustrates deferred processing; check the chosen provider’s current window and supported workload. The caching guide owns implementation details.

Three worked workflows

1. Interactive coding

  • Input: a bounded bug or feature, the relevant repository revision, and named files. Search before loading code; merge related coding work into one scoped task.
  • Acceptance rule: required behavior passes its tests, the diff stays in scope, and review finds no unresolved correctness issue.
  • Cost intervention: let tools handle search, formatting, and checks. Feed the agent relevant excerpts and errors, keep useful instructions stable, and escalate only when checks or ambiguity justify it.
  • Stopping condition: stop when checks and diff review pass. If the retry or time budget runs out, record a rejected result and the remaining blocker instead of continuing an open-ended loop.

2. Bulk structured extraction

  • Input: a fixed set of source records, a JSON schema, and representative examples.
  • Acceptance rule: each accepted row passes schema validation and agrees with its source; missing or ambiguous fields follow the declared rejection rule.
  • Cost intervention: parse and normalize deterministically first. Send only semantic or rejected rows to the model, reuse eligible prefixes, and batch independent records when the deadline allows it.
  • Stopping condition: finish when every submitted record is accepted or explicitly rejected. Retry only rejected rows within the budget; include their costs and report coverage against the original set.

3. Source-backed research

  • Input: one decision question, a date cutoff, named primary-source candidates, and a bounded research budget.
  • Acceptance rule: every consequential claim has a supporting source passage and date; conflicting or unavailable evidence is labeled unresolved. A link alone does not count as support.
  • Cost intervention: search and fetch targeted sections, deduplicate sources, and use the model to compare evidence and draft the answer. Reserve stronger reasoning for contradictions or consequential judgment, then manually check the claim-bearing passages.
  • Stopping condition: stop when the question is supported to the declared standard or the source/time budget is exhausted. Report the evidence gap; do not replace missing evidence with repeated model guesses.

Measure Cost per Accepted Result

CAR = (model/tool cost + human review and cleanup hours × loaded hourly rate) / accepted tasks

Use equal submitted task coverage: both lanes receive the same task set, acceptance rules, deadline, and retry policy. Count the cost of every attempt, including failures and retries, and all review and cleanup time. Report accepted/submitted counts alongside CAR so a lane cannot look efficient by quietly skipping hard tasks. Count each accepted task once.

Allocate subscription cost over a declared period using a consistent share of measured use attributable to this workflow. Report the full seat bill and unused capacity separately; do not charge a whole seat to one tiny experiment or value included allowance as portable API dollars. For local inference, allocate hardware depreciation or rental, power, setup, and maintenance over the same period and attributable use. Include human setup time once, without double-counting it as both tool cost and review. Show the allocation assumptions and how a different utilization level changes the conclusion.

Cash and time are separate outputs. Record the actual bill and review minutes before assigning a value to time. Without a defensible hourly rate, report cash and review minutes separately; combined CAR remains undefined. A lower combined CAR may mean time saved rather than money returned to your account. For a renewal decision, also compare the full future bill; an already-paid seat is a sunk cost only for the current period.

Hypothetical comparison—not an AIHackers result: both lanes receive the same four submitted tasks. Lane A spends $0.60 across its attempts, accepts two of four tasks, and needs 72 minutes of review and cleanup at $60/hour. Its CAR is ($0.60 + 1.2 × $60) / 2 = $36.30. Lane B spends $4.00, accepts all four tasks, and needs 24 minutes at the same rate. Its CAR is ($4.00 + 0.4 × $60) / 4 = $7.00. These invented model/tool totals include all attempts; the example has no subscription or local costs. The higher model bill buys a lower accepted-result cost under these assumptions.

Zero accepted results means no finite CAR. Report the spend, review time, and 0/N acceptance; mark CAR undefined. If the evaluation or required evidence is missing, use not-run or incomplete rather than inventing a number.

CAR is not API token price, subscription quota consumption, benchmark cost per task, or a provider’s self-reported session cost. Small samples, subjective acceptance, differing task difficulty, and cost allocation limit comparisons. Show latency and failure notes too. The benchmark-literacy guide helps turn a shortlist into a fair evaluation.

DeepSeek is still not-run

AIHackers’ DeepSeek CAR remains not-run. The dated connectivity preflight establishes tooling access, not accepted-result performance. No savings finding follows without task decisions, receipts, and review time.

Choose your next step

Current buying guidance

The September 27 value shortlist applies this method to current working tiers and model routes. Use AI Value for the current published guide and Free Stack for no-payment and trial boundaries. For a specific route, check its owner: OpenCode for harness and billing choices, Z.AI for supported-tool plan rules, and the local-inference guide for hardware tradeoffs. Compare those candidates with the method above before paying.

Evaluate another low-cost model

Use the dated MiMo guide, DeepSeek value analysis, GPT-6 Luna pricing and budget comparison to find candidates. Their own review dates and access conditions apply. Choose two accessible routes, hold the task and acceptance rule fixed, then count retries and review time. A shortlist is not an AIHackers finding that one model is cheapest for your work.

For premium work, the Opus 5.5 review examines capability, output efficiency and the limits of launch enthusiasm. Compare it with GPT-6 Sol and Astra on your acceptance criteria. Keep launch credits and reset offers separate from the recurring cost of the workflow.

Before changing a Codex plan, check the reset state

Check the account’s actual meters and renewal terms. The reset chronology retains account-specific plan-change reports; they establish neither a universal defect nor a workaround.

Check where your Claude subscription is billed

Use the bill checklist for matched totals and renewal steps.

Claude Code: reduce session context overhead

The command guide covers context inspection, clearing and compaction. Measure accepted work after changing your workflow.

Timing and offers

The Deals ledger owns offer status and eligibility, including ChatGPT referrals and the Kimi draw. Verify the offer and ordinary renewal terms when you act. Promotions and referrals never change AIHackers’ rankings, recommendations, benchmarks, or CAR. Expired offers are history, not usable savings options.

Five-minute monthly audit

Bookmark this section for your next renewal. Copy the checklist into a note; change one thing at a time so you can tell whether it helped.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
Review date / next renewal:
Actual subscription and API bills (one currency):
Unused or overlapping seats to review:
Matched web/app renewal totals, including tax and fees:
Which limit blocked useful work:
Natural reset time / banked-reset expiry / work due before each:
Weekly clock already running or starts with use? Evidence, not assumed rollover:
Long sessions resumed or model/effort/tools changed:
Cache reads/writes observed (unknown if unavailable):
Repeated failures and their causes:
Obsolete agent instructions to test removing:
Accepted/submitted tasks / review minutes / total attempts:
One change to test / spending and retry cap / keep-or-revert rule:
Verified offer or fallback worth checking:
Next review date:

Review record: September 27, 2026. This revision checks the billing example, cache mechanics, reset rules, plan wording and instruction guidance against the linked primary sources. Dollar examples are listings or vendor calculations, not AIHackers receipts or experiments. Our priority order and remedies are editorial analysis; site-owned DeepSeek CAR remains not-run.

Maintenance: scheduled monthly editorial review, plus rechecks when a consequential source changes. Next planned review: October 2026. A reminder is not a completed fact-check; use the displayed review date and recheck account-specific prices and clocks before acting. Supporting articles keep their own review dates. Report a correction.

Frequently asked questions

What is the best first step for reducing LLM costs?

Write the acceptance test before choosing a model. Cost per accepted result includes retries and human cleanup, so the cheapest token route can be the more expensive workflow.

Are subscription limits the same as API credits?

No. Included subscription allowance, purchased product credits, promotional access, and metered API billing are separate ledgers with different rules.

When should an LLM task use Batch processing?

Use Batch for asynchronous, independent work that can tolerate the provider’s completion window and that has a machine-checkable output contract.

Does AIHackers have a DeepSeek cost-per-accepted-result finding?

No. Connectivity and tooling preflight are not accepted-result evidence. AIHackers’ DeepSeek CAR remains not-run.