One packing-list batch passed; the two browser artifacts failed. This $1 showcase ran three independent tasks through DeepSeek’s direct API: a cat runner, a subscription calculator and twenty fictional messages transformed into orders and CSV. The accepted JSON and CSV are below. There is no playable game, working calculator or gameplay clip from these runs.
The three requests reported $0.0145671 through OpenCode. That is telemetry, not reconciled billing. $0.973212 remains reserved against the original candidate’s $1 ceiling while charges are unresolved. The recorded review subtotal is 1.81 agent minutes, covering independent inspection and the short browser check. Additional main-agent inspection and editorial/QA setup were not timed; human review time was not measured. API dollars and review minutes stay separate.
- API spend
- $0.014567 reported*
- Recorded review
- 1.81 agent minutes
- Accepted tasks
- 1/3 task results
01 / game
Cat runner
One HTML file: a fenced path, flowerpots, a fixed obstacle sequence and a cat that jumps, collides and restarts.
DeepSeek V4.1 Flash · direct API
failed-no-artifact
The response reached the 8,192-token ceiling with all tokens spent on reasoning and returned no HTML. No game or clip can be offered.
API: $0.005035 reported · unreconciledReview: 0.30 agent minutes1 initial attempt · 0 repairs
Prompt Source: none published Receipt / Run record
Acceptance checks and limits
- fail: Complete offline HTML and visible game — Zero source bytes; no HTML was returned.
- not-run: Keyboard, click and touch jumping — No artifact.
- not-run: Deterministic, clearable obstacles — No artifact.
- not-run: Distance scoring and collision — No artifact.
- not-run: Jump and landing squash/stretch — No artifact.
- not-run: Complete instant reset — No artifact.
- not-run: Narrow layout and keyboard restart — No artifact.
- One initial generation request; no repair reached the model.
- A repair request was refused before generation: retained reservations would exceed the $1 original-candidate lane.
- Gameplay, mobile controls and ten-second clip: not-run because no accepted game exists.
02 / calculator
Subscription calculator
Monthly and annual entries, cancellation choices and local saving, with arithmetic in USD cents.
DeepSeek V4.1 Flash · direct API
failed-incomplete-html
The response ended mid-expression inside an unclosed script block. Chromium displayed the static form but Add created no entry and the total stayed $0.00. The source is available for inspection; it is not offered as a working calculator.
API: $0.005044 reported · unreconciledReview: 0.31 agent minutes1 initial attempt · 0 repairs
Prompt Source Receipt / Run record
Acceptance checks and limits
- fail: Usable add/remove, periods and cancellation controls — Unfinished script is not executed; Add creates no entries.
- fail: $15/month plus $120/year equals $300/year — The calculator remains at $0.00 because entries cannot be added.
- not-run: Cancellation and removal update totals — No working entry handlers.
- not-run: Reload preserves entries and cancellation — No working save handlers.
- not-run: Invalid inputs preserve valid totals — Unfinished artifact cannot validate inputs.
- not-run: Unavailable storage remains usable — Unfinished artifact is already unusable.
- fail: Integer-cent arithmetic, offline operation and narrow keyboard use — Source is incomplete. Static inspection also found no safe aggregate bound for many large entries.
- Exact generated source retained without repairs or editing; downloaded as plain text.
- Chromium 148 offline file: no external requests; no console/page errors because the unclosed script is not executed.
- Static layout fits 1440, 390 and 320 pixels; this does not establish calculator usability.
- A repair request was refused before generation: retained reservations would exceed the $1 original-candidate lane.
- Persistence, cancellation, invalid-input and unavailable-storage behavior remain not-run.
03 / transformation
Twenty messages → packing list
One fictional batch, three products, amendments, aliases, cancellations and missing history. Orders, CSV and quantities must agree.
DeepSeek V4.1 Flash · direct API
accepted
One fictional batch passed the frozen contract: five orders, twelve CSV lines, all three product totals and coverage of all twenty messages. Missing history stays a clarification instead of an invented order.
API: $0.004488 reported · unreconciledReview: 1.20 agent minutes1 initial attempt · 0 repairs
| Product | To pack | Cancelled |
|---|---|---|
| Trail Notes Notebook | 5 | 1 |
| USB-C Cable, 2 m | 5 | 2 |
| Enamel Camp Mug | 4 | 7 |
Missing history: That earlier order is not part of the supplied message history. Please provide the order ID and its current lines and quantities so I can set it to 3 additional C-to-C leads (SKU-CBL-02) and remove the Trail Notes Notebook (SKU-NBK-01) line.
Download JSONDownload CSV Prompt Source Receipt / Run record
Acceptance checks and limits
- pass: Structured schema and catalog identities — Exactly seven required top-level fields, valid customer/product IDs and five sorted orders.
- pass: Amendments, aliases and cancellations — Final line quantities, statuses and audit source IDs match the frozen key.
- pass: Canonical CSV — Twelve rows; required order and quoting; no final newline.
- pass: Product quantities — Open/cancelled totals: notebooks 5/1, cables 5/2, mugs 4/7.
- pass: Missing-history clarification — M20 asks for order ID and current lines; no unsupported order is created.
- pass: Every message covered once — M01 through M20 in order, with exact dispositions, order IDs and product references.
- pass: Independent exact reconciliation and source identity — Offline validator passes; independent review matches raw text events to source bytes and reconciles the key.
- One accepted batch counts as one task; this is one candidate and one fictional dataset.
- JSON has no executable code; CSV is extracted verbatim from its returned csv field.
- No comparative model ranking or general order-processing accuracy is established.
- Human review minutes and full CAR remain unmeasured/not-run.
What the receipt establishes
The version 1 manifest links each result to its frozen prompt, source identity, attempts, effective request settings, token counters, checks and spend state. The shared generation profile denies all tools and adds a small visible AI-generated label to HTML. Codex authored the task pack, safety tooling, validators and this page. The downloaded outputs are the candidate’s bytes, without editor repairs. The CSV is extracted verbatim from the returned JSON’s csv field.
We froze identical route-neutral briefs, acceptance rules, fictional inputs and a private packing-list key before generation. Each task received one initial attempt. The packing result passed both the frozen-key validator and independent message-by-message reconciliation. One accepted batch counts as one task, not twelve CSV rows or twenty accepted messages. Existing cat projects were not attributed to the tested model.
The cat response spent all 8,192 generated tokens on reasoning and returned no HTML. The calculator used 4,707 reasoning tokens and 3,485 visible tokens, ending mid-expression in an unclosed script block. Chromium did not execute that block: Add created no entry and the total stayed $0.00. Cancellation, reload, invalid-input and unavailable-storage behavior therefore remain unverified. The packing response stopped normally after 4,511 reasoning and 2,429 visible tokens.
The original runner labelled both length-limited responses “completed.” Review corrected their public artifact outcomes and preserved that original label in the run records. Future runner classification is fixed. This is a disclosed tooling correction; no generated source was polished.
Why repairs stopped
Every task allows an initial attempt and at most two repairs. A repair may contain observed failures and that candidate’s previous output, with no other candidate’s code or hidden key. A failed output does not trigger a route change.
The ceiling is $1 per original candidate lane and $3 overall, including billed failures, reasoning and auxiliary generation calls. Before each request, the runner reserved $0.324404: the full documented input-context bound at the highest applicable input rate plus the capped output at peak pricing, rounded up. This deliberately conservative bound is much larger than the reported usage cost. The current provider pricing distinguishes peak and off-peak rates; our reservations use peak rates throughout.
All three reservations remain held. The cat and calculator repair requests were refused before model generation, because another worst-case reservation would exceed their original lane’s $1 ceiling. Unused ceilings in the other lanes cannot be transferred into this candidate’s lane. No deposit or automatic top-up was purchased or enabled by the experiment.
The local cap fixture verified an outgoing 8,192-token limit, reasoning accounting, and exactly one provider request under a forced HTTP 500; a second dispatch was rejected locally. Paid requests used OpenCode 1.18.33, a separate tool-disabled profile and omitted sampling/effort overrides. DeepSeek documents thinking enabled and high effort as defaults; those defaults are a provider statement, while the omitted fields and cap are recorded adapter observations. The native default output limit was overridden. These results measure this constrained run, not the model at its full output allowance.
A balance check confirmed sufficient existing DeepSeek allocation without publishing account values. We do not show remaining provider inference allowance: unresolved charges and an account balance are different from an experiment ceiling. Direct-price examples would be hypothetical. Full-CAR remains not-run; these three artifacts do not supply the full site-task evaluation, reconciled billing or a loaded human hourly rate.
The two access gaps
DeepSeek V4.1 Flash, MiMo V2.6 Flash and Kimi K3 were selected for compact browser code and structured text tasks. Only DeepSeek had verified existing allocation through the audited OpenCode credential store. DeepSeek’s API terms, effective April 29, permit downstream applications and assign outputs, with AI disclosure required. The route was deepseek/deepseek-flash; V4.1 Flash is the provider’s documented mapping for that API ID.
Direct MiMo and Kimi allocation was not verified. OpenCode Zen’s October 7 roster lists MiMo V2.6 Flash Free, then LongCat 2.5 Preview Free and Big Pickle as proposed substitutes. Free access is temporary. MiMo and Big Pickle data may be used for improvement; LongCat’s provider is described as using zero retention and no training. Big Pickle’s underlying identity is undisclosed. None was called in this experiment.
The OpenCode terms, effective August 15, contain internal-use and automated Output-extraction restrictions alongside an ownership clause. Permission for this automated public showcase remains unresolved, so those replacements stayed unrun. This is an access/permission gap, not a model-quality failure. NVIDIA’s trial terms, section 1.2, exclude production use of trial output; no NVIDIA-generated artifact appears here.
Browser and evidence limits
The failed calculator was inspected as an offline file in Chromium 148, with external requests blocked, at 1440, 390 and 320 pixels. Its static form fits those widths; that does not make it usable. There were no console errors because the unfinished script never executed. Keyboard/touch gameplay and the ten-second clip are not-run, because no game exists. Firefox, WebKit, physical devices and screen-reader operation were not tested.
Prompts, reviewed source downloads and run records remain readable without this page’s JavaScript. Accepted games, if a later version earns one, load on demand in sandboxed frames. Accepted calculators open separately for local saving. The source hashes prevent an edited artifact from being presented as the tested output. Raw prompts, filtered events, stderr and access/budget evidence remain private; public records omit credentials, account identifiers, private paths and reasoning text.
Use the cost-saving playbook to define acceptance and count retries and review, or AI Value for dated buying guidance. The OpenCode guide owns general harness and access information. This page owns these runs, their failures and the accepted fictional packing batch.