# Independent fictional-inbox review

Reviewer: fresh independent Codex reviewer; not the corpus, policy, runner, or draft author. Review scope: fictional fixtures, protocol, transport and spend controls, fixed-denominator scoring, raw result traceability, and draft-only reader claims. No inference, authenticated discovery, credential access, account-balance inspection, commits, merge, or production work was performed by this reviewer.

## Dataset verdict before final evaluation

At 2026-10-08T12:10:15Z, all 200 message/reference pairs had been read against the fictional `policy.md` without opening model or baseline result artifacts. Dataset review **passes after two policy-grounded corrections**. Human label validation and real-inbox acceptance remain **not performed**.

- `C007`: the requested access revocation was explicitly already completed; only an optional new quote remained current. The original `review/routine` reference changed to `sales/routine`. Message text did not change.
- `H004`: the original billing-dashboard loading failure overlapped billing workflow ownership and technical app-fault ownership. The message now explicitly states paid subscription and working sign-in; the loading fault is `technical/degraded`, replacing `billing/degraded`.
- Both findings were sent before any heldout/challenge result inspection. Original development corpus and proposed protocol are preserved under `history/`. No dev reference labels changed. No label correction was derived from model performance.
- The original proposal required review and freeze before every scored call, but one smoke and 40 dev calls had already occurred. The revised protocol explicitly records this unmet development gate and requires review and freeze before final heldout/challenge and stability calls. The earlier gate is not claimed as passed.

Structural checks passed: 200 unique IDs and distinct messages; dev40/heldout120/challenge40; 179 groups containing 21 two-case counterfactual groups (2 dev, 9 heldout, 10 challenge), with every group wholly inside one split; all controlled scenarios and enum values valid; every priority matches the fixed impact mapping. The 12 stability IDs exist, are unique, and appear in the same preselected sequence in both diagnostic rounds. Corpus arrival order owns ties; the probe reverses both question and option order and cannot isolate their effects.

`buildRequest` was inspected directly: API state contains only `{message}`. The fixed company policy and Choice rubric are question instructions. IDs, groups, split, scenario, expected label, annotation, neighboring messages, and result notes are absent from model input. The rules baseline contains text rules, not ID/reference lookups. This inspection does not independently prove authoring chronology; the retained author record and run manifests provide that provenance.

Dataset and instruction hashes reviewed at this checkpoint:

| File | SHA-256 |
| --- | --- |
| `policy.md` | `f817cc6325660dc1ef8a7dd3215360bb12607f7e87dd7c9688ab9c8cfcc24b3e` |
| `corpus.json` | `c9febfb493c7274dc6a43460cb31dd7981de6ceaf7ef0d8dcefb576e85fac39e` |
| `protocol.json` | `2d0c6f9c728cfb08c09e5e572344eef6a33a3bb69f090ff55b0ac9d060bb5ab3` |
| `questions.json` | `23f69c597c2a42f2b67b8ff17e62b507587a895b949852e7a8b18a7afbad0f9a` |

## Runner and scoring checkpoint

The offline stub suite passed 15/15 tests and 79 assertions. It tested exclusive locking, serialization of library reservations, cumulative $1 admission across ledger sessions, retention of failed holds, subtraction of outside reservations, private-file evidence checks, redacted HTTP failures, strict response validation/fallback, no network in offline mode, and stop-on-first-failure without retries. Temporary test directories were removed by test teardown.

Direct source review confirms one fixed ledger path per experiment checkout across CLI invocations/output directories/variants, a private exclusive lock, a durable $0.01 reservation before transport, a hard cumulative $1 cap, and conservative retained holds for failed or unknown requests. Successful validated usage alone settles a hold. Raw HTTP uses the fixed host, refuses redirects, has a bounded timeout, and makes no automatic retry. Account funds and keys do not enter manifests or requests; allowlisted headers exclude account-balance headers. Library request validation also enforces the message-only schema. These controls assume the operator preserves the ledger and does not create a separate checkout to reset its scope.

Independent primary-source recheck on October 8 supports the exact `jev-1.13.0` pin and published input price of $0.042 per million, free output, accepted versioned IDs, and 64k combined request budget with 32k state-plus-longest-question limit. [Models](https://docs.typesafe.ai/models), [HTTP contract](https://docs.typesafe.ai/api). Promotional issuance is discretionary; credit is consumed before purchased credit and can have additional expiry/revocation terms. [Terms, section 8.2](https://typesafe.ai/legal/mca). Calculated usage cost is not a payment receipt, verified balance, current new-user offer, or CAR.

Scoring approval **passes after source corrections**. Findings sent to integration owner and resolved before final freeze:

1. Missing/pending requested cases must remain errors in fixed report-set denominators; completed-row denominator alone is insufficient.
2. Invalid/failed predictions must rank below all valid priorities, including valid `review` predictions.
3. Report top20 precision, per-field confidence exclusions, counterfactual both-case exact correctness, named heldout/challenge/combined sets, valid-pair disagreements and invalid-pair counts, and repeat-original versus reversed comparison.
4. Keep earlier development artifacts and source identities distinct from final frozen scoring; no pool of different prompts, lanes, or repeated observations.

The revised scorer uses the entire intended case set for fixed denominators, counts missing/invalid cases as errors, and puts invalids below every valid priority. Named sets keep heldout120, challenge40, and descriptive combined160 distinct. Top20 precision, separate confidence exclusions, both-case exact correctness, fixed stability IDs, valid-pair disagreement, invalid/missing pairs, and repeat-versus-reversed comparisons are present. Successful-only confidence bins remain separate from fixed-denominator accuracy. The baseline reports no invented confidence and computes any missing deterministic rows under the frozen rules, with those IDs disclosed.

The revised offline suite passed **17/17 tests, 91 assertions**. Separate reviewer-authored in-memory probes passed for missing-case denominators, failure ordering after valid review, confidence exclusions, fixed stability pair denominators, and named report sets. These probes made no API requests and wrote no result fixtures.

| Reviewed source | SHA-256 |
| --- | --- |
| `jev.ts` | `2d18c96038a301ede3b4923aac72a1b447399bc2623bae5f5bc3e38002a7c4a8` |
| `evaluate.ts` | `3cadd100910ae30a56987439a50ba8d29a1be6f3c63cba4ca34d0bca121b03f6` |
| `rules.ts` | `5adb4c2e8f6f553dc489aee0f7408d0595c8c49a9aeae9f041fa07ef7018e5d3` |
| `run.ts` | `61622b393f8d44834ad66897de4fa8d10d71f123c9201594217d08bed68be20a` |
| `report.ts` | `e815e7e9906b593c4523cee014f33d550fe8f696a538efe06ae141be9a6531c4` |
| `runner.test.ts` | `62a2f7b010ce3f6b00a375d69f8ed7120f3d927387451473fde8943e98ee4fc1` |

After dataset approval, the reviewer opened all **41** development raw attempts: one D001 smoke and dev40. Stored request/response hashes, independently normalized answers, message-only inputs, and header allowlist match their rows. No artifact mismatch was found. Smoke exact match was 1/1; dev Jev was 39/40 and original rules 15/40; final rules development fit was 40/40. The smoke used an earlier corpus/protocol snapshot and remains separate. Reported usage calculates smoke cost $0.000102102 plus dev40 cost $0.004082274; these are calculated token costs, not receipts. No heldout/challenge output was opened during dataset or pre-freeze scoring review.

## Final frozen-result evidence

`freeze.json` records 2026-10-08T12:15:44.189Z. Every one of its 11 file hashes matches current bytes; heldout, challenge, repeat, and reversed manifests match the frozen fixture, question, and source hashes and were created after that freeze. `types.ts` is `0ad95e8d6580130dd043ec662f584b8699db0eaea4314aef49fd4a8d48570997`; its final scorer version is `inbox-scoring-1.1`. Policy, labels, questions, rules, and scoring were not revised from final outputs.

All **225** retained raw attempts were read and independently checked against their normalized rows: exact input text and question/criteria order, request and stored response hashes, raw/parsed JSON equality, pinned model, strict schema, derived priority, HTTP 200, zero retries, reported usage, calculated cost, and header allowlist. No mismatch was found. All 243 original-to-public copies (225 attempt artifacts plus six sets of manifest/results/report files) match byte for byte. The original dev-v1 reports retain their historical 0.75 confidence edge; final metrics use the frozen 0.8 edge and do not combine the legacy scorer as a final score.

The copied ledger contains 225 reservations and 225 settlements, 450 questions, no unresolved hold, and $0.022973202 calculated cumulative cost. Attempt IDs cover exactly the retained attempts. The 546,981 input and 21,656 output tokens independently sum to the reported accounting; input usage multiplied by the published rate reproduces the total. No account balance or provider payment receipt was examined by this reviewer. Cash purchases initiated by the task are reported separately as $0; actual promotional debit remains unreconciled. Build-agent compute, hosting, and human effort are excluded, so this is not a total project-cost or CAR result.

Independent counting reproduced:

| Set | Jev exact | Rules exact | Jev queue / impact / priority | Jev serious misses | Rules serious misses | Jev unnecessary urgent | Rules unnecessary urgent |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Heldout primary | 115/120 | 40/120 | 117/120; 118/120; 118/120 | 0/34 | 18/34 | 0/86 | 1/86 |
| Challenge diagnostic | 39/40 | 12/40 | 39/40; 40/40; 40/40 | 0/16 | 8/16 | 0/24 | 3/24 |
| Combined descriptive | 154/160 | 52/160 | 156/160; 158/160; 158/160 | 0/50 | 26/50 | 0/110 | 4/110 |

All 160 primary final responses are format-valid; malformed responses, HTTP/network failures, refusals, missing attempts, and confidence exclusions are zero in this observed run. Zero errors in these categories is an observation, not a service reliability guarantee. The six semantic mismatches are `H003`, `H043`, `H068`, `H090`, `H102`, and `C007`; they remain scored errors, with their raw answers and author references retained. No label adjustment was made from them.

Heldout top20 captures 20/34 serious cases for Jev versus 16/34 for rules; precision is 20/20 versus 16/20. Challenge top20 captures 16/16 versus 10/16; precision is 16/20 versus 10/20. Combined top20 captures 20/50 versus 17/50; precision is 20/20 versus 17/20. Exact selected IDs and arrival ties match an independent sort. Arrival is authored rather than sampled, so these are descriptive review-budget measures.

Each fixed stability comparison has 12/12 attempted and valid pairs, zero invalid or missing pairs, and 0/12 classification disagreements: main versus repeat-original, main versus order-reversed, and repeat-original versus order-reversed. This does not establish universal determinism or isolate question-order from choice-order effects. The primary accuracy remains the first frozen observation; repeats do not replace it.

Full frozen-scorer replay matches every saved metric and case detail for both systems in heldout, challenge, and combined sets. Confidence bins/ECE and nearest-rank latency percentiles were also independently recalculated. Heldout queue/impact ECE is 0.04175/0.02200; challenge is 0.06000/0.03700; combined is 0.0426875/0.02575. All bins retain their sample sizes and separate targets; no confidence threshold was selected. Heldout client p50/p95 is 285.083/337.585 ms (n=120); challenge is 279.798/320.586 ms (n=40); combined is 284.233/327.236 ms (n=160). These timings include transport/body read/normalization and exclude admission/disk writes.

Evidence identity: `freeze.json` SHA-256 `5df3d28730a86a7d15b5504cded5c473984a316eb161dfed8b360a3e8041b09a`; `summary.json` SHA-256 `3a66e7b7c43658510bb2f631b0332ab092c5706b9b2d4bd73b003cc74bf7785e`. Overlay primary160/repeat12/reversed12 arrays, summary, and freeze match the source package. The replay has 224 selectable observations (200 first observations plus 24 diagnostics); all 225 raw responses are retained, including the separate smoke. Development baseline rows are explicitly recomputed under frozen 1.1; original 1.0 rows remain archived. All 160 recorded final baseline classifications match direct execution of frozen `classifyRules`.

Recorded-result and frozen-source review **passes**. This review does not establish real-inbox generalization, human operational acceptance, production readiness, accepted-task economics, or publication approval beyond the requested draft scope.

## Final reader-source verdict

Reader-source review **passes** after correction of one factual sentence: the request shares message state once, while the full policy appears in both question instructions. The reviewed draft prominently distinguishes fictional company/messages/policy from actual recorded API answers. Its joint/per-field fractions, review load (21/120 versus 85/120), every named error, urgent-case denominators, top20 coverage, repeat/order results, latency scope, token totals, and $0.1021 per 1,000-request arithmetic projection match retained evidence. The projection remains explicitly unmeasured; priority follows impact rather than adding an independent success measure.

The new article and bounded changes to the builder guide, decision-model chooser, cost-saving playbook, Deals ledger, and coding-agent brief keep direct-route Jev results separate from their older manually authored visual fixtures and unrun OpenRouter/Decisions/Clef workloads. All six final error IDs are visible and retained. Neither the zero observed serious misses nor the synthetic reference agreement is described as human acceptance or a production result. The next Decisions/Clef comparison is future work on the same frozen sets; no comparative result is invented.

Historical $5-credit announcements are separated from current eligibility and the owner’s account report. The October 8 post-run report of displayed $5, no billing history, and “credits expiring in 9 days” is explicitly owner-reported, not a browser observation by the reviewer. It does not establish zero debit, an unrounded balance, a universal lifetime, or an exact calendar expiry. The actual account debit remains unreconciled; no receipt or cash-cost saving is manufactured.

The README now supplies exact offline paths to the bundled primary160, repeat12, and reversed12 records. Bundled `primary-results.json` matches the separate downloadable primary array byte for byte. Bun-only runner tests and replay require no third-party package install or credential. The replay source uses one bounded static fetch and `textContent`, exposes no editable message field or inference route, preserves arrival ties, and selects the heldout set by default. Failed/missing records remain visibly distinct from valid Jev answers.

Reviewed final overlay hashes (paths relative to `research/drafts/jev-inbox-triage/site/`):

| File | SHA-256 |
| --- | --- |
| `content/posts/jev-inbox-triage/index.md` | `9fb7665b892cb0917b5d04c21e1c432b2a2f698c54e67dada5510252be2f6492` |
| `content/posts/jev-practical-builder-guide/index.md` | `d12bdf95ab89c1686afd1bc4204f945a0270cd629de2badd8e75261200b1a4ee` |
| `content/posts/jev-practical-builder-guide/build-with-jev.txt` | `54e507708c69acb9ab4556893b4a603df38fc9c6494a9f0bb05777e905c5dee8` |
| `content/compare/decision-models.md` | `fab0fbf549e3ff46dd5043087b63df1596bdff0c176ffa8843524a18ad0a8677` |
| `content/value/llm-cost-saving-playbook.md` | `f483a63a8355c5160636e962f4133c90e2032f7fe6eb92260affa5c4f2e2518d` |
| `content/value/deals/_index.md` | `1932bdfb4b7d14356c53f235c425dde0fa1c226901cb4ad44055f4e8342ccb63` |
| `static/inbox-triage/replay.js` | `463da04c23d5255eb9cbdd61aee10aff8ed9e4b203b4c10801ecfdf1efb9341a` |
| `layouts/shortcodes/inbox-replay.html` | `9bab56158b44137a9415cee68dc71747e27983b7a24f1f9bc3fd816397b7bf20` |

The reviewer made no API calls and launched no browser/server. The 17-test stub suite and independent offline result calculations passed; the integration owner owns archive extraction, rendered preview, normal-build exclusion, standard checks, hosted quality, candidate commit, draft PR, and their exact-head evidence. Earlier preview failures and their logs remain retained; their eventual correction must be recorded alongside final validation. This source verdict authorizes only the requested draft review milestone, with isolated preview and **no merge or production publication**.
