A useful inbox triage system identifies the team that can restore work and the current operational impact. A deadline or an angry subject line should not decide that alone. We built a small read-only example: fictional customer messages go through a fixed policy, code derives priority, and a human keeps control of the response.

The company, messages and policies are fictional. The Jev responses and measurements come from actual recorded API runs. The reference labels are Codex-authored and have not been validated by human support operators. This recorded LatticeDesk example suggests routing only: nothing sends a reply, issues a refund, changes an account or executes message instructions.

On the 120 held-out messages, Jev matched both reference labels in 115 cases (95.8%); the frozen rule baseline matched 40 (33.3%). Jev gave every one of the 34 reference-blocked cases urgent priority. This is useful evidence for a shadow-mode trial on the same kind of work, with a small synthetic sample and a deliberately simple baseline.

For Jev’s primitives and integration options, use the builder guide. This article follows one concrete workflow rather than surveying use cases. The decision-model chooser keeps other providers’ contracts separate.

The small job

Each message asks two independent questions: which queue owns the current work, and what is its current impact? Jev returns a bounded Choice for each. Code maps impact to priority, validates the returned labels and keeps missing or invalid responses in review.

DecisionAllowed valuesImportant boundary
Queuebilling, account, technical, sales, reviewChoose the restoring team; equally important unrelated problems need review.
Current impactblocked, degraded, routine, unclearJudge stated present harm. Tone, an amount of money and a deadline do not establish an outage.
Code-derived priorityurgent, high, normal, reviewRespectively maps from blocked, degraded, routine and unclear. The model does not choose a support action.

A viable workaround means degraded rather than blocked. A resolved incident or quoted historical outage is routine unless another current problem is stated. A known owner with unspecified harm retains its queue and gets unclear impact. The full fictional policy is part of the reproducible example, not a real service commitment.

Inspect the saved decisions

The replay opens on the held-out set. Choose a corpus set and message, compare the reference and baseline with the saved answer, then switch from arrival order to suggested priority. The original arrival number stays visible. This is a replay of recorded data; it accepts no new text and makes no API calls.

LatticeDesk / fictional support inbox

Arrival order → suggested queue

This replays saved responses to author-written messages. It makes no inference requests and takes no support actions.

This interactive replay requires JavaScript; the article and downloads remain available.

Without JavaScript, the article contains the results and you can download the saved inputs and responses. No response is generated in your browser.

Author reference, rule baseline and recorded Jev answer are separate evidence. Confidence is a returned model value, not a measured accuracy guarantee. Calculated cost is input usage × the documented rate, not a reconciled account debit.

Download the recorded replay JSON

What we measured

The corpus has 200 Codex-authored messages: 40 development cases, 120 held-out cases for the primary evaluation, and 40 deliberately difficult challenge cases. Related variants stay within one split. Only each message enters model state; IDs, scenario names, split, annotations and reference answers stay local. The policy and Choice criteria are supplied as question instructions.

The chronology matters. One smoke call and the 40-case development run preceded independent label review. The author-handoff protocol proposal was written after that development run and initially required review before any scored call; that historical gate was not met. Its correction requires review before the final evaluation. We retained that post-development proposal and the original corpus instead of calling the earlier gate passed. The earlier smoke/development manifests retain different protocol hashes, but their executed protocol snapshots were not retained; the proposal is not presented as those missing inputs. A separate Codex reviewer read all 200 labels without seeing model outputs, corrected C007’s reference to sales/routine after noticing the account task was already completed, and clarified H004’s current dashboard-loading fault with paid access and sign-in working. The final held-out and challenge evaluation uses the corrected, frozen corpus and retained frozen protocol.

One rule-baseline revision used the development cases and fit all 40 development references. That 40/40 is development fit, not evidence of generalization. The Jev questions were unchanged. The replay uses the frozen rule revision when comparing development messages; the original development run with its earlier rule baseline is retained in the evidence. Neither held-out nor challenge results should be used to choose a better prompt or quietly revise a label.

The primary denominator is fixed at 120 held-out cases; the 40 challenge cases are reported separately, with the combined 160 shown only as a description of this sample. Missing, failed and malformed answers count as errors. Queue and impact are separate model decisions; priority agreement follows the impact-to-priority map, so it is not a third independent test.

The held-out run completed October 8, 2026, on the direct TypeSafe route with jev-1.13.0, two Choice questions per request and frozen inbox-rules-1.1 / inbox-scoring-1.1. All 120 responses passed format/model/usage validation; no held-out case is missing and no retry was made.

Held-out resultRecorded JevFrozen rules
Queue agreement117/120 (97.5%)75/120 (62.5%)
Impact / derived-priority agreement118/120 (98.3%)61/120 (50.8%)
Both labels agree115/120 (95.8%)40/120 (33.3%)
Blocked cases without urgent priority0/3418/34
Unnecessary urgent priority0/861/86
Queue or priority asks for review21/12085/120

The rules often abstained rather than assigning a wrong team with confidence. This metric counts an abstention as a non-match whenever the reference provides an answer; it does not mean the baseline would discard the message. Its weak held-out result describes this small, development-fitted ruleset, not every possible deterministic implementation or a human support team.

With a review budget of 20 messages, arrival order surfaced 6 of the 34 blocked cases. Rule priority surfaced 16; Jev priority surfaced 20. Those are 17.6%, 47.1% and 58.8% of the blocked set, respectively. Jev’s first 20 were all reference-blocked, but 14 blocked cases necessarily remained beyond that budget. Twenty is a descriptive cutoff here, not an optimized production policy.

The challenge set gave 39/40 joint matches for Jev and 12/40 for rules, with no Jev missed-urgent flags among its 16 blocked references and eight for rules. Combined, the two sets produced 154/160 matches (96.25%) versus 52/160 (32.5%). That combined number describes this constructed sample; it does not replace the held-out result or turn the challenge cases into a real-world prevalence estimate.

For reproducibility, inspect the full report and confusion matrices, freeze record and exact hashes, blinded independent review, dataset review and retained earlier protocol/corpus history. The source archive includes all 225 original call records; the main evaluation remains 160 unique messages.

The reference answers assess agreement with this authored policy. They do not prove customer satisfaction, escalation safety or reduced staffing. Synthetic examples can expose a failure mechanism, but cannot estimate real support traffic without a representative independent sample.

Successful decisions and errors

A useful success, H008: a payment-state dispute had disabled the workspace, preventing a service crew from reading the day’s assignments. The rule baseline selected billing but left impact unclear. Jev returned billing/blocked, which code mapped to urgent. That agreed with the reference without requiring an angry message or the word “urgent.”

Negation mattered in H086: “No one is locked out anymore” preceded a request to change intentionally restricted viewer permissions next month. Rules marked it blocked; Jev returned account/routine, agreeing with the reference that this was future administration rather than a current lockout.

Jev still made five held-out errors. H003 requested a corrected tax address on a receipt while work continued normally. Both systems over-prioritized this routine billing administration as degraded/high. In H090, a misspelled receipt correction went to technical rather than the billing reference; Jev correctly kept its impact routine. H043, H068 and H102 contain the other reference disagreements. Inspect them in the replay instead of treating a valid enum as a correct answer. H068’s cosmetic work and H102’s historical ticket admit policy-boundary interpretations; these scores retain the frozen author references and measure reference disagreement, not independently proven real-world mistakes.

The challenge error, C007, was confident. It requested access revocation and a new quote, then explicitly said access had already been revoked and the quote was optional. The frozen reference was sales/routine; Jev selected account/routine with queue confidence 0.97. This is the case whose reference was corrected during blinded review before the final run. The resulting failure remains in the score. A high-confidence gate could have accepted the wrong owner even though impact was right.

No HTTP, network or format failure occurred in the 225 recorded calls. That is an observation about this run; behavior under outages, rate limits and load was not measured. The runner’s error fallback still retains the message for review.

A small repeat and ordering probe

Twelve cases were called again with the same questions, then again with both question order and Choice criteria order reversed. Each round produced 12/12 valid pairs with no queue, impact or derived-priority disagreement against the main result; the two rounds also agreed with each other. They remain separate diagnostics and do not replace or average into main accuracy.

This tests two ordering changes together, so it cannot isolate their individual effects. Twelve pairs on one model/day do not establish determinism. Returned confidence and probabilities are preserved in the records rather than treated as a promise of identical future answers.

Latency and confidence

For the 120 held-out calls, measured client p50 was 285.1 ms and p95 was 337.6 ms. Challenge calls had p50 279.8 ms and p95 320.6 ms. The rule baseline’s local compute latency was not measured; it made no provider requests.

Timing measures client elapsed time from the HTTP request through response-body reading, JSON parsing and validation. It includes network and provider time, excludes budget admission and disk writes, and cannot isolate model compute. Requests are sequential with concurrency one, no automatic retries and a 30-second timeout. Percentiles use nearest rank. This is one client’s recorded sequence, not a throughput, load or cold-start benchmark.

Choice confidence describes the concentration of the returned probabilities. It does not measure accuracy on this inbox. Inspect confident errors and unclear inputs before picking a routing threshold. The example applies no automatic-action threshold. TypeSafe confidence documentation.

Cost: three separate numbers

TypeSafe’s model page, checked October 8, 2026, lists Jev 1.13 at $0.042 per million input tokens, with output free. The calculation is reported input tokens × 0.042 / 1,000,000. Count the shared message state and both question instructions; the full policy appears in each question. A token count for the message alone understates this example. TypeSafe models.

Recorded scopeCallsReported input tokensCalculated inference value
Held-out + challenge evaluation160388,943$0.016335606
Development + separate smoke4199,628$0.004184376
Repeat + ordering diagnostics2458,410$0.002453220
Entire recorded experiment225546,981$0.022973202

All calls returned usage; the shared budget ledger has no unresolved reservation. The 21,656 reported output tokens have zero calculated charge at the documented rate. Cash purchases by this task were $0: no card, refill or credit purchase was made. The actual account debit and promotional-balance change remain unreconciled. Provider responses supplied token usage without a billing receipt. Do not read the calculated total as an observed charge or as a verified remaining balance.

Post-run account report, October 8: the owner said the dashboard still displayed $5, displayed no billing history, and showed “credits expiring in 9 days.” This is an owner report, not an independently verified browser observation or itemized debit receipt. The unchanged display does not prove zero usage cost or a particular unrounded balance. The exact expiry instant was not supplied; the countdown establishes neither a universal credit lifetime nor a calendar cutoff. Check your own account’s expiry before budgeting a later run.

The final 160 cases averaged about 2,431 reported input tokens per request because the full policy and two questions are substantial inputs. At that same average, model/route price and zero-retry pattern, 1,000 messages calculate to about $0.1021. This is an arithmetic projection, not a measured batch or a saving. Local rules made no paid API call; their engineering and compute cost was not measured. Build-agent compute, hosting, review and cleanup time are also excluded.

LedgerWhat it means here
Calculated inference valueReturned input usage multiplied by the documented price. It is not a reconciled bill.
Promotional creditAn account balance governed by the provider’s terms. A balance change needs before/after account evidence.
Cash spendingAn actual paid charge. Existing promotional credit is not cash paid for this run.

The owner reported on October 8 that a TypeSafe signup with Gmail automatically received $5; the signup date was not supplied. That one report does not establish a Gmail requirement, recurring credits or current new-account eligibility. TypeSafe’s September 20 launch announcement advertised $5, but its September 27 notice said new signups no longer receive free credits. No later restoration was verified in this review. The announcement text was retrieved from X’s first-party syndication records; the X webpages themselves could not be opened.

Treat this as historical promotional access, not a current free-tier recommendation. TypeSafe’s MCA section 8.2 makes promotional credits discretionary, allows additional terms including expiry, and says promotional credit is consumed before purchased credit. Check your own balance and applicable terms before running anything paid. The Deals ledger retains the status; the cost-saving playbook explains the separate ledgers.

Full cost per accepted result remains undefined without measured review and cleanup time and a defensible time valuation. Queue/impact agreement and calculated API value are narrower outputs; neither establishes a complete CAR or a saving against another model.

Run the example locally

Download the Bun experiment source archive and extract it in a new working directory. It includes the policy, questions, corpus, deterministic rules, runner, scoring and tests. No third-party JavaScript package installation is needed.

1
2
tar -xzf experiment.tar.gz
bun experiments/inbox-triage/run.ts --mode offline --split dev

This command runs the 40 development messages through the local rule baseline, writes a fresh result directory and prints its path. It makes no Jev requests: its rows are offline, provider latency and token cost are absent, and the Jev report is not-run. The recorded replay JSON contains the separate actual calls shown above.

To recalculate the saved comparison without inference, download the primary rows, repeat rows and reversed-order rows beside your extracted source, then use a fresh output filename:

1
2
3
4
5
6
bun experiments/inbox-triage/report.ts \
  --primary results-primary.json \
  --split evaluation \
  --repeat results-repeat.json \
  --reversed results-reversed.json \
  --out report-replayed.json

For an explicitly authorized live check of just the first development case:

1
2
bun --env-file=/absolute/path/to/ignored.env \
  experiments/inbox-triage/run.ts --mode live --split dev --ids D001

The private environment file supplies JEV_API_KEY or TYPESAFE_API_KEY and JEV_BALANCE_FILE. Follow the included README to record current, verified promotional funds and the model/price evidence; an example balance shape is not evidence. The runner fixes the direct endpoint and jev-1.13.0, uses a 30-second timeout, sends one request at a time, refuses redirects and performs zero automatic retries. All live variants share a cumulative $1 ceiling and reserve $0.01 before each attempt; missing usage or failed calls retain their reservation for reconciliation. Keep the ledger rather than resetting it to create headroom.

Use the recorded replay first. A live rerun requires your own authorized API access, a fixed spend cap and the same input/rubric versions. Never paste a real inbox into this demonstration: replace personal information and review your provider’s applicable data terms before building an integration. Direct API documentation.

The decision to make next

This example justifies a shadow-mode Jev trial for a small text-only inbox with a stable policy. It improved urgent-case prioritization and reference agreement over this frozen ruleset at a very small calculated token value. It also made a confident ownership mistake and does not establish the cost or safety of operating a real inbox.

OpenAI Decisions, Clef and a generative-model baseline were not run. No cross-provider ranking, human-time saving, production reliability or complete CAR follows from this result. Existing competent rules or another already-integrated model may still be the simpler option for your workflow.

For the next controlled comparison, run Decisions and Clef on these same frozen 120 held-out and 40 challenge messages, with the fixed 12-case repeat/order subset. Preserve their raw responses and unavoidable API differences, and evaluate each provider’s confidence separately. Keep the existing references and report new model versions as separate run identities.

Keep exact rules, permissions and effects in code. A reasonable next experiment uses representative messages with independently reviewed labels, a separate development set for any tuning, and a held-out evaluation. Track missed blocked work, false escalations and how many messages still need review. Start in shadow mode and keep the current queue as the fallback.

Further reading

Evidence scope: TypeSafe pricing and the historical credit announcements were checked October 8, 2026; source coverage and retrieval limits are retained. The task results were recorded on the same day, using the linked frozen inputs and reports.