# Fictional inbox authoring and review record

Authoring completed on 2026-10-08 (`+08:00`). Status at author handoff: **ready for independent review; not yet frozen**. The labels are Codex-authored reference judgments under a fictional policy. They are not human-validated customer labels, and author checks are not independent review.

## Policy before labels

The corpus author wrote `policy.md` first and measured its SHA-256 before creating the reference labels:

`f817cc6325660dc1ef8a7dd3215360bb12607f7e87dd7c9688ab9c8cfcc24b3e`

LatticeDesk is invented. No actual inbox, customer message, company policy, credential, account identifier, support commitment, or vendor pricing was copied into this corpus. The author made no Jev API requests and did not inspect model or baseline predictions while authoring or revising these fixtures.

The policy deliberately separates queue ownership from current harm. Payment-related suspension belongs to billing; identity, membership, and permissions belong to account; a fault with known valid access belongs to technical; purchase and plan-fit questions belong to sales. Equally important unrelated requests and unknown ownership belong to review. A known owner can still have unclear impact. A review-owned message can still be blocked.

Priority follows the fixed impact mapping in code. Tone, status, a deadline alone, quoted old failures, and customer-supplied routing instructions do not change that mapping. Present work with a viable workaround is degraded; an explicit current task with no viable route is blocked. Resolved incidents and harmless administration are routine.

## Author choices

Each message was written individually. The author did not expand one template across customers or take a production ticket sample. The cases vary workflows, affected roles, payment states, access prerequisites, workaround availability, current versus historical harm, spelling, quoted material, and unrelated requests. Some terminology necessarily repeats because all tickets concern one fictional software product.

The development split has 40 messages. The primary heldout split has 120 messages. The diagnostic challenge split has 40 messages and intentionally concentrates misleading instructions and subtle counterfactuals. Splits are deliberately authored rather than random samples of a real arrival distribution.

There are 21 paired counterfactual groups (42 messages): two pairs in dev, nine in heldout, and ten in challenge. A pair changes a specified cause, workaround, resolution, context, or tone while retaining the relevant surrounding situation. Both members stay in the same split. The other 158 messages have singleton groups. Paired messages are intentionally similar; they are not independent prevalence observations.

The author selected the stability subset before any results were read: `H013`, `H026`, `H055`, `H066`, `H081`, `H097`, `C002`, `C008`, `C011`, `C023`, `C032`, `C040`. The protocol keeps the case sequence fixed and reverses both question order and choice order in the order probe. That probe cannot isolate those two effects separately.

Only `message` is ticket input. IDs, groups, splits, scenario labels, expected labels, and annotations are evaluation metadata and must never enter a model request. The separate fictional policy is allowed instruction input. Scoring uses fixed denominators, retains failures, and orders top-20 suggestions by priority then corpus arrival order. Provider confidence remains separate for queue and impact; no joint probability is manufactured.

## Author self-review and fixture checks

The initial authoring script was saved at 2026-10-08 19:58:56 `+08:00`, immediately before its initial corpus write in the same command. The author then clarified three messages before seeing any predictions: `D022` made cosmetic wording refer to a board heading rather than a welcome email; `D028` stated a concrete fifteen-minute onboarding delay; `H044` stated that a trial workspace was being used for a live shift checklist. `H044`'s rationale was clarified as well. No reference labels changed. The resulting corpus filesystem modification time was 2026-10-08 20:01:35 `+08:00`. The integration owner separately checked that this corpus hash matched the development-run manifest; the earlier development smoke retained its own input snapshot.

The author parsed both JSON files using Bun and checked all 200 required object shapes, enum membership, priority mapping, nonempty text, unique IDs, distinct message text, exact split counts, group nonleakage, and the existence of the 12 stability IDs. Those fixture checks passed. They check structure and author intent; they cannot establish that every reference judgment is unambiguous.

The following tables retain the initial author handoff counts, before independent review. The final reviewed counts appear below.

| Split | Cases | Billing | Account | Technical | Sales | Review |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| dev | 40 | 9 | 11 | 10 | 5 | 5 |
| heldout | 120 | 27 | 25 | 35 | 20 | 13 |
| challenge | 40 | 10 | 10 | 10 | 2 | 8 |
| Total | 200 | 46 | 46 | 55 | 27 | 26 |

| Split | Blocked | Degraded | Routine | Unclear |
| --- | ---: | ---: | ---: | ---: |
| dev | 12 | 5 | 19 | 4 |
| heldout | 34 | 16 | 57 | 13 |
| challenge | 16 | 4 | 14 | 6 |
| Total | 62 | 25 | 90 | 23 |

All 11 controlled scenario labels appear: ordinary (68), calm-serious (7), loud-low-impact (7), resolved (6), quoted-history (7), mixed (15), missing-context (18), negation (9), typos (9), policy-injection (12), and counterfactual (42). Scenario is a single primary slice label; categories can overlap conceptually, so these slice names are not exhaustive descriptions of each message.

Author handoff hashes, **not a freeze declaration**:

| File | SHA-256 |
| --- | --- |
| `policy.md` | `f817cc6325660dc1ef8a7dd3215360bb12607f7e87dd7c9688ab9c8cfcc24b3e` |
| `corpus.json` | `d1c6c1ca3638df839fc3060c70406dbaa778cd72ae4c8488c5c4e2ffa18f215b` |
| `protocol.json` | `824c3078e6793f5470077ce7a6b8866ece94d9fa5ae3d371e03b00b7314313d4` |

## Independent review and freeze

Independent dataset review: **completed by a fresh Sol 6.1 reviewer**, who read all 200 cases without model or baseline results. Human label validation: **not performed**. Output review and operational acceptance are separate. See [independent-review.md](independent-review.md) and the final `freeze.json` for the authoritative hashes and evaluation gate.

The reviewer corrected two ambiguities before any heldout/challenge inference. `H004` now explicitly establishes paid subscription and functioning sign-in, so its dashboard loading fault belongs to technical/degraded. `C007` states access revocation was already completed, leaving only a current optional quote, so its owner is sales/routine. The original corpus is retained in `history/corpus-development-v1.json`; the original proposed protocol is retained in `history/protocol-proposed-v1.json`.

| Final reviewed split | Cases | Billing | Account | Technical | Sales | Review |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| dev | 40 | 9 | 11 | 10 | 5 | 5 |
| heldout | 120 | 26 | 25 | 36 | 20 | 13 |
| challenge | 40 | 10 | 10 | 10 | 3 | 7 |
| Total | 200 | 45 | 46 | 56 | 28 | 25 |

One smoke and 40 development calls preceded independent review. The original proposed protocol's broader before-any-scored-call gate was unmet; this is disclosed rather than relabeled as a pass. The user's task permits development tuning and requires freeze before final evaluation. No development reference label changed; one rules revision used development data only. Final evaluation and stability calls are gated by independent review and an exact source freeze.

The reviewer should read the policy and all labels, flag unsupported harm or ambiguous ownership, check the paired changes, verify input isolation and split boundaries, and record findings without access to heldout/challenge outputs. Resolve substantive findings before freezing exact policy, corpus, protocol, prompts, and rules hashes. Preserve this author record and any pre-freeze run snapshots. Do not revise reference labels because a scored model disagrees.

These authored fixtures can support a bounded comparison under this policy. They cannot establish performance on real tickets, actual arrival prevalence, production safety, human acceptance, or cost per accepted result (CAR).
