Use code for exact rules, a decision model for a bounded judgment about meaning, and a generative model when you need new text or an explanation. A queue ID is a decision; drafting the reply is generation. Start with the simplest route that meets your acceptance criteria.
Which can you try now? Jev has documented TypeSafe and OpenRouter routes. Clef has Workers AI routes and downloadable weights. Ordinary GPT-6 Luna has a documented API. OpenAI announced its separate Decisions API in limited preview on September 29; a planned broad release does not establish your access today. Confirm access before designing around it. OpenAI announcement.
This is a documentation-led chooser checked October 3, 2026, not a hands-on benchmark. We ran no model inference. The recommendations below are application design judgments; provider specifications and limitations are linked where they matter.
Choose the job before the model
| Option | Input/output and fit | Access, cost and evidence |
|---|---|---|
| Code | Exact fields → results. Rules, dates, IDs and permissions; ambiguous meaning needs another step. | Your code; compute/maintenance costs. Design recommendation. |
| Jev | Text → Choice/Score/Noul. Semantic routing; no native images or new prose. | Hosted TypeSafe/OpenRouter; input tokens, free output. Docs; inference not-run. |
| Clef / Clef Flash | Text/JSON/images → decisions. Labels and routing; explanations need generation. | Hosted Workers AI or local weights; input tokens or hardware/operations. Docs; inference not-run. |
| OpenAI Decisions | Text/images → finite answers. Bounded routing; contract unresolved. | Limited preview announced; eligibility/price unresolved. Vendor announcement. |
| Generative GPT-6 Luna | Text/images → text or structured output. Extraction/explanations; validate meaning. | OpenAI API; input/cache/output tokens. Pricing owner. Docs; inference not-run. |
Jev’s API reference, Clef’s hosted model reference and Flash reference, and Luna’s model card establish these contracts. A listing is not evidence of successful access from your account.
Choose Jev for a text-only decision loop with a documented small typed interface. Choose Clef when native images or control over a local deployment are material requirements. Choose generative Luna when the same step must extract arbitrary fields or write the explanation: adding a second model solely for a label can complicate an otherwise simple workflow. Keep OpenAI Decisions as an evaluation candidate until your account has access and an authoritative implementation contract.
Already using Luna? Keep your current workflow as the baseline. Compare rules, that workflow and one decision route on the same labeled requests before adding a service. Start read-only with review; this article establishes a local-weights path only for Clef.
What “Choice”, “Score” and “Noul” mean
In Jev, Choice picks among alternatives you supply and returns their probabilities. Score returns a probability-weighted position across ordered rubric levels; it need not equal one discrete level. Noul estimates the probability that a yes/no proposition is true. Choice and Score also expose confidence; Noul has no separate confidence field. TypeSafe primitives.
A supplied URL can be a Choice candidate. Jev can select it; it cannot invent a new URL or write arbitrary prose. Conversely, Luna supports Structured Outputs, so general generation does not have to be unstructured. A valid enum or JSON schema establishes shape, not correctness. Handle refusals and incomplete responses before using generated fields.
Keep each provider route separate
Jev direct: POST https://api.typesafe.ai/v1/systemone, with jev-1.13.0 or the moving alias jev-latest. OpenRouter Jev: POST https://openrouter.ai/api/alpha/decisions, with typesafe/jev-1.13 or ~typesafe/jev-latest. OpenRouter also documents /api/v1/systemone. Its Decisions/SystemOne routes serve TypeSafe Jev; they are not OpenAI Decisions. Credentials, billing and logging belong to the chosen route. OpenRouter integration.
Jev’s direct limits are 64k tokens for state plus all questions and 32k for state plus the longest question. OpenRouter lists a 32,000-token context. Do not transfer the direct packing limit to the broker route. No Jev weights for self-hosting were identified in the reviewed documentation. Use the existing Jev builder guide for its request example, agent handoff and synthetic routing workbench. TypeSafe models.
Clef hosted: Workers AI uses @cf/cloudflare/clef or @cf/cloudflare/clef-flash; the request’s model field is respectively clef or clef-flash. REST wraps the result in Cloudflare’s response envelope; a Workers binding returns the inference result. Do not copy a parser across those interfaces without checking it. REST interface.
The hosted image field accepts embedded PNG/JPEG/WebP images, not remote URLs: up to four, at most 4 MiB and 16 megapixels each, 8 MiB total decoded and a 13 MiB whole request. Both cards list 65,536 context tokens and warn that long text state is truncated. Make source completeness an application check. Clef schema.
Clef local: Cloudflare publishes 27B Clef and 9B Clef Flash weights and code under Apache-2.0, with local multimodal examples including video frame arrays. Local execution needs compatible hardware, a serving implementation and operational ownership; downloading weights is not turnkey deployment. The hosted descriptions mention video, but the reviewed hosted schema exposes images and no video field. We therefore leave hosted video integration unresolved. Cloudflare release, local model card.
OpenAI Decisions: the announcement establishes Luna-backed questions with finite answers and text/image context. It does not supply a verified endpoint/schema, confidence semantics, pricing or current eligibility in our reviewed sources. We provide no executable request and borrow no ordinary Luna prices. These sources do not establish an end-to-end latency benchmark or SLA.
Three small workflows to evaluate
All inputs and permitted labels below are synthetic, conceptual examples. They are not model responses or offline-validated API code. Your application supplies the labels; the model advises which fits.
A. Route a support request
Input: “I changed plans yesterday; now checkout fails.” Code first applies explicit rules such as account suspension and known incident IDs. Give the remaining text and queue definitions to Jev or Clef.
Permitted outputs: billing, technical, sales, needs_review. Code validates the returned ID and queues the request; it does not issue refunds or change account permissions. If queues overlap, a result is missing, the request times out or confidence fails a tested threshold, route to review. A confident wrong queue remains possible.
Next step: label a small set of real, redacted support requests, including mixed intents and missing context. Compare rules alone, rules plus a decision model, and your existing Luna workflow before adding a second service.
B. Triage a registry claim against a passage
Input: claim “The programme launched nationwide”; passage “The bank expanded its limited pilot to two additional cities.” Code checks the entity ID, retrieval date, source URL and exact quote presence before semantic assessment.
Permitted outputs: supports, contradicts, insufficient_evidence. The human reference label for this fixture is insufficient_evidence: expansion of a pilot does not establish nationwide launch. This is our label, not an observed model answer.
Code attaches the source and decision to a research queue. Missing context, stale sources, truncated input and disagreement go to review. Legal conclusions and publication remain independently reviewed even when confidence is high. A passage classifier cannot authenticate its source or establish facts outside it.
Next step: build a held-out set of source/claim pairs with human labels and retrieval checks. Review false support decisions separately; their cost differs from an unnecessary escalation.
C. Classify a screenshot
Input: an invented screenshot showing a checkout error. Permitted outputs: checkout_error, login_screen, other, unreadable. Code checks file type/size, then submits an embedded image to a documented hosted Clef route, or uses a separately validated local image pipeline. OpenAI Decisions is an image-capable announced candidate, subject to access and contract verification. Jev requires a separate text extraction/perception step.
Code queues the image for an approved handler and keeps its source reference, retrieval record and image hash attached. A label does not authenticate the image. Cropping, unreadable text, over-limit images, API failure and text inside the image that tries to redirect the workflow go to review. If the task is “extract every error message and explain it,” generative Luna’s image input and structured text output can be the simpler single step.
Next step: collect redacted screenshots with human labels, including blank and misleading images. Test perception errors separately from routing errors; a valid label is not proof that the image was understood.
Confidence, cost and privacy decide production fit
Confidence is a signal to evaluate, not permission. Jev’s Choice/Score confidence measures concentration of its probability distribution, not measured accuracy on your workload. Its known limitations include option-order bias and distracting information. Test shuffled choices, adversarial instructions, missing correct answers and confident mistakes. Clef’s published local adapter computes confidence differently from TypeSafe; the hosted formula is unresolved. Re-evaluate thresholds for every model and route. TypeSafe confidence, Clef local code.
Set review thresholds using labeled development data and check them on a held-out set. Measure false acceptance, missed escalation, review volume and p50/p95 end-to-end latency. Missing answers, malformed values, refusals, timeouts and exhausted bounded retries must have an explicit fallback. Keep permissions, allowlists and approval gates in code at every threshold.
Prices checked October 3: Jev 1.13 lists $0.042 per million input tokens with free output; hosted Clef lists $0.24 and Clef Flash $0.09 per million input tokens. Workers AI bills through its Neuron system, with a daily allowance and paid overage; Workers execution/storage and local serving are separate costs. TypeSafe rates, OpenRouter Jev rates, Workers AI pricing.
For 1,000 requests × 1,000 billed input tokens, Jev’s listed inference arithmetic is $0.042. It is not a workload bill or a quality comparison. Count the rubric, repeated state, retries, image accounting, generative output/reasoning, downstream work and human review. Use cost per accepted result rather than token price alone. Keep detailed Luna rates on the pricing owner; they do not price OpenAI Decisions. Older GPT-5.6 Luna findings do not evaluate GPT-6 Luna. No universal speed, accuracy or savings winner is established here.
Privacy follows the whole route. TypeSafe’s no-training policy does not establish default zero retention; enterprise ZDR is separately offered. OpenRouter’s storage opt-ins and its provider’s retention policy both matter. TypeSafe privacy, enterprise terms, OpenRouter collection and provider logging.
Workers AI’s data-use policy does not erase AI Gateway logs, which are enabled by default per gateway, or application logs. Local Clef can keep inference under your control; telemetry, storage and network configuration determine whether inputs stay local. Ordinary OpenAI API abuse logs normally retain content for up to 30 days, with legal/safety exceptions; Responses storage and approved retention controls have additional conditions. Do not transfer those details into an undocumented Decisions contract. Workers AI data usage, AI Gateway logging, OpenAI API data controls.
Practical verdict: keep exact decisions in code. For bounded semantic judgments, evaluate Jev for hosted text and Clef when images or local control matter. Use generative Luna when producing the explanation is part of the job. Wait for verified OpenAI Decisions access and documentation before committing an integration. Start read-only, preserve review, and agree a bounded task set and spending cap before paid inference.