TypeSafe Jev
Hosted TypeSafe API; text decisions. Separate provider from AutoTrust.
Official sourceCategories / Decision Models
Decision models turn supplied context into bounded answers and probabilities for code to act on, rather than composing free-form prose. Start here when the output must be a choice, a rubric score or a yes/no judgment.
Last checked: October 2, 2026
Use bounded outputs for routing a ticket, classifying a record, scoring a rubric, checking a policy gate or verifying a proposed action. In an agent hot path, your code can branch on a probability without parsing an essay. Unlike a fixed classifier, these systems let you describe the decision choices in the request; that does not remove the need to evaluate each new task.
Four profiles, grouped by deployment path—not a performance ranking.
Hosted TypeSafe API; text decisions. Separate provider from AutoTrust.
Official sourceOpen weights with a typed decision path and a separate generation path. Operate and validate the runtime yourself.
Official sourceWorkers AI or Apache-2.0 open weights; typed text/image decisions.
Official sourceApache-2.0 weights and text/image inference code; plan for CUDA memory.
Official sourceChoice selects from named options with a distribution. Score uses an ordered rubric. Noul expresses a yes/no probability. Check each provider’s field names and runtime; an analogous primitive is not proof of wire-level API compatibility.
Define thresholds using held-out examples, including uncertain and out-of-domain inputs. A probability is not permission: enforce authorization and deterministic safety checks outside the model.
Use a general chat or generative model for essays, creative assets and open-ended dialogue. A decision interface is for a bounded output contract, even when its underlying model also supports generation. It is a component, not a complete automation agent or coding environment.
Vendor evaluations differ in datasets, prompts, hardware and aggregation. They cannot be combined into an AI Tool Finder ranking. We have not independently reproduced them; compare deployment fit first, then test the same workload and failure costs.