Tool guides

Tool profile · Official-source review

vLLM Decision 2.0: Six Self-Hosted Decision Models

Decision 2.0 is the vLLM Semantic Router team's model family for bounded text decisions, not version 2.0 of the vLLM inference engine.

Checked October 6, 2026 · Documentation review, no hands-on product test

Choose a variant by workload and operating limits

The official collection contains six model variants. Keep them as one family when comparing deployment choices; the name on a checkpoint is not always its exact parameter count. Official collection.

Original model cards; names, actual parameters and context
Model card / nameActual parametersContext tokens
Kai-0.6B0.60B8,192
Eos-0.8B0.75B16,384
Sol-2B1.88B16,384
Nox-4B4.21B16,384
Lux-9B7.94B16,384
Vega-27B29.37B32,768

Values above come from the individual linked cards. No latency or quality ordering is implied by table position.

Documented interface and runtime

The cards describe text/JSON system_one inference with Choice, yes/no and Score probabilities. This review did not establish an image request path across the family; a base model's capabilities do not supply a supported image interface here. Nox interface · Lux interface.

Cards declare Apache-2.0 and document Transformers 5.17 or later, PyTorch, safetensors and custom model code; Vega additionally lists peft. Review that code and dependencies before deployment. These are downloadable models, not an included free hosted API. Vega requirements.

What a deployment comparison should measure

Start with the maximum request length you actually need and an available memory envelope. Evaluate candidates on identical held-out inputs, with the same answer options and error costs. Include loading time, serving memory, latency distribution and fallback frequency.

Published single-GPU measurements are author results for particular settings. Do not transfer one variant's latency, context or behavior to another. Community ONNX/MLX conversions may be useful alternatives, but they are not the original team's release merely because they share a model name.

This is an evaluation outline, not a completed test. Use the Decision Models selection guide to compare local operation with hosted or Workers AI routes.

Evidence and next step

Sources are linked beside the claims they support. Workflow suggestions are editorial examples, not completed tests; no account, installation, paid generation or benchmark was used for this profile.

Start with the official product or project and confirm the exact access, version and terms needed for your task. How we review.