Command Code Agr and Agr-flash: English Text Decisions
Agr and Agr-flash are a self-hosted decision-model family for English text and JSON. Compare their request limits and deployment needs; the Flash model is not intended for safety decisions.
Checked October 6, 2026 · Documentation review, no hands-on product test
One family, two deployment tradeoffs
The Command Code project supports typed Choice, Noul and Score answers through Python and a System One server route. Its integration with the TypeSafe SDK does not make it a TypeSafe product. Inputs are English text/JSON, not images. Official repository.
| Variant | Base | Token limits | Deployment note |
|---|---|---|---|
| Agr | Gemma 4 31B IT base | 16,384 / 16,384 / 32,768 | 61.43 GB bf16 weights; author example uses a 96 GB GPU |
| Agr-flash | SmolLM2-360M base | 6,144 / 2,048 / 12,288 | 0.73 GB weights; not intended for safety decisions |
Specs: Agr model card · Agr-flash model card. Weight-file size is not total inference memory, and an author's deployment example is not a verified minimum.
Output contracts still need evaluation
Choice and Score accept 2–255 options or levels. The project distinguishes its confidence field from simply taking the largest option probability, and notes that option order can affect output. Review the runtime's exact contract before swapping an existing integration. Interface and evaluation notes.
Both model cards declare Apache-2.0. You still operate the runtime, provide hardware and review any dependencies. In particular, preserve the author's explicit warning that Agr-flash is not meant for safety decisions. Flash limitations.
How to compare them for your task
An original evaluation brief: route English support tickets into billing, technical, account or review queues using the same labeled examples and option definitions. Measure errors that matter to the queue, along with memory and operating cost. A compact model is only useful if its mistakes are acceptable for that job.
The author's Decision Index is a chance-corrected skill score multiplied by 100, not raw accuracy. It is not our benchmark and cannot establish that either model is universally better. Benchmark definition.
Keep authorization outside the score. For other hosting and input options, see the Decision Models selection guide.
Evidence and next step
Sources are linked beside the claims they support. Workflow suggestions are editorial examples, not completed tests; no account, installation, paid generation or benchmark was used for this profile.
Start with the official product or project and confirm the exact access, version and terms needed for your task. How we review.