Tool guides

Tool profile · Official-source review

Command Code Agr and Agr-flash: English Text Decisions

Agr and Agr-flash are a self-hosted decision-model family for English text and JSON. Compare their request limits and deployment needs; the Flash model is not intended for safety decisions.

Checked October 6, 2026 · Documentation review, no hands-on product test

One family, two deployment tradeoffs

The Command Code project supports typed Choice, Noul and Score answers through Python and a System One server route. Its integration with the TypeSafe SDK does not make it a TypeSafe product. Inputs are English text/JSON, not images. Official repository.

Current family limits; token columns are state / question / request
VariantBaseToken limitsDeployment note
AgrGemma 4 31B IT base16,384 / 16,384 / 32,76861.43 GB bf16 weights; author example uses a 96 GB GPU
Agr-flashSmolLM2-360M base6,144 / 2,048 / 12,2880.73 GB weights; not intended for safety decisions

Specs: Agr model card · Agr-flash model card. Weight-file size is not total inference memory, and an author's deployment example is not a verified minimum.

Output contracts still need evaluation

Choice and Score accept 2–255 options or levels. The project distinguishes its confidence field from simply taking the largest option probability, and notes that option order can affect output. Review the runtime's exact contract before swapping an existing integration. Interface and evaluation notes.

Both model cards declare Apache-2.0. You still operate the runtime, provide hardware and review any dependencies. In particular, preserve the author's explicit warning that Agr-flash is not meant for safety decisions. Flash limitations.

How to compare them for your task

An original evaluation brief: route English support tickets into billing, technical, account or review queues using the same labeled examples and option definitions. Measure errors that matter to the queue, along with memory and operating cost. A compact model is only useful if its mistakes are acceptable for that job.

The author's Decision Index is a chance-corrected skill score multiplied by 100, not raw accuracy. It is not our benchmark and cannot establish that either model is universally better. Benchmark definition.

Keep authorization outside the score. For other hosting and input options, see the Decision Models selection guide.

Evidence and next step

Sources are linked beside the claims they support. Workflow suggestions are editorial examples, not completed tests; no account, installation, paid generation or benchmark was used for this profile.

Start with the official product or project and confirm the exact access, version and terms needed for your task. How we review.