Categories / Decision Models

Workers AI + open weights

Cloudflare Clef: Hosted and Open Decision Models

Cloudflare Clef and Clef-flash return bounded decisions with probabilities for routing, scoring and agent gates. They are available through Workers AI and as Apache-2.0 open weights.

Official-source review · Last checked October 2, 2026 · No independent product benchmark

Two models, one decision workflow

Use @cf/cloudflare/clef (27B) or @cf/cloudflare/clef-flash (9B). Cloudflare documents a 64K context window, up to 64 questions per request and up to four images alongside text state. Workers AI supports bindings, REST and AI Gateway. Workers AI release details.

Output contract and deployment

Cloudflare documents System One/Jev API compatibility: Choice selects an option with probabilities; Score evaluates an ordered rubric; Noul returns a yes/no probability. Endpoint and model changes still need integration testing. Clef API schema · Clef-flash API schema.

Choose hosted Workers AI to avoid operating inference servers, or inspect the Apache-2.0 weights linked from the official announcement for self-hosting. Check current Workers AI billing and limits; open weights do not make inference infrastructure free.

Reported results are not our ranking

Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash in its published benchmark. These are vendor-run measurements, not an end-to-end latency guarantee. AI Tool Finder did not independently reproduce the decision benchmark or latency results. Published methodology and results.

Cloudflare’s RL fine-tuning work is seeking design partners; do not treat the announcement as a generally available self-service fine-tuning product.

A practical decision gate

Define the allowed outcomes before sending state. Log the model version, question, distribution and downstream action; route uncertain cases to review. For destructive actions, authorization must remain outside the model. Validate costs and error rates on your own held-out examples before putting a model in a latency-sensitive agent path.

Official sources

Choose your next step

Decision Models