The flagship of the vLLM Semantic Router team's Decision 1.0 family. Qwen3.5-9B with a shared candidate head returns a probability for every supplied answer to choice, yes/no and score questions, over a 16,384-token input.
Respan's behaviour-scoring model for evals, guardrails and monitoring. For each plain-language behaviour you define, it reads a conversation or agent trace and returns the probability the behaviour is present, absent or not observable, in one forward pass.
Decides
choice, score, noul, classify, route
noul, classify
Architecture
decision
span
Fine-tuned from
qwen/qwen3.5-9b
—
License
apache-2.0
proprietary
Availability
Open weights
Hosted API
Hosted by
—
Respan
Input price
—
$0.020/MTok
Decision accuracy
77.4%
—
Calibration error
—
—
Valid action rate
—
—
Median latency
—
—
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between decision and span-01?
decision is from vLLM Semantic Router and span-01 from Respan. decision has open weights you can download and run; span-01 is only available as a hosted API. Both answer noul and classify questions. Only decision answers choice, score and route. decision is licensed apache-2.0; span-01, proprietary.
Which is more accurate, decision or span-01?
Only decision publishes an accuracy figure (77.4% on vLLM-SR decision benchmark (54 tasks, 3,766 decisions, weighted)); span-01 does not, so there is no comparison to make without your own test.
Which is cheaper, decision or span-01?
decision: Free (open weights). span-01: $0.02 / $0 per 1M. Open weights cost nothing per call beyond your own hardware.
Can I run decision or span-01 locally?
decision yes — systemone pull vllm-semantic-router/decision downloads its weights. The other is only served as a hosted API.