The flagship of the vLLM Semantic Router team's Decision 1.0 family. Qwen3.5-9B with a shared candidate head returns a probability for every supplied answer to choice, yes/no and score questions, over a 16,384-token input.
Invergent's decision model for text and images. A fine-tune of Gemma-4-26B-A4B (about 4B parameters active per token) that answers Choice, Noul and Score questions, with an optional thinking mode for harder questions.
Decides
choice, score, noul, classify, route
choice, score, noul, classify, route
Architecture
decision
rune
Fine-tuned from
qwen/qwen3.5-9b
google/gemma-4-26b-a4b-it
License
apache-2.0
apache-2.0
Availability
Open weights
Open weights + hosted API
Hosted by
—
Invergent
Input price
—
—
Decision accuracy
77.4%
—
Calibration error
—
—
Valid action rate
—
—
Median latency
—
—
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between decision and rune?
decision is from vLLM Semantic Router and rune from Surogate (Invergent). decision has open weights you can download and run; rune has open weights and a hosted API. Both answer choice, score, noul, classify and route questions. rune reads up to 262K tokens of state, against 16K tokens for decision. decision is the smaller model, at 9.0B parameters to 26B.
Which is more accurate, decision or rune?
Only decision publishes an accuracy figure (77.4% on vLLM-SR decision benchmark (54 tasks, 3,766 decisions, weighted)); rune does not, so there is no comparison to make without your own test.
Which is cheaper, decision or rune?
decision: Free (open weights). rune: Hosted, price not published, or free to self-host. Open weights cost nothing per call beyond your own hardware.
Can I run decision or rune locally?
Yes, both: systemone pull vllm-semantic-router/decision and systemone pull surogate/rune download the weights.