A 0.4B decision model made from the first 20 layers of Qwen3-0.6B plus a small attention head that compares options. It answers choice, noul and score questions in one pass, and choice order cannot change its answer by construction. Non-commercial licence.
Omar (kouhxp)'s CPU decision model: a Qwen3.5-0.8B fine-tune shipped as GGUF for llama.cpp that answers yes/no, choice and score questions with a probability per option plus a 'none of these' reject probability, served by a local Jev-style HTTP runtime.
Decides
choice, score, noul
choice, score, noul, classify, route
Architecture
bev-decider
gutsy
Fine-tuned from
qwen/qwen3-0.6b
qwen/qwen3.5-0.8b
License
cc-by-nc-4.0
apache-2.0
Availability
Open weights
Open weights
Hosted by
—
—
Input price
—
—
Decision accuracy
74.7%
73.2%
Calibration error
—
—
Valid action rate
—
—
Median latency
—
—
p95 latency
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between bev-decider and gutsy?
bev-decider is from Avishek Biswas and gutsy from Omar (kouhxp). Both have open weights you can download and run. Both answer choice, score and noul questions. Only gutsy answers classify and route. gutsy reads up to 8K tokens of state, against 2K tokens for bev-decider. bev-decider is the smaller model, at 478M parameters to 800M. bev-decider is licensed cc-by-nc-4.0; gutsy, apache-2.0.
Which is more accurate, bev-decider or gutsy?
They report on different suites — bev-decider 74.7% on avbiswas/bev-decision test split (5,000 held-out questions over 2,617 states), the maker's own, gutsy 73.2% on JevBench public set (231 items; 169 correct), the maker's own run with Q8_0 — so the numbers do not rank them. Test both on your own labelled examples.
Which is cheaper, bev-decider or gutsy?
bev-decider: Free (open weights). gutsy: Free (open weights). Open weights cost nothing per call beyond your own hardware.
Can I run bev-decider or gutsy locally?
Yes, both: systemone pull avishek-biswas/bev-decider and systemone pull kouhxp/gutsy download the weights.
—
Evaluation suite
avbiswas/bev-decision test split (5,000 held-out questions over 2,617 states), the maker's own
JevBench public set (231 items; 169 correct), the maker's own run with Q8_0