Omar (kouhxp)'s CPU decision model: a Qwen3.5-0.8B fine-tune shipped as GGUF for llama.cpp that answers yes/no, choice and score questions with a probability per option plus a 'none of these' reject probability, served by a local Jev-style HTTP runtime.
A local, Jev-compatible decision model from Rizzo AI Academy. XHToken's Spark-X2.5-4B with a merged typed-decisions LoRA, run on llama.cpp; yes/no, choice and score questions share one prefill of the state and are read from the answer-letter logits. No text is generated.
Decides
choice, score, noul, classify, route
choice, score, noul, classify, route
Architecture
gutsy
rizzo-flow
Fine-tuned from
qwen/qwen3.5-0.8b
xhtoken/spark-x2.5-4b
License
apache-2.0
apache-2.0
Availability
Open weights
Open weights
Hosted by
—
—
Input price
—
—
Decision accuracy
73.2%
64.8%
Calibration error
—
0.112
Valid action rate
—
—
Median latency
—
195 ms
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between gutsy and rizzo-flow?
gutsy is from Omar (kouhxp) and rizzo-flow from Rizzo AI Academy. Both have open weights you can download and run. Both answer choice, score, noul, classify and route questions. gutsy is the smaller model, at 800M parameters to 4.0B.
Which is more accurate, gutsy or rizzo-flow?
They report on different suites — gutsy 73.2% on JevBench public set (231 items; 169 correct), the maker's own run with Q8_0, rizzo-flow 64.8% on LocalLLaMA/typed-decisions test split (400 cases, 2,000 decisions), Q8_0 — so the numbers do not rank them. Test both on your own labelled examples.
Which is cheaper, gutsy or rizzo-flow?
gutsy: Free (open weights). rizzo-flow: Free (open weights). Open weights cost nothing per call beyond your own hardware.
Can I run gutsy or rizzo-flow locally?
Yes, both: systemone pull kouhxp/gutsy and systemone pull rizzo-ai-academy/rizzo-flow download the weights.
Evaluation suite
JevBench public set (231 items; 169 correct), the maker's own run with Q8_0
LocalLLaMA/typed-decisions test split (400 cases, 2,000 decisions), Q8_0