French medical decision models by Bofeng Huang: pass a patient message or clinical note, a question and candidate answers, and get one probability per answer from a single forward pass. Choice, score and noul. A research model, not a medical device.
Omar (kouhxp)'s CPU decision model: a Qwen3.5-0.8B fine-tune shipped as GGUF for llama.cpp that answers yes/no, choice and score questions with a probability per option plus a 'none of these' reject probability, served by a local Jev-style HTTP runtime.
Decides
choice, score, noul
choice, score, noul, classify, route
Architecture
docto-decision
gutsy
Fine-tuned from
qwen/qwen3.5-4b
qwen/qwen3.5-0.8b
License
apache-2.0
apache-2.0
Availability
Open weights
Open weights
Hosted by
—
—
Input price
—
—
Decision accuracy
88.9%
73.2%
Calibration error
0.079
—
Valid action rate
—
—
Median latency
47 ms
—
p95 latency
—
—
Figures are from each model’s manifest; accuracy and latency are what the publishers report, on their own suites and hardware. Add a third model.
Questions
What is the difference between docto-decision and gutsy?
docto-decision is from Bofeng Huang and gutsy from Omar (kouhxp). Both have open weights you can download and run. Both answer choice, score and noul questions. Only gutsy answers classify and route. gutsy is the smaller model, at 800M parameters to 4.7B.
Which is more accurate, docto-decision or gutsy?
They report on different suites — docto-decision 88.9% on Docto Decision Bench fr v0.1 (12 tasks, mostly silver labels; the maker's own benchmark), gutsy 73.2% on JevBench public set (231 items; 169 correct), the maker's own run with Q8_0 — so the numbers do not rank them. Test both on your own labelled examples.
Which is cheaper, docto-decision or gutsy?
docto-decision: Free (open weights). gutsy: Free (open weights). Open weights cost nothing per call beyond your own hardware.
Can I run docto-decision or gutsy locally?
Yes, both: systemone pull bofeng-huang/docto-decision and systemone pull kouhxp/gutsy download the weights.
Evaluation suite
Docto Decision Bench fr v0.1 (12 tasks, mostly silver labels; the maker's own benchmark)
JevBench public set (231 items; 169 correct), the maker's own run with Q8_0