Plumbify (Qwen3.5-9B)
Naveenraj Kamalakannan's trained decision head (a "plumb") inside open LLMs, served by vLLM. It answers choice, yes/no and score questions in one forward pass over the KV cache, with a calibrated probability per option and a conformal set. The base weights stay unchanged.
On a frozen base model, Plumbify trains a decision head (62M parameters on the 9B) that reads hidden states at five layers, plus a rank-32 LoRA active only while the model reads a decision, then fits a temperature and a 90% conformal threshold. Each repository holds the unchanged base weights beside the head and adapter. The model still chats; setting plumb.mode to "system1" returns the plumb's answer with nothing generated. Needs vLLM 0.30.0 with the plumbify plugin and one NVIDIA GPU; trained on English decisions only. Six plumbed models: Qwen3-1.7B, Qwen3.5-4B, Qwen3.5-9B (this entry), Gemma 4 12B, Qwen3.5-27B and Qwen3.5-35B-A3B. The maker's figures are system scores: at the default trust of 0.7 an unsure plumb answer goes back to the LLM, so they are not head-only. No calibration error is published. A personal project, unrelated to the plumb-4b model skipped on 30 September.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one probability
At a glance
| Parameters | 9B |
| Base model | Qwen/Qwen3.5-9B |
| Maker | Totum Labs |
| Released | 2026-10-01 |
| License | apache-2.0 |
| Reported accuracy | 77.9% |
| Reported latency | 1.6 s p50 per request with thinking off, including the model's reply; a single decision in the guide took 58 ms |
Get the weights
pip install systemonemodels
systemone pull totum-labs/plumbify
The files are served from the maker's Hugging Face repository, totum-labs/Qwen3.5-9B-plumb, and verified against the checksums recorded here.