# Totum Labs: plumbify

> Naveenraj Kamalakannan's trained decision head (a "plumb") inside open LLMs, served by vLLM. It answers choice, yes/no and score questions in one forward pass over the KV cache, with a calibrated probability per option and a conformal set. The base weights stay unchanged.

- Page: https://systemonemodels.ai/totum-labs/plumbify
- API: https://api.systemonemodels.ai/v1/models/totum-labs/plumbify
- Download: `pip install systemonemodels && systemone pull totum-labs/plumbify`

## Facts

| | |
|---|---|
| Maker | Totum Labs (https://systemonemodels.ai/totum-labs) |
| Decides | choice, score, noul |
| Architecture | plumb |
| Base model | qwen/qwen3.5-9b |
| Parameters | 9.0B |
| Licence | apache-2.0 |
| Availability | Open weights |
| Released | 2026-10-01 |
| Latest version | 2026.09 |

## Reported evaluation

Suite: Plumbify held-out decisions (447: 150 DecisionBench held-out families, 105 hard-skill templates, 192 unseen general templates), the maker's own; plumbed Qwen3.5-9B with LLM fallback, thinking off. Numbers are the publisher's own.

- Decision accuracy: 77.9%

## Model card

<!-- generated by scripts/seed_catalog.py; edit content/models/catalog.yaml -->

# Plumbify (Qwen3.5-9B)

Naveenraj Kamalakannan's trained decision head (a "plumb") inside open LLMs, served by vLLM. It answers choice, yes/no and score questions in one forward pass over the KV cache, with a calibrated probability per option and a conformal set. The base weights stay unchanged.

On a frozen base model, Plumbify trains a decision head (62M parameters on the 9B) that reads hidden states at five layers, plus a rank-32 LoRA active only while the model reads a decision, then fits a temperature and a 90% conformal threshold. Each repository holds the unchanged base weights beside the head and adapter. The model still chats; setting plumb.mode to "system1" returns the plumb's answer with nothing generated. Needs vLLM 0.30.0 with the plumbify plugin and one NVIDIA GPU; trained on English decisions only. Six plumbed models: Qwen3-1.7B, Qwen3.5-4B, Qwen3.5-9B (this entry), Gemma 4 12B, Qwen3.5-27B and Qwen3.5-35B-A3B. The maker's figures are system scores: at the default trust of 0.7 an unsure plumb answer goes back to the LLM, so they are not head-only. No calibration error is published. A personal project, unrelated to the plumb-4b model skipped on 30 September.

## What it decides

- **choice** — picks one option from a set
- **score** — places the input on an ordered scale
- **noul** — answers a yes/no question with one probability

## At a glance

| | |
|---|---|
| Parameters | 9B |
| Base model | `Qwen/Qwen3.5-9B` |
| Maker | Totum Labs |
| Released | 2026-10-01 |
| License | apache-2.0 |
| Reported accuracy | 77.9% |
| Reported latency | 1.6 s p50 per request with thinking off, including the model's reply; a single decision in the guide took 58 ms |

## Get the weights

```bash
pip install systemonemodels
systemone pull totum-labs/plumbify
```

The files are served from the maker's Hugging Face repository, [`totum-labs/Qwen3.5-9B-plumb`](https://huggingface.co/totum-labs/Qwen3.5-9B-plumb), and verified against the checksums recorded here.

## Read more

- [Model card](https://huggingface.co/totum-labs/Qwen3.5-9B-plumb)
- [Code](https://github.com/therealnaveenkamal/plumbify)
- [Qwen3.5-27B plumb](https://huggingface.co/totum-labs/Qwen3.5-27B-plumb)

---

*This page was opened by System One for Totum Labs, who can claim the organisation and take it over at any time.*

---

From System One Models — https://systemonemodels.ai/ · every System One model: https://systemonemodels.ai/system-one-models
