Jeeves-9B
PostHog's open 9B decision model on Qwen3.5-9B. A merged LoRA and a pointer head write a reasoning chain per question, then return a calibrated probability for every option. Serves Jev's /v1/systemone contract (choice, noul, score); thinking can be switched off.
Borderline for this registry: by default Jeeves generates a reasoning chain (up to 2,560 tokens) for each question before the pointer head decides, so it is not one forward pass per question; the chain is internal unless asked for. With think=false it answers from the prompt alone in about 0.3 s and scores 0.804 on PostHog's test split, against 0.889 with thinking. Trained with LoRA SFT on 12 public datasets plus synthetic policy data, then CISPO RL, with one temperature fitted on dev. Ships block-4 and block-8 diffusion drafters for speculative decoding; the FP8 export is PostHog/jeeves-fp8. The training code and a drop-in SDK are MIT on GitHub; the weights are Apache-2.0. On the third-party JevBench board PostHog reports 0.935 on the 231 public items and 0.865 on the hard tier, its own run (sealed tier not run). The maker notes knowledge questions trail Jev (MMLU 0.793), a slow tail with full chains (17.1 s p90 on an H100) and English-only evaluation. Author: Nicholas P. Waltz.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one probability
At a glance
| Parameters | 9B |
| Base model | Qwen/Qwen3.5-9B |
| Maker | PostHog |
| Released | 2026-09-29 |
| License | apache-2.0 |
| Reported accuracy | 88.9% |
| Reported latency | 3.3 s median / 17.1 s p90 with full thinking, about 0.3 s without, on one H100 in FP8 (325 dev questions) |
Get the weights
pip install systemonemodels
systemone pull posthog/jeeves
The files are served from the maker's Hugging Face repository, PostHog/jeeves, and verified against the checksums recorded here.