Jeff 1
A rank-16 LoRA on Qwen3-4B-Instruct-2507 that answers choice, noul and score questions over labels named at call time, reading probabilities from the model's own token distribution rather than generated text. Tuned for checking a claim against supplied evidence.
A PEFT adapter (11.8M trainable parameters), so the base model is needed. Choice uses first-token scoring when options start with different tokens; Score uses whole-sequence scoring, one extra pass per level. Trained on 12,119 rows for fact-checking against supplied evidence; it does not retrieve sources. On Gestalt Labs' own audit (9,730 human-labelled fact-check rows from VitaminC, FEVER, SciFact and Climate-FEVER, also used for error analysis) it scores 0.818 accuracy and ECE 0.081; the maker's own re-measurement of Jev 1.13.0 on the same rows gives 0.828 and 0.093. Weak on not-enough-info cases, and it marks refuted claims as supported more often. English only; latency not measured. The adapter is Apache-2.0, but about 7,000 training labels came from Jev 1.13.0 and 413 rows from SciFact (CC BY-NC 2.0); check both before commercial use. Ships a train-your-own guide and a local /v1/systemone server. Not affiliated with TypeSafe AI.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one probability
At a glance
| Parameters | 4B |
| Base model | Qwen/Qwen3-4B-Instruct-2507 |
| Maker | Gestalt Labs |
| Released | 2026-09-19 |
| License | apache-2.0 |
| Reported accuracy | 81.8% |
Get the weights
pip install systemonemodels
systemone pull gestalt-labs/jeff
The files are served from the maker's Hugging Face repository, GestaltLabs/Jeff-1, and verified against the checksums recorded here.