Wald-Q4B (Wald-4B v1.2)
A 4B decision model fine-tuned from Qwen3.5-4B-Base. A state and typed questions (choice, yes/no, score) go in; a temperature-calibrated probability for every option comes out of one pass, through a self-hosted /v1/systemone-compatible server. Apache-2.0 weights and code.
v1.2 (main since 1 October) is v1.1 plus a merged rank-16 robustness LoRA and is one-pass only. v1.1 (tag v1.1) has optional thinking efforts that generate up to 512 tokens before reading the options again; pin a revision. The shipped server needs one NVIDIA GPU with vLLM; the maker's GGUF builds run through its own wald-serve on llama.cpp, while chat front-ends return text, not probabilities. All numbers are self-run: the public JevBench items (a third-party benchmark) were used as a development scoreboard, so 204/231 is not held out; the maker's own Decision Index 0.2.1 run (54.59, v1.1 with high effort) awaits validation by that third-party board. Calibration was fitted on the maker's development data. Some training text was written by Claude Haiku, and the maker's provenance notes say part of the training data has non-commercial terms. Independent of TypeSafe; the server adapts Kev and simple-jev code under Apache-2.0. First published as Harry19081/Wald-4B.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one probability
- classify — assigns a category from a fixed taxonomy
- route — sends the input to one of several destinations
At a glance
| Parameters | 4.2B |
| Base model | Qwen/Qwen3.5-4B-Base |
| Maker | ORG2 AI |
| Released | 2026-10-01 |
| License | apache-2.0 |
| Reported accuracy | 88.3% |
| Reported latency | 33 ms p50 / 168 ms p95 per decision on one RTX PRO 6000, effort none (measured on v1.1) |
Get the weights
pip install systemonemodels
systemone pull org2ai/wald
The files are served from the maker's Hugging Face repository, org2ai/Wald-4B, and verified against the checksums recorded here.