Valen-4B
A multimodal decision model from the Valen team: Qwen3.5-4B plus a two-layer MLP-Mixer decision head that returns probability distributions for Choice, Noul and Score questions about a shared state of text, images or video, without generating answer text.
Trained in two SFT stages (decision-head warm-up, then joint LoRA, mixer and visual-merger training); videos are sampled at 16 frames. Several questions can share one state or be scored separately. Loading needs trust_remote_code. On its own VisualDecisionBench (released alongside) the maker reports 83.06% image and 87.36% video accuracy for the 4B; its own run on JevBench's public items (a third-party benchmark) gives 83.12%. The card notes that the released merged weights can differ slightly in probability from the training checkpoints behind those numbers; no calibration figures are published. Siblings Valen-2B and 0.8B, and an earlier Valen-Preview-0923. Base models and source datasets keep their own licences.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one probability
At a glance
| Parameters | 4.5B |
| Base model | qwen/qwen3.5-4b |
| Maker | Valen Team |
| Released | 2026-10-06 |
| License | apache-2.0 |
| Reported accuracy | 83.1% |
Get the weights
pip install systemonemodels
systemone pull valen-team/valen
The files are served from the maker's Hugging Face repository, Valen-Team/Valen-4B, and verified against the checksums recorded here.
Read more
This page was opened by System One for Valen Team, who can claim the organisation and take it over at any time.