InnerJev-4B-Full
Open decision models from Shanghai Innovation Institute and Fudan University researchers: Qwen3.5-4B and Qwen3.8-27B fine-tunes, self-distilled from their own reasoning, that read option-label logits in one forward pass and give a probability per option.
Released with arXiv 2610.03935 (2 October 2026) by Feiyu Duan, Jiayu Lin and co-authors in Zhongyu Wei's group. The method, Reasoning-to-Readout Self-Distillation, has the same-size Qwen model reason twice and trains the student to give the averaged answer distribution without thinking. The prompt ends at "The correct answer is (" and the probabilities are the logits of single-token option labels at temperature 1: a label-token readout, not a dedicated head; nothing is generated. Inputs over 6,000 tokens are rejected. Text only. On the authors' JEVal benchmark the paper reports 71.92% (ECE 0.0843) for InnerJev-4B and 79.24% (ECE 0.0149) for InnerJev-27B, without saying whether these are the LoRA or full-rank checkpoints. Siblings InnerJev-4B-LoRA, 27B-Full and 27B-LoRA, all standalone BF16 models. Training code and data are public; the data keeps its upstream terms. A trial hosted API is available by request form.
What it decides
- choice — picks one option from a set
- score — places the input on an ordered scale
- noul — answers a yes/no question with one probability
At a glance
| Parameters | 5.2B |
| Base model | qwen/qwen3.5-4b |
| Maker | Jiayu Lin (Shanghai Innovation Institute) |
| Released | 2026-10-02 |
| License | apache-2.0 |
| Reported accuracy | 71.9% |
| Reported latency | 58.7 ms median per question on one H100 80 GB, one request at a time |
Get the weights
pip install systemonemodels
systemone pull jiayu-lin/innerjev
The files are served from the maker's Hugging Face repository, jylin001206/InnerJev-4B-Full, and verified against the checksums recorded here.