System One Studio was called Laya Studio when this post was written.
Laya arrived on 18 September 2026 from Convai Innovations, three days after Jev, as the first open-weights answer to it. It is small (421M parameters), Apache-2.0, and runs on a CPU; within a week it had been ported to Core ML, MLX, GGUF, ONNX and LiteRT by other people, which tells you how much demand there was for a decision model you could actually hold.
How it decides
Laya is an encoder, not a language model. The backbone is ModernBERT-large, fully fine-tuned, with a decision head trained from scratch. The state, the question and every option are packed into one sequence, one masked marker per option, and a single bidirectional pass scores every marker at once. The softmax over those scores is the answer: for a Choice question, a distribution over the options; for Score, over the levels; for Noul, one probability.
Because every option is in the sequence, adding options costs tokens, not passes; because there is no decoding, latency is a forward pass: 39.5 ms per question on a T4 for the English checkpoint, 32.8 ms for the multilingual one, 193–464 ms on a CPU.
Training uses Convai's RLCD recipe with proper scoring rules, which is the part that makes the probabilities usable as probabilities rather than ranks.
Three checkpoints
The repository holds three models, and it matters which one you pull:
- English base (421M, 512-token context): the general model.
- Multilingual (322M on mmBERT-base, 1,024 tokens, extendable to 8,192).
- Typed-decisions: the same architecture trained on the typed-question format. On the typed-decisions suite the base checkpoint scores near chance (36.2%); this one scores 76.6%.
Reported calibration error is 0.081 after temperature fitting, 0.466 as shipped, so fit a temperature on your own data before you trust a threshold.
pip install systemonemodels
systemone pull convai-innovations/laya --variant typed-decisions
Serving
laya-serve exposes TypeSafe's /v1/systemone contract, so a client written
for Jev works against Laya by changing the base URL. The MLX port (laya-mlx)
loads the same folder on Apple silicon.
Fine-tuning it
This is where Laya is most useful. A fine-tune with LoRA takes 15–20 minutes on an M-series Mac with System One Studio: point it at a dataset of (state, question, answer) rows, hold out a test split, and it measures accuracy and calibration before and after, then writes a model card from the numbers it recorded. The snake-game policy on this registry was made that way: 98.8% held-out accuracy, expected calibration error 0.002, 36 ms per decision, up from a 15.8% base.
systemone push in a System One Studio workspace finds every run and export and
publishes them with their evaluation, so the fine-tune's page shows the same
numbers the studio measured. Laya's page lists the
fine-tunes that name it as their base.