# Convai Innovations: laya-typed-decisions-noulxp

> Unofficial NoulXP package of Convai Innovations' Laya typed-decisions (Apache-2.0), fp32-o23: checked against the model's own answers, none changed. Not made by Convai Innovations.

- Page: https://systemonemodels.ai/convai-innovations/laya-typed-decisions-noulxp
- API: https://api.systemonemodels.ai/v1/models/convai-innovations/laya-typed-decisions-noulxp
- Download: `pip install systemonemodels && systemone pull convai-innovations/laya-typed-decisions-noulxp`

## Facts

| | |
|---|---|
| Maker | Convai Innovations (https://systemonemodels.ai/convai-innovations) |
| Decides | choice, score, noul, classify, route |
| Architecture | laya |
| Base model | convai-innovations/laya |
| Parameters | 421M |
| Context | 1K tokens |
| Licence | apache-2.0 |
| Availability | Open weights |
| Latest version | 0.1.0 |

## Reported evaluation

Suite: Measured by System One Models on 2026-10-05: one question per request, 4 threads of a Xeon Gold 6342. Numbers are the publisher's own.

- Median latency: 126 ms
- p95 latency: 2185 ms

## Model card

# Laya typed-decisions: NoulXP package (unofficial)

**Unofficial.** This is a NoulXP package of [Laya typed-decisions](https://huggingface.co/convaiinnovations/laya/tree/main/typed-decisions) by [Convai Innovations](https://www.convaiinnovations.com), made by System One Models and published here, in Convai Innovations' organization, on its behalf. Convai Innovations did not make or review it. The original, its card and its own code: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya/tree/main/typed-decisions), revision `7b928d828b7b`. The original on System One Models: [convai-innovations/laya](https://systemonemodels.ai/convai-innovations/laya).

**If you made this model:** this repository is in your organization here; claim it at https://systemonemodels.ai/convai-innovations to run it yourself, or write to ceo@systemonemodels.tech and we will correct this package or take it down, the same day.

Laya typed-decisions is Convai's Laya checkpoint fine-tuned on the typed-decisions training split: ModernBERT-large with Laya's decision head, 1,024 tokens. A NoulXP package runs it on any NoulXP runtime (`noulxp run`, `noulxp serve`, the System One Engine) with no model-specific code: nothing in it executes.

## Lineage

answerdotai/ModernBERT-large (Answer.AI, Apache-2.0) → Laya typed-decisions (Convai Innovations, Apache-2.0) → this package (CPU and GPU).

## What we changed

- **Format:** ONNX graph, opset 23 (attention as ONNX's fused Attention operator), with the NoulXP files that describe the input, calibration and answers.
- **Precision:** none: float32, the published weights byte for byte.
- **Nothing else.** The input construction, calibration and answers are the model's own. `noulxp/conformance.jsonl` holds what the model's own code answers to NoulXP's 52 requests (recorded on its published weights in float32, on a CPU), and this package reproduces it.

## Our measurements (2026-10-05; the CPU test split 2026-10-07)

Measured by System One Models, not by the maker. Fidelity is measured against the model's own answers: NoulXP's 52-request conformance set and the typed-decisions test split (LocalLLaMA/typed-decisions, 400 requests, 2,000 questions).

| Where | NoulXP's 52 requests | Typed-decisions test split (2,000 questions) |
|---|---|---|
| CPU (Intel Xeon Gold 6342, x86) | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | 0 of 2000 top answers changed (ties included), max \|Δp\| 2.2e-06 (onnxruntime 1.30.0 on the CPU, against the model's own code) |
| NVIDIA A40, fast | pass: 52/52 cases within 0.01, max \|Δp\| 0.0010, 0 top answers changed | 0 of 2000 top answers changed, max \|Δp\| 0.0010 |

| | Result |
|---|---|
| Typed-decisions test accuracy, against the split's labels | 76.60% (the model's own, as this package reproduces it) |
| CPU latency, one question per request, 4 threads of an Intel Xeon Gold 6342 (2.8 GHz) | p50 126.48 ms, p95 2184.59 ms |
| Memory of that process at its peak | 3.3 GB |
| GPU throughput, one NVIDIA A40, 32 requests per pass (fast) | 82.03 decisions/s |

Every variant we measured. A variant is published only if no top answer changes, on the 52 requests and on the test split, and every probability stays within NoulXP's 0.01:

| Variant | Size | 52 requests, CPU | Test split | CPU p50 / p95 | A40 |
|---|---|---|---|---|---|
| fp32 | 850 MB | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | the reference: this package on the CPU | 138.32 / 3587.51 ms | 48.06/s (exact) |
| fp32-o23 | 849 MB | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | on an NVIDIA A40 at exact (float32 products): 0 of 2000 top answers changed, max \|Δp\| 4.2e-06 | 126.48 / 2184.59 ms | 82.03/s (fast) |
| fp16 | 1.01 GB | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | not run on the CPU | 123.93 / 3130.69 ms | fails (fast: 2 changed, max 0.0048 on the test split; exact: 2 changed, max 0.0048 on the test split) |
| fp16-o23 | 1.00 GB | pass: 52/52 cases within 0.01, max \|Δp\| 0.0032, 0 top answers changed | not run on the CPU | 131.96 / 2493.62 ms | fails (fast: 1 changed, max 0.0021 on the test split; exact: 1 changed, max 0.0021 on the test split) |
| int8 | 657 MB | FAIL: 1/52 cases within 0.01, max \|Δp\| 0.3694, 26 top answers changed | not run (fails the 52) | 70.24 / 2436.45 ms | not run (CPU variant) |
| int8-o23 | 656 MB | FAIL: 2/52 cases within 0.01, max \|Δp\| 0.3100, 28 top answers changed | not run (fails the 52) | 61.64 / 1384.88 ms | not run (CPU variant) |
| int8-block | 705 MB | FAIL: 45/52 cases within 0.01, max \|Δp\| 0.0214, 0 top answers changed | not run (fails the 52) | 135.11 / 3413.32 ms | not run (CPU variant) |
| int8-block-o23 | 703 MB | FAIL: 48/52 cases within 0.01, max \|Δp\| 0.0171, 0 top answers changed | not run (fails the 52) | 132.42 / 2408.25 ms | not run (CPU variant) |

## The maker's numbers

Reported by Convai Innovations, not measured by us: Convai reports 0.766 accuracy on the typed-decisions benchmark (2,000 decisions; noul 0.857, choice 0.733, score 0.723), above the 0.735 teacher ceiling.

## Where it runs on System One

On the CPU. This graph uses ONNX's fused Attention operator (opset 23), which onnxruntime runs on CUDA from version 1.30; the System One Engine's GPU image has an older onnxruntime, so the engine serves this package on the CPU until that image is updated. Elsewhere, onnxruntime-gpu 1.30 or newer runs it on an NVIDIA GPU.

## Use it

```bash
pip install systemonemodels
systemone run noulxp convai-innovations/laya-typed-decisions-noulxp --request request.json   # on this machine
```

Or with the files of this repository's `noulxp/` folder and the reference runtime (`pip install "noulxp[onnx]"`):

```bash
noulxp check noulxp/                       # replays the conformance file on this machine
noulxp run noulxp/ --request request.json  # one request
noulxp serve noulxp/                       # POST /v1/systemone
```

## Licence

Apache-2.0, as the original. `noulxp/LICENSE` and `noulxp/NOTICE` say what the package contains and what we changed. answerdotai/ModernBERT-large is Apache-2.0-licensed.

---

From System One Models — https://systemonemodels.ai/ · every System One model: https://systemonemodels.ai/system-one-models
