# Convai Innovations: laya-noulxp-cpu

> Unofficial NoulXP package of Convai Innovations' Laya (English) (Apache-2.0), fp32-o23: checked against the model's own answers, none changed. Not made by Convai Innovations.

- Page: https://systemonemodels.ai/convai-innovations/laya-noulxp-cpu
- API: https://api.systemonemodels.ai/v1/models/convai-innovations/laya-noulxp-cpu
- Download: `pip install systemonemodels && systemone pull convai-innovations/laya-noulxp-cpu`

## Facts

| | |
|---|---|
| Maker | Convai Innovations (https://systemonemodels.ai/convai-innovations) |
| Decides | choice, score, noul, classify, route |
| Architecture | laya |
| Base model | convai-innovations/laya |
| Parameters | 421M |
| Context | 512 tokens |
| Licence | apache-2.0 |
| Availability | Open weights |
| Latest version | 0.1.0 |

## Reported evaluation

Suite: Measured by System One Models on 2026-10-05: one question per request, 4 threads of a Xeon Gold 6342. Numbers are the publisher's own.

- Median latency: 116 ms
- p95 latency: 859 ms

## Model card

# Laya (English): NoulXP package (unofficial)

**Unofficial.** This is a NoulXP package of [Laya (English)](https://huggingface.co/convaiinnovations/laya) by [Convai Innovations](https://www.convaiinnovations.com), made by System One Models and published here, in Convai Innovations' organization, on its behalf. Convai Innovations did not make or review it. The original, its card and its own code: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya), revision `7b928d828b7b`. The original on System One Models: [convai-innovations/laya](https://systemonemodels.ai/convai-innovations/laya).

**If you made this model:** this repository is in your organization here; claim it at https://systemonemodels.ai/convai-innovations to run it yourself, or write to ceo@systemonemodels.tech and we will correct this package or take it down, the same day.

Laya (English) is the English checkpoint at the root of Convai's Laya repository: ModernBERT-large with a decision head that scores one marker per option, trained with RLCD. A NoulXP package runs it on any NoulXP runtime (`noulxp run`, `noulxp serve`, the System One Engine) with no model-specific code: nothing in it executes.

## Lineage

answerdotai/ModernBERT-large (Answer.AI, Apache-2.0) → Laya (English) (Convai Innovations, Apache-2.0) → this package (the CPU package).

## What we changed

- **Format:** ONNX graph, opset 23 (attention as ONNX's fused Attention operator), with the NoulXP files that describe the input, calibration and answers.
- **Precision:** none: float32, the published weights byte for byte.
- **Nothing else.** The input construction, calibration and answers are the model's own. `noulxp/conformance.jsonl` holds what the model's own code answers to NoulXP's 52 requests (recorded on its published weights in float32, on a CPU), and this package reproduces it.

## Our measurements (2026-10-05; the CPU test split 2026-10-07)

Measured by System One Models, not by the maker. Fidelity is measured against the model's own answers: NoulXP's 52-request conformance set and the typed-decisions test split (LocalLLaMA/typed-decisions, 400 requests, 2,000 questions).

| Where | NoulXP's 52 requests | Typed-decisions test split (2,000 questions) |
|---|---|---|
| CPU (Intel Xeon Gold 6342, x86) | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | 0 of 2000 top answers changed (ties included), max \|Δp\| 7.7e-06 (onnxruntime 1.30.0 on the CPU, against the model's own code) |

| | Result |
|---|---|
| Typed-decisions test accuracy, against the split's labels | 36.15% (the model's own, as this package reproduces it) |
| CPU latency, one question per request, 4 threads of an Intel Xeon Gold 6342 (2.8 GHz) | p50 115.87 ms, p95 859.02 ms |
| Memory of that process at its peak | 3.3 GB |

Every variant we measured. A variant is published only if no top answer changes, on the 52 requests and on the test split, and every probability stays within NoulXP's 0.01:

| Variant | Size | 52 requests, CPU | Test split | CPU p50 / p95 | A40 |
|---|---|---|---|---|---|
| fp32 | 850 MB | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | the reference: this package on the A40 at exact (float32 products) | 129.59 / 1124.44 ms | 48.58/s (exact) |
| fp32-o23 | 849 MB | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | on an NVIDIA A40 at exact (float32 products): 0 of 2000 top answers changed, max \|Δp\| 1.3e-05 | 115.87 / 859.02 ms | 55.88/s (exact) |
| fp16 | 1.01 GB | pass: 52/52 cases within 0.01, max \|Δp\| 5.0e-05, 0 top answers changed | not run on the CPU | 126.09 / 1101.43 ms | 64.22/s (fast) |
| fp16-o23 | 1.00 GB | pass: 52/52 cases within 0.01, max \|Δp\| 0.0029, 0 top answers changed | not run on the CPU | 119.85 / 903.49 ms | fails (fast: 2 changed, max 0.0068 on the test split; exact: 2 changed, max 0.0068 on the test split) |
| int8 | 657 MB | FAIL: 2/52 cases within 0.01, max \|Δp\| 0.6796, 17 top answers changed | not run (fails the 52) | 63.66 / 724.68 ms | not run (CPU variant) |
| int8-o23 | 656 MB | FAIL: 2/52 cases within 0.01, max \|Δp\| 0.9083, 26 top answers changed | not run (fails the 52) | 58.31 / 489.01 ms | not run (CPU variant) |
| int8-block | 705 MB | FAIL: 32/52 cases within 0.01, max \|Δp\| 0.0638, 2 top answers changed | not run (fails the 52) | 137.22 / 1251.35 ms | not run (CPU variant) |
| int8-block-o23 | 703 MB | FAIL: 33/52 cases within 0.01, max \|Δp\| 0.0731, 2 top answers changed | not run (fails the 52) | 129.82 / 994.55 ms | not run (CPU variant) |

## The maker's numbers

Reported by Convai Innovations, not measured by us: Convai reports 0.362 on the typed-decisions benchmark (2,000 decisions) for this base English checkpoint, near chance there, and recommends its fine-tuned typed-decisions checkpoint for that task; mean calibration error 0.466 as shipped, 0.081 after refitting temperatures.

## Where it runs on System One

On the CPU. This graph uses ONNX's fused Attention operator (opset 23), which onnxruntime runs on CUDA from version 1.30; the System One Engine's GPU image has an older onnxruntime, so the engine serves this package on the CPU until that image is updated. Elsewhere, onnxruntime-gpu 1.30 or newer runs it on an NVIDIA GPU.

## Use it

```bash
pip install systemonemodels
systemone run noulxp convai-innovations/laya-noulxp-cpu --request request.json   # on this machine
```

Or with the files of this repository's `noulxp/` folder and the reference runtime (`pip install "noulxp[onnx]"`):

```bash
noulxp check noulxp/                       # replays the conformance file on this machine
noulxp run noulxp/ --request request.json  # one request
noulxp serve noulxp/                       # POST /v1/systemone
```

## Licence

Apache-2.0, as the original. `noulxp/LICENSE` and `noulxp/NOTICE` say what the package contains and what we changed. answerdotai/ModernBERT-large is Apache-2.0-licensed.

---

From System One Models — https://systemonemodels.ai/ · every System One model: https://systemonemodels.ai/system-one-models
