# Convai Innovations: laya-multilingual-noulxp

> Unofficial NoulXP package of Convai Innovations' Laya multilingual (Apache-2.0), fp32-o23: checked against the model's own answers, none changed. Not made by Convai Innovations.

- Page: https://systemonemodels.ai/convai-innovations/laya-multilingual-noulxp
- API: https://api.systemonemodels.ai/v1/models/convai-innovations/laya-multilingual-noulxp
- Download: `pip install systemonemodels && systemone pull convai-innovations/laya-multilingual-noulxp`

## Facts

| | |
|---|---|
| Maker | Convai Innovations (https://systemonemodels.ai/convai-innovations) |
| Decides | choice, score, noul, classify, route |
| Architecture | laya |
| Base model | convai-innovations/laya |
| Parameters | 322M |
| Context | 1K tokens |
| Licence | apache-2.0 |
| Availability | Open weights |
| Latest version | 0.1.0 |

## Reported evaluation

Suite: Measured by System One Models on 2026-10-05: one question per request, 4 threads of a Xeon Gold 6342. Numbers are the publisher's own.

- Median latency: 44.9 ms
- p95 latency: 742 ms

## Model card

# Laya multilingual: NoulXP package (unofficial)

**Unofficial.** This is a NoulXP package of [Laya multilingual](https://huggingface.co/convaiinnovations/laya/tree/main/multilingual) by [Convai Innovations](https://www.convaiinnovations.com), made by System One Models and published here, in Convai Innovations' organization, on its behalf. Convai Innovations did not make or review it. The original, its card and its own code: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya/tree/main/multilingual), revision `7b928d828b7b`. The original on System One Models: [convai-innovations/laya](https://systemonemodels.ai/convai-innovations/laya).

**If you made this model:** this repository is in your organization here; claim it at https://systemonemodels.ai/convai-innovations to run it yourself, or write to ceo@systemonemodels.tech and we will correct this package or take it down, the same day.

Laya multilingual is the multilingual checkpoint of Convai's Laya repository: mmBERT-base with Laya's decision head, for 100+ languages, 1,024 tokens as shipped. A NoulXP package runs it on any NoulXP runtime (`noulxp run`, `noulxp serve`, the System One Engine) with no model-specific code: nothing in it executes.

## Lineage

jhu-clsp/mmBERT-base (JHU CLSP, MIT) → Laya multilingual (Convai Innovations, Apache-2.0) → this package (CPU and GPU).

## What we changed

- **Format:** ONNX graph, opset 23 (attention as ONNX's fused Attention operator), with the NoulXP files that describe the input, calibration and answers.
- **Precision:** none: float32, the published weights byte for byte.
- **Nothing else.** The input construction, calibration and answers are the model's own. `noulxp/conformance.jsonl` holds what the model's own code answers to NoulXP's 52 requests (recorded on its published weights in float32, on a CPU), and this package reproduces it.

## Our measurements (2026-10-05; the CPU test split 2026-10-07)

Measured by System One Models, not by the maker. Fidelity is measured against the model's own answers: NoulXP's 52-request conformance set and the typed-decisions test split (LocalLLaMA/typed-decisions, 400 requests, 2,000 questions).

| Where | NoulXP's 52 requests | Typed-decisions test split (2,000 questions) |
|---|---|---|
| CPU (Intel Xeon Gold 6342, x86) | pass: 52/52 cases within 0.01, max \|Δp\| 5.1e-05, 0 top answers changed | 0 of 2000 top answers changed (ties included), max \|Δp\| 7.2e-06 (onnxruntime 1.30.0 on the CPU, against the model's own code) |
| NVIDIA A40, exact | pass: 52/52 cases within 0.01, max \|Δp\| 5.1e-05, 0 top answers changed | 0 of 2000 top answers changed, max \|Δp\| 6.2e-06 |

| | Result |
|---|---|
| Typed-decisions test accuracy, against the split's labels | 35.20% (the model's own, as this package reproduces it) |
| CPU latency, one question per request, 4 threads of an Intel Xeon Gold 6342 (2.8 GHz) | p50 44.94 ms, p95 741.74 ms |
| Memory of that process at its peak | 2.3 GB |
| GPU throughput, one NVIDIA A40, 32 requests per pass (exact) | 117.71 decisions/s |

Every variant we measured. A variant is published only if no top answer changes, on the 52 requests and on the test split, and every probability stays within NoulXP's 0.01:

| Variant | Size | 52 requests, CPU | Test split | CPU p50 / p95 | A40 |
|---|---|---|---|---|---|
| fp32 | 681 MB | pass: 52/52 cases within 0.01, max \|Δp\| 5.1e-05, 0 top answers changed | the reference: this package on the A40 at exact (float32 products) | 46.32 / 1085.42 ms | 99.81/s (exact) |
| fp32-o23 | 680 MB | pass: 52/52 cases within 0.01, max \|Δp\| 5.1e-05, 0 top answers changed | on an NVIDIA A40 at exact (float32 products): 0 of 2000 top answers changed, max \|Δp\| 6.2e-06 | 44.94 / 741.74 ms | 117.71/s (exact) |
| fp16 | 1.10 GB | pass: 52/52 cases within 0.01, max \|Δp\| 5.1e-05, 0 top answers changed | not run on the CPU | 50.69 / 1140.21 ms | fails (fast: 3 changed, max 0.0069 on the test split; exact: 3 changed, max 0.0069 on the test split) |
| fp16-o23 | 1.10 GB | pass: 52/52 cases within 0.01, max \|Δp\| 0.0018, 0 top answers changed | not run on the CPU | 46.94 / 908.84 ms | fails (fast: 2 changed, max 0.0075 on the test split; exact: 2 changed, max 0.0075 on the test split) |
| int8 | 990 MB | FAIL: 3/52 cases within 0.01, max \|Δp\| 0.8095, 24 top answers changed | not run (fails the 52) | 32.07 / 901.78 ms | not run (CPU variant) |
| int8-o23 | 989 MB | FAIL: 3/52 cases within 0.01, max \|Δp\| 0.8939, 24 top answers changed | not run (fails the 52) | 26.0 / 507.22 ms | not run (CPU variant) |
| int8-block | 1.01 GB | FAIL: 22/52 cases within 0.01, max \|Δp\| 0.1177, 1 top answers changed | not run (fails the 52) | 50.26 / 1183.4 ms | not run (CPU variant) |
| int8-block-o23 | 1.01 GB | FAIL: 20/52 cases within 0.01, max \|Δp\| 0.1268, 2 top answers changed | not run (fails the 52) | 45.5 / 817.44 ms | not run (CPU variant) |

## The maker's numbers

Reported by Convai Innovations, not measured by us: Convai reports 0.342 on the typed-decisions benchmark for this checkpoint (its table), mean calibration error 0.314 as shipped and 0.106 after refitting temperatures, and about 2.2x the speed of the English checkpoint.

## Where it runs on System One

On the CPU. This graph uses ONNX's fused Attention operator (opset 23), which onnxruntime runs on CUDA from version 1.30; the System One Engine's GPU image has an older onnxruntime, so the engine serves this package on the CPU until that image is updated. Elsewhere, onnxruntime-gpu 1.30 or newer runs it on an NVIDIA GPU.

## Use it

```bash
pip install systemonemodels
systemone run noulxp convai-innovations/laya-multilingual-noulxp --request request.json   # on this machine
```

Or with the files of this repository's `noulxp/` folder and the reference runtime (`pip install "noulxp[onnx]"`):

```bash
noulxp check noulxp/                       # replays the conformance file on this machine
noulxp run noulxp/ --request request.json  # one request
noulxp serve noulxp/                       # POST /v1/systemone
```

On an NVIDIA GPU, serve it at `--precision exact` (float32 products, TF32 off): that is the setting measured above.

## Licence

Apache-2.0, as the original. `noulxp/LICENSE` and `noulxp/NOTICE` say what the package contains and what we changed. jhu-clsp/mmBERT-base is MIT-licensed.

---

From System One Models — https://systemonemodels.ai/ · every System One model: https://systemonemodels.ai/system-one-models
