Prism Norwegian (prism-no)

Norwegian UPOS tagging, morphological features, and lemmatization with calibrated confidences, built for on-device, fully offline use. One compact model (17.6 M parameters) covers both written standards — Bokmål (nb) and Nynorsk (nn) — in a single set of weights; mixed input is fine. The artifact manifest also declares the BCP 47 macrolanguage no (since 0.2.3) and, since 0.2.4, the segmentation inventory the language-independent Prism runtimes (library 0.6.0+) read from the artifact. 0.2.5 re-lowers the int8 -fast programs with a grouped XNNPACK partitioner — ≈ 2× fp32 CPU speed at byte-identical weights and identical output (0.2.2–0.2.4 were manifest-only, byte-identical to 0.2.2; every published quality number still applies unchanged). Use 0.2.5 with Prism 0.6.0 or later; older runtimes keep working with any of these artifacts.

It beats UDPipe 2.17 on UPOS and lemmas on the official UD test splits — at about one twentieth of UDPipe's model size (fast artifact; one tenth for fp32), running fully offline on a laptop CPU.

This repository mirrors the versioned release artifacts of the Prism project. Prism ships native runtimes for Swift, C++, C, and Java/Kotlin that read these artifacts directly.

Which folder to use

Folder Size When to use
prism-no-0.2.5-fast/ ≈ 45 MB Recommended. int8; ≈ 2× faster, development-split quality within 0.014 pp of fp32
prism-no-0.2.5/ ≈ 94 MB Bit-exact fp32 reference behind the published benchmark

An application bundles exactly one folder. The folder is everything a Prism runtime needs; point the tagger API at its local path:

let tagger = try PrismTagger(artifactURL: artifactFolder)                // Swift
prism::tagger::Tagger tagger("prism-no-0.2.5-fast");                     // C++
try (var tagger = PrismTagger.load(Path.of("prism-no-0.2.5-fast"))) {}   // Java

Quick starts for every binding: the project README. The artifact contract (programs, model.ptd weights, tokenizer and label schemas, checksums) is documented in docs/INTEGRATION.md.

Note: these are ExecuTorch programs with the decoding policy and calibration baked in — not transformers-loadable checkpoints. vocabulary.json is a standard Hugging Face tokenizer.json and loads with the tokenizers library.

Quality

Evaluated exactly once on the untouched official UD test splits against UDPipe 2.17 (gold tokenization, official CoNLL definitions):

Test F1 Prism UDPipe 2.17
Bokmål UPOS 98.76% 98.57%
Bokmål Lemmas 98.98% 98.87%
Bokmål UFeats 97.20% 97.59%
Nynorsk UPOS 98.77% 98.60%
Nynorsk Lemmas 98.68% 98.56%
Nynorsk UFeats 96.94% 97.38%

Prism wins UPOS and lemmas on both written standards and stays behind only on exact morphology bundles — from a model a twentieth of UDPipe's size.

fast versus fp32

The frozen test evaluation above is fp32; the fast artifact is quality-gated on the development split (67,619 tokens across both standards — the test splits are evaluated exactly once and stay reserved for the fp32 benchmark). Accuracy with the identical production decoding policy:

Task Standard fp32 fast Delta
UPOS nb 99.1724% 99.1641% -0.0082 pp
UPOS nn 98.8384% 98.8448% +0.0064 pp
UFeats exact nb 97.9021% 97.8883% -0.0137 pp
UFeats exact nn 95.3408% 95.3312% -0.0096 pp
Lemma nb 99.2301% 99.2246% -0.0055 pp
Lemma nn 98.8672% 98.8608% -0.0064 pp

Every delta is at most 0.014 percentage points — an order of magnitude below seed-to-seed training variance. int8 costs no measurable quality, and 0.2.5 changes only the delegate structure, not the weights, so these numbers are unchanged.

Speed

Measured with the reproducible benchmark suite on the checked-in CC0 example texts that ship inside the Prism repository (data/examples/) — anyone can rerun the identical protocol on their own hardware. Apple M4 Max, CPU only, release builds, median of 3 repetitions. The benchmark document is the Bokmål text repeated seven times: 385 sentences / 6,335 tokens, roughly a book chapter:

C++ end-to-end, 6,335-token document fp32 fast (int8, 0.2.5)
Raw text in, tagged sentences out 3.54 s (1,853 tokens/s) 1.80 s (3,763 tokens/s)
Pretokenized input 3.56 s (1,842 tokens/s) 1.80 s (3,775 tokens/s)
Tagger load (cold start) 32 ms 35 ms

So the fast artifact tags a full chapter in under two seconds and a single ~900-token text in about half a second. The grouped int8 lowering in 0.2.5 lifts the fast artifact from ≈ 1.81× to ≈ 1.96× fp32 versus the previous per-op lowering, at identical output. The model-independent stages are effectively free next to the forward pass: runtime segmentation ≈ 8 M tokens/s, byte-level BPE ≈ 2.2–2.6 M tokens/s.

Reproduce it:

git clone https://github.com/dmlux/Prism.git
cmake -S cpp -B cpp/build -DCMAKE_BUILD_TYPE=Release -DPRISM_BENCHMARKS=ON
cmake --build cpp/build --target prism_benchmarks_norwegian
cpp/build/prism_benchmarks_norwegian --benchmark_repetitions=3 --benchmark_min_time=1x

Recorded run and protocol: docs/benchmarks/prism-no-0.2.5.md; earlier records: docs/benchmarks/prism-no-0.2.2.md.

Model description

A 17.6 M-parameter encoder student (16-layer NorBERT4-xsmall backbone, hidden 192) distilled from a NorBERT4-large teacher, with a character CNN feeding morphology and lemma heads, a structured morphology decoder, and per-head temperature calibration (UPOS ECE 0.0017). Trained on the official UD gold treebanks plus teacher-labeled silver text. Full technical reference: docs/ARCHITECTURE.md.

Training data and attribution

This model exists thanks to openly licensed Norwegian resources:

  • UD Norwegian-Bokmaal and UD Norwegian-Nynorsk treebanks (Universal Dependencies contributors, based on the Norwegian Dependency Treebank by the National Library of Norway) — CC BY-SA 4.0
  • NBdigital (sbr-43) and municipal documents (sbr-60), National Library of Norway, Språkbanken — CC0
  • Nynorsk Wikipedia, Wikimedia contributors — CC BY-SA 4.0 (text never redistributed)
  • Backbone: ltg/norbert4-xsmall; distillation teacher and silver labeler: ltg/norbert4-large (Language Technology Group, University of Oslo) — Apache 2.0

Pinned revisions and checksums travel inside each artifact (manifest.json, LICENSES/).

License

Model weights: CC BY-SA 4.0. Using or bundling the unmodified artifact — including commercially, in closed-source applications — is fine (keep the attribution); redistributed modified weights must stay open. Prism source code is Apache 2.0.

Downloads last month
42
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support