vibert-capu, ONNX int8
The capitalisation-and-punctuation restorer used in relay in front of its Vietnamese→English translator. The recogniser emits no capitals, so a name is not visibly a name; running this first takes hard proper-noun survival in the English from 42/57 to 54/57 on 203 real recogniser clauses (twelve rescued, none broken, p=0.0005), at 19 ms per clause.
Why this exists. The upstream repository publishes pytorch_model.bin and nothing else — no
ONNX graph, no vocab.txt, no tokenizer.json. This 115 MB int8 graph was exported with
tools/capu/export_onnx.py in the app repository and has no other source. The tokenizer is a
WordPiece over FPTAI/vibert-base-cased, generated at install from that repo's vocab.txt.
Traps. normalizer.lowercase must be false — that is an override, not the default, and
if it reverts the model receives lowercased, diacritic-folded Vietnamese. onnxruntime version
changes this model's punctuation: 1.29 and 1.24.3 agree on every argmax but differ in the second
decimal of the logits, which flips marginal end-of-clause decisions. Pin the version.
$TRANSFORM_VERB_VB_VBN is not a verb form here; it converts number words to digits.
Model tree for dknguyen2304/vibert-capu-onnx
Base model
dragonSwing/vibert-capu