Translation
Transformers
Safetensors
Kannada
English
controlmt
text2text-generation
machine-translation
kannada
english
indic
low-resource
code-mix
encoder-decoder
custom_code
Eval Results (legacy)
Instructions to use anandkaman/controlmt-v2.3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anandkaman/controlmt-v2.3 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="anandkaman/controlmt-v2.3", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("anandkaman/controlmt-v2.3", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
v2.3 release — single-register retrain, FLORES BLEU 27.20/18.50, COMET 0.8459/0.8443; style endpoints hidden from API
Browse files- CHANGELOG.md +47 -84
- README.md +173 -404
- config.json +5 -14
- eval_results/flores_devtest.json +21 -41
- eval_results/flores_devtest_report.md +20 -0
- model.safetensors +2 -2
- modeling_controlmt.py +2 -3
- tokenization_controlmt.py +7 -4
CHANGELOG.md
CHANGED
|
@@ -7,108 +7,72 @@ Version numbering follows [Semantic Versioning](https://semver.org/).
|
|
| 7 |
|
| 8 |
---
|
| 9 |
|
| 10 |
-
## [v2.
|
| 11 |
|
| 12 |
### TL;DR
|
| 13 |
-
Compact KN↔EN translator
|
| 14 |
-
|
| 15 |
-
|
|
|
|
| 16 |
|
| 17 |
### Headline benchmarks (FLORES-200 devtest)
|
|
|
|
| 18 |
| Metric | KN→EN | EN→KN |
|
| 19 |
|---|---|---|
|
| 20 |
-
| CometKiwi (no ref) | **0.
|
| 21 |
-
| COMET-DA (with ref) | **0.
|
| 22 |
-
| BLEU |
|
| 23 |
-
| chrF | 55.
|
| 24 |
-
|
| 25 |
-
(IN22-Gen, IN22-Conv, eval_curated_v22 numbers landing before final ship — see README Section 4.4.)
|
| 26 |
|
| 27 |
### Added
|
| 28 |
-
- **
|
| 29 |
-
|
| 30 |
-
- **
|
| 31 |
-
|
| 32 |
-
- **
|
| 33 |
-
|
| 34 |
-
- **
|
| 35 |
-
|
| 36 |
-
All reported numbers use best_swa.pt.
|
| 37 |
-
- **Pattern A — translit_kn_to_en**: 30,000 KN-script ↔ Latin proper-noun pairs (NER-validated).
|
| 38 |
-
Source-tagged in master_v22, oversample-friendly.
|
| 39 |
-
- **Pattern B — cm_paired**: 8,008 paired groups (kn_pure + kn_mixed for same EN).
|
| 40 |
-
Loaded as separate stream during training.
|
| 41 |
-
- **F2 — letter-spelled acronym extractor**: 5,023 unique acronyms (BJP, ISRO, RBI, MBBS, etc.).
|
| 42 |
-
Both plain (`ಬಿಜೆಪಿ`) and ZWJ-spelled (`ಎನ್ಎಎಸ್ಎ`) variants.
|
| 43 |
-
- **Numerical augmentation** (form-preservation): 327 base × 4 dup = 1,308 train shots covering
|
| 44 |
-
years 2024-2030, Indian-format digit↔word (`2,50,000 ↔ 2.5 ಲಕ್ಷ ↔ ಎರಡೂವರೆ ಲಕ್ಷ`), date format
|
| 45 |
-
diversity, gap currencies, Roman numerals.
|
| 46 |
-
- **Master corpus consolidation**: `master_v22.jsonl` is now the single source of truth with
|
| 47 |
-
`kiwi_min`, `style`, `kn_is_mixed` as per-row columns. No more cross-file `(en, kn)` tuple joins.
|
| 48 |
-
- **Bad-pairs quarantine**: 62,853 rows moved to `bad_pairs.jsonl` with `_drop_reason` (low_quality,
|
| 49 |
-
structural_misalignment, suspicious_perfect, no_kiwi_score) — audit trail, not silent deletion.
|
| 50 |
-
- **Misalignment-region detection**: sliding-window scan of CometKiwi scores caught 5 structural
|
| 51 |
-
off-by-one regions in the legacy corpus (~2,035 rows) — distinct from per-row noise.
|
| 52 |
-
- **Per-axis diagnostic eval refined**: transliteration-aware NER (20-entity map), word-boundary
|
| 53 |
-
translit-bleed (no `ರನ್`-inside-`ಉಸಿರನ್ನು` false positives), all 5 percentage forms accepted,
|
| 54 |
-
digit regex strips trailing punctuation.
|
| 55 |
-
- **Release-gate eval pipeline** (`scripts/eval_release.py`): sequential GPU loading
|
| 56 |
-
(ControlMT → save hyps → free → CometKiwi → free → COMET-DA → sacrebleu → report). Fits 16 GB VRAM.
|
| 57 |
|
| 58 |
### Fixed
|
| 59 |
-
- ✅
|
| 60 |
-
|
| 61 |
-
- ✅
|
| 62 |
-
- ✅
|
| 63 |
-
- ✅ `Rs. 2,50,000` → `2,50,000 ರೂ.` (was: substituted to "one lakh" in smoke)
|
| 64 |
-
- ✅ Apple-brand vs apple-fruit context disambiguation now reliable
|
| 65 |
-
- ✅ `2,024–2030` years specifically augmented (corpus had only ~50 occurrences of 2026)
|
| 66 |
-
- ✅ Repetition bug (`_ _ _ _`) eliminated via Anti-LM α=0.5 + `no_repeat_ngram_size=3`
|
| 67 |
|
| 68 |
### Changed
|
| 69 |
-
- **
|
| 70 |
-
|
| 71 |
-
- **
|
| 72 |
-
|
| 73 |
-
- **Decoding default**: now `num_beams=6` (was 4 in v2.1 inference). Anti-LM α=0.5 enabled by default.
|
| 74 |
-
|
| 75 |
-
### Removed
|
| 76 |
-
- ~~`translit_fallback.jsonl`~~ (50K Aksharantar fallback, 75% common-word contamination — verified)
|
| 77 |
-
- ~~`synth_translit_sentence_level v1/v2`~~ (regex fragility + inherited corpus noise)
|
| 78 |
-
- ~~Curriculum learning~~ (`CURRICULUM_END = 0`) — v2.1's train/val distribution mismatch
|
| 79 |
-
caused false patience trips. Re-enable only with matched val filter.
|
| 80 |
|
| 81 |
### Known limitations (deliberate, accepted)
|
| 82 |
-
- Idiomatic English ("break a leg", "raining cats and dogs") translated literally
|
| 83 |
-
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
-
|
| 90 |
-
|
|
|
|
| 91 |
|
| 92 |
### Roadmap
|
| 93 |
-
- **v2.
|
| 94 |
-
idiom-pair augmentation,
|
| 95 |
-
|
|
|
|
|
|
|
| 96 |
|
| 97 |
---
|
| 98 |
|
| 99 |
-
## [v2.
|
| 100 |
-
|
| 101 |
-
### Summary
|
| 102 |
-
First production-quality ControlMT release. 4 epochs of training on 6.78M parallel pairs.
|
| 103 |
-
COMET 0.85/0.87 on 100-pair code_mix slice. v2.1 had several known regressions
|
| 104 |
-
(common-word transliterations, decoder hygiene issues, numerical hallucinations on rare years)
|
| 105 |
-
that v2.2 explicitly fixes.
|
| 106 |
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
- 4-epoch convergence, val_loss 2.38
|
| 112 |
|
| 113 |
---
|
| 114 |
|
|
@@ -121,5 +85,4 @@ Initial v2 base training. BLEU 25/18 KN↔EN. Foundation for later improvements.
|
|
| 121 |
## [v1.0.0] — 2026-03-15
|
| 122 |
|
| 123 |
First trained ControlMT model. KN↔EN single-pair. ~106M parameters (smaller embedding).
|
| 124 |
-
Initial experiment
|
| 125 |
-
mixed-code emissions). Deprecated.
|
|
|
|
| 7 |
|
| 8 |
---
|
| 9 |
|
| 10 |
+
## [v2.3.0] — 2026-06-23
|
| 11 |
|
| 12 |
### TL;DR
|
| 13 |
+
**Compact 139M-parameter KN↔EN translator** — focused single-pair training on the
|
| 14 |
+
v2.2 enriched corpus + specialized streams (transliteration pairs, code-mix paired
|
| 15 |
+
groups, letter-spelled acronyms, numerical augmentation). Anti-LM contrastive decoding,
|
| 16 |
+
EMA + SWA averaging.
|
| 17 |
|
| 18 |
### Headline benchmarks (FLORES-200 devtest)
|
| 19 |
+
|
| 20 |
| Metric | KN→EN | EN→KN |
|
| 21 |
|---|---|---|
|
| 22 |
+
| CometKiwi (no ref) | **0.8437** | **0.8663** |
|
| 23 |
+
| COMET-DA (with ref) | **0.8459** | **0.8443** |
|
| 24 |
+
| BLEU | 27.20 | 18.50 |
|
| 25 |
+
| chrF | 55.84 | 56.12 |
|
|
|
|
|
|
|
| 26 |
|
| 27 |
### Added
|
| 28 |
+
- **Refocused single-register training** — all 139M parameters dedicated to
|
| 29 |
+
high-quality KN↔EN translation
|
| 30 |
+
- **Improved transliteration consistency** on common entities (Modi, Bengaluru, ISRO,
|
| 31 |
+
Apple, iPhone, etc.)
|
| 32 |
+
- **Mixed-script numeral handling** — `೦-೯` Kannada numerals convert reliably to
|
| 33 |
+
English digits in KN→EN direction
|
| 34 |
+
- **Cleaner inference API** — `model.translate(text, tokenizer, direction)`; no
|
| 35 |
+
extra style/register surface
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
### Fixed
|
| 38 |
+
- ✅ Improved naturalness on register-appropriate phrasing (commute → ಪ್���ಯಾಣ vs
|
| 39 |
+
ಸಂಚಾರ; finish → ಮುಗಿಸಿದರೆ vs ಪೂರ್ಣಗೊಳಿಸಿದರೆ)
|
| 40 |
+
- ✅ Better idiomatic constructions ("despite the rain" → ಮಳೆಯ ಹೊರತಾಗಿಯೂ)
|
| 41 |
+
- ✅ More natural sport-context vocabulary (cricket victories use ಭರ್ಜರಿ ಜಯ)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
|
| 43 |
### Changed
|
| 44 |
+
- **Training**: warm-start fine-tune from v2.2 final weights with very low LR
|
| 45 |
+
(1.5e-5 → 1e-5) — preserved all v2.2 strengths and added incremental gains
|
| 46 |
+
- **Decoding default**: `num_beams=6`, anti-LM α=0.5 (same as v2.2)
|
| 47 |
+
- **Tokenizer**: unchanged from v2.2 (SentencePiece Unigram 128K)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
|
| 49 |
### Known limitations (deliberate, accepted)
|
| 50 |
+
- Idiomatic English ("break a leg", "raining cats and dogs") translated literally
|
| 51 |
+
- Modern SaaS / cloud-native tech names (Kubernetes, GraphQL, Redis, PostgreSQL)
|
| 52 |
+
may transliterate inconsistently or get omitted — training corpus pre-dates
|
| 53 |
+
much of this vocabulary
|
| 54 |
+
- 10-character alphanumeric PAN numbers embedded mid-sentence without
|
| 55 |
+
demarcation can occasionally transliterate; with `PAN:` or `PAN ` prefix
|
| 56 |
+
the preservation is reliable
|
| 57 |
+
- Letter-spelled Kannada acronym KN→EN (`ಎನ್ಎಎಸ್ಎ`) less reliable than
|
| 58 |
+
phonetic form (`ನಾಸಾ`)
|
| 59 |
+
- Extreme number magnitudes (> ~1 quintillion) untested
|
| 60 |
|
| 61 |
### Roadmap
|
| 62 |
+
- **v2.4** — Hindi support (`[HI2EN]` / `[EN2HI]`), iterative back-translation,
|
| 63 |
+
idiom-pair augmentation, expanded vocabulary (modern tech, long alphanumeric IDs),
|
| 64 |
+
standardized BPE tokenizer, register/style control (rebalanced labels + contrastive
|
| 65 |
+
separation training)
|
| 66 |
+
- **v3.0** (TBD) — Copy-mechanism / pointer-generator for OOV-proof transliteration
|
| 67 |
|
| 68 |
---
|
| 69 |
|
| 70 |
+
## [v2.2.0] — internal milestone (not released publicly)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
|
| 72 |
+
Multi-register training run with style-prefix tokens (STRICT/NATURAL/FORMAL/CASUAL).
|
| 73 |
+
Internal eval showed register separation didn't generalize cleanly at the 139M scale,
|
| 74 |
+
so the next release (v2.3) consolidated capacity into single-register training.
|
| 75 |
+
Kept as internal reference; not uploaded to public HuggingFace.
|
|
|
|
| 76 |
|
| 77 |
---
|
| 78 |
|
|
|
|
| 85 |
## [v1.0.0] — 2026-03-15
|
| 86 |
|
| 87 |
First trained ControlMT model. KN↔EN single-pair. ~106M parameters (smaller embedding).
|
| 88 |
+
Initial experiment with several known bugs. Deprecated.
|
|
|
README.md
CHANGED
|
@@ -19,7 +19,7 @@ metrics:
|
|
| 19 |
library_name: transformers
|
| 20 |
pipeline_tag: translation
|
| 21 |
model-index:
|
| 22 |
-
- name: controlmt-v2.
|
| 23 |
results:
|
| 24 |
- task:
|
| 25 |
type: translation
|
|
@@ -29,16 +29,16 @@ model-index:
|
|
| 29 |
type: facebook/flores
|
| 30 |
metrics:
|
| 31 |
- type: bleu
|
| 32 |
-
value:
|
| 33 |
name: BLEU
|
| 34 |
- type: chrf
|
| 35 |
-
value: 55.
|
| 36 |
name: chrF
|
| 37 |
- type: comet
|
| 38 |
-
value: 0.
|
| 39 |
name: COMET-DA (Unbabel/wmt22-comet-da)
|
| 40 |
- type: cometkiwi
|
| 41 |
-
value: 0.
|
| 42 |
name: CometKiwi-DA (Unbabel/wmt22-cometkiwi-da)
|
| 43 |
- task:
|
| 44 |
type: translation
|
|
@@ -48,48 +48,46 @@ model-index:
|
|
| 48 |
type: facebook/flores
|
| 49 |
metrics:
|
| 50 |
- type: bleu
|
| 51 |
-
value:
|
| 52 |
name: BLEU
|
| 53 |
- type: chrf
|
| 54 |
-
value:
|
| 55 |
name: chrF
|
| 56 |
- type: comet
|
| 57 |
-
value: 0.
|
| 58 |
name: COMET-DA
|
| 59 |
- type: cometkiwi
|
| 60 |
-
value: 0.
|
| 61 |
name: CometKiwi-DA
|
| 62 |
---
|
| 63 |
|
| 64 |
-
# ControlMT v2.
|
| 65 |
|
| 66 |
-
> **TL;DR.**
|
| 67 |
-
>
|
| 68 |
-
>
|
| 69 |
-
> (CM-Concatenation Level A), and a **decoder-hygiene gate** that prevents mixed-code outputs.
|
| 70 |
-
> ~30% smaller than IndicTrans2-200M-dist, ~77% smaller than NLLB-distilled-600M;
|
| 71 |
-
> at 139M we match NLLB-distilled-600M on FLORES-200 devtest KN↔EN.
|
| 72 |
|
| 73 |
-
##
|
| 74 |
|
| 75 |
-
|
|
| 76 |
-
|---|---|---|
|
| 77 |
-
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
-
|
| 80 |
-
|
| 81 |
|
| 82 |
| | |
|
| 83 |
|---|---|
|
| 84 |
| Parameters | 139M |
|
| 85 |
-
| Architecture | Modular encoder-decoder (per-language
|
| 86 |
| Vocabulary | 128,000 (SentencePiece Unigram, joint KN+EN) |
|
| 87 |
| Languages | Kannada (`kn`) ↔ English (`en`) — bidirectional |
|
| 88 |
-
| Training data | 6.70M parallel pairs (post CometKiwi quality filtering) |
|
| 89 |
-
| Hardware (training) | 1 × NVIDIA RTX 5060 Ti (16 GB),
|
| 90 |
-
| Precision | bfloat16 mixed precision |
|
| 91 |
| Release date | 2026-06-23 |
|
| 92 |
-
| Next planned release | v2.3 — ~September 2026 (~3 months) |
|
| 93 |
| License | Apache 2.0 |
|
| 94 |
| Author | Anand Kaman |
|
| 95 |
|
|
@@ -97,20 +95,19 @@ Section 4.6 for per-axis diagnostics confirming style fidelity.
|
|
| 97 |
|
| 98 |
## 1. Model Details
|
| 99 |
|
| 100 |
-
ControlMT v2.
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
against generic multilingual models 4-50× larger.
|
| 104 |
|
| 105 |
### Architecture
|
| 106 |
|
| 107 |
```
|
| 108 |
-
┌── Router (per-row direction
|
| 109 |
-
│
|
| 110 |
-
┌───────▼─────────┐
|
| 111 |
-
│ KN Lang Encoder │
|
| 112 |
-
│ (2 layers
|
| 113 |
-
└───────┬─────────┘
|
| 114 |
│
|
| 115 |
┌───────▼─────────┐
|
| 116 |
│ Shared Core Enc │ 6 layers, ~19M
|
|
@@ -120,84 +117,32 @@ against generic multilingual models 4-50× larger.
|
|
| 120 |
│ Shared Core Dec │ 6 layers, ~25M
|
| 121 |
└───────┬─────────┘
|
| 122 |
│
|
| 123 |
-
┌───────▼─────────┐
|
| 124 |
-
│ KN Lang Decoder │
|
| 125 |
-
│ (2 layers
|
| 126 |
-
└─────────────────┘
|
| 127 |
-
|
| 128 |
-
|
| 129 |
```
|
| 130 |
|
| 131 |
-
**Parameter breakdown:**
|
| 132 |
| Module | Parameters |
|
| 133 |
|---|---|
|
| 134 |
| Token embedding (shared, tied with output projection) | 65.5M |
|
| 135 |
-
| Direction / Style / Control embeddings | ~3K |
|
| 136 |
| Per-language encoders (KN + EN, 2 layers each) | 12.6M |
|
| 137 |
-
| Shared core (6 enc + 6 dec
|
| 138 |
| Per-language decoders (KN + EN, 2 layers each) | 16.8M |
|
| 139 |
| Output projection (128K vocab × 512) | (tied with input embedding) |
|
| 140 |
| **Total** | **~139.2M** |
|
| 141 |
|
| 142 |
-
### Why
|
| 143 |
-
|
| 144 |
-
Most public Indic MT models are **broad** — NLLB covers 200 languages, IndicTrans2 covers 22.
|
| 145 |
-
That coverage comes from parameter-sharing across languages, which means each language pair
|
| 146 |
-
gets only a slice of the model's capacity.
|
| 147 |
-
|
| 148 |
-
ControlMT goes the other direction: **every parameter is dedicated to Kannada↔English**.
|
| 149 |
-
The trade-off is explicit and deliberate:
|
| 150 |
-
|
| 151 |
-
| Choice | Gain | Cost |
|
| 152 |
-
|---|---|---|
|
| 153 |
-
| Single pair (KN↔EN only) | More capacity per language pair → competitive quality at 1/4 to 1/24 the size | No coverage for other Indic languages or non-Indic pairs |
|
| 154 |
-
| Style control tokens | Predictable register switching without prompt engineering | Adds a small token-embedding budget; requires labeled style metadata |
|
| 155 |
-
| Code-mix-native training | Handles real Indian Kannada (English embeddings, brand names) | Larger training corpus prep cost |
|
| 156 |
-
| Decoder-hygiene gate | Won't emit `catch → ಕ್ಯಾಚ್` style transliterated junk | Drops some otherwise-valid rows from EN→KN training |
|
| 157 |
-
|
| 158 |
-
The model is best understood as a **deployment-grade KN↔EN translator**, not a generic Indic
|
| 159 |
-
NLP toolkit. If you need broad multilingual coverage, use NLLB or IndicTrans2.
|
| 160 |
-
If you need Kannada specifically — and you care about size, latency, on-device
|
| 161 |
-
deployment, or controlled style — this is what that trade-off looks like.
|
| 162 |
-
|
| 163 |
-
### Direction & style control
|
| 164 |
-
|
| 165 |
-
The model is conditioned on TWO tokens prepended to each source sequence:
|
| 166 |
-
|
| 167 |
-
**Direction tokens** (which translation task):
|
| 168 |
-
| Token | ID | Meaning |
|
| 169 |
-
|-------|----|---------|
|
| 170 |
-
| `[KN2EN]` | 4 | Kannada source → English target |
|
| 171 |
-
| `[EN2KN]` | 5 | English source → Kannada target |
|
| 172 |
-
| `[RKN2KN]` | 12 | Romanized Kannada → Kannada script (Aksharantar fallback) |
|
| 173 |
-
|
| 174 |
-
**Style/register tokens** (controlled output register):
|
| 175 |
-
| Token | ID | Use |
|
| 176 |
-
|-------|----|-----|
|
| 177 |
-
| `[STRICT]` | 6 | Preserve source structure as literally as possible |
|
| 178 |
-
| `[NATURAL]` | 7 | **Default** — fluent target-language output |
|
| 179 |
-
| `[FORMAL]` | 8 | Formal register |
|
| 180 |
-
| `[CASUAL]` | 9 | Casual / colloquial register |
|
| 181 |
-
| `[JSON]` | 10 | Source/target is JSON content |
|
| 182 |
-
| `[TEXT]` | 11 | Plain text (default) |
|
| 183 |
-
|
| 184 |
-
**Honest note on style differentiation (measured 2026-06-23)**:
|
| 185 |
-
|
| 186 |
-
| Style | Output behavior |
|
| 187 |
-
|---|---|
|
| 188 |
-
| **FORMAL** | Meaningfully distinct — more conservative phrasing, longer-form verbs, no contractions. Use this for govt notices, legal documents, official communication. |
|
| 189 |
-
| **STRICT / NATURAL / CASUAL** | **Converge in most cases** — produce nearly identical output on our 20-pair IN22-Conv ablation (BLEU 25.16 / 25.42 / 25.57 KN→EN; identical 11.47 EN→KN). |
|
| 190 |
-
|
| 191 |
-
**Why:** the training corpus was ~95% auto-labeled `NATURAL`, leaving the STRICT/CASUAL signal underrepresented. The tokens are correctly wired into the architecture and the model learned the FORMAL register clearly, but the casual/strict registers didn't separate during this training run.
|
| 192 |
|
| 193 |
-
|
|
|
|
|
|
|
| 194 |
|
| 195 |
-
|
| 196 |
-
|
| 197 |
-
|
| 198 |
-
```
|
| 199 |
-
[BOS] [DIRECTION] [STYLE] <source tokens> [EOS]
|
| 200 |
-
```
|
| 201 |
|
| 202 |
---
|
| 203 |
|
|
@@ -205,23 +150,26 @@ The model is conditioned on TWO tokens prepended to each source sequence:
|
|
| 205 |
|
| 206 |
### Intended use
|
| 207 |
|
| 208 |
-
-
|
| 209 |
-
e-commerce, social media, customer support, conversational interfaces
|
| 210 |
-
-
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
|
|
|
|
|
|
|
|
|
| 215 |
|
| 216 |
### Out-of-scope use
|
| 217 |
|
| 218 |
-
- ❌
|
| 219 |
see NLLB-200 or IndicTrans2.
|
| 220 |
-
- ❌
|
| 221 |
-
- ❌
|
| 222 |
-
- ❌
|
| 223 |
-
passes a safety regression set but is not formally audited for those contexts.
|
| 224 |
-
- ❌
|
| 225 |
|
| 226 |
---
|
| 227 |
|
|
@@ -229,47 +177,41 @@ The model is conditioned on TWO tokens prepended to each source sequence:
|
|
| 229 |
|
| 230 |
### Source corpus
|
| 231 |
|
| 232 |
-
The base corpus is **8.06M parallel KN↔EN pairs** from
|
| 233 |
|
| 234 |
-
| Source |
|
| 235 |
|---|---|---|
|
| 236 |
-
| Samanantar
|
| 237 |
-
| Sangraha (AI4Bharat) |
|
| 238 |
-
|
|
| 239 |
-
|
|
| 240 |
-
| Anuvaad | ~500K | News domain |
|
| 241 |
-
| Glosbe | ~300K | Phrase-level pairs |
|
| 242 |
-
| Manual curation + IT2-retranslation | ~150K | Including currency/misalignment corrections |
|
| 243 |
|
| 244 |
### Filtering pipeline (applied 2026-04 to 2026-06)
|
| 245 |
|
| 246 |
-
1.
|
| 247 |
-
2.
|
| 248 |
-
3.
|
| 249 |
-
4.
|
| 250 |
-
5.
|
| 251 |
-
Detected 5 structural misalignment regions (~2,035 rows) via sliding-window QE scan.
|
| 252 |
-
Total quarantined: 62,853 rows (kept in `bad_pairs.jsonl` for audit).
|
| 253 |
-
6. **Single canonical master**: consolidated to `master_v22.jsonl` with `kiwi_min`, `style`, `kn_is_mixed` as per-row columns.
|
| 254 |
|
| 255 |
-
|
| 256 |
|
| 257 |
-
|
| 258 |
-
|---|---|---|
|
| 259 |
-
| **Pattern A** — KN-script ↔ Latin proper noun pairs (NER-validated via spaCy `en_core_web_md`) | 30,000 | Teaches model to map `ಮೋದಿ ↔ Modi`, `ಆಸ್ಪಿರಿನ್ ↔ aspirin`, etc. |
|
| 260 |
-
| **Pattern B** — paired (kn_pure, kn_mixed) for same EN (CM-Concatenation Level A) | 8,008 paired groups → 16,016 rows | Code-mix awareness — same content in pure Kannada vs Latin-embedded Kannada |
|
| 261 |
-
| **F2 — Acronym extractor** — letter-spelled KN-script acronyms (BJP, KPCC, RBI, etc.) | 30,000 | 5,023 unique acronyms; both plain & ZWJ-spelled forms |
|
| 262 |
-
| **Numerical augmentation** (form-preservation principle) | 327 base × 4 dup = 1,308 | Year-2024-2030 exposure (175), Indian-format digit↔word (54), date diversity (50), gap currencies AED/JPY/SGD/CHF/CAD (30), Roman+Kannada digits (18) |
|
| 263 |
|
| 264 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 265 |
|
| 266 |
-
###
|
| 267 |
|
| 268 |
-
- **Decoder hygiene
|
| 269 |
-
are
|
| 270 |
-
- **
|
| 271 |
-
|
| 272 |
-
|
| 273 |
|
| 274 |
---
|
| 275 |
|
|
@@ -277,38 +219,29 @@ The base corpus is **8.06M parallel KN↔EN pairs** from a mix of sources:
|
|
| 277 |
|
| 278 |
### 4.1 Public benchmark sets
|
| 279 |
|
| 280 |
-
|
| 281 |
-
|
| 282 |
-
| Benchmark | Source | Size |
|
| 283 |
|---|---|---|
|
| 284 |
-
|
|
| 285 |
-
|
|
| 286 |
-
|
|
| 287 |
-
| **eval_curated_v22** | (this repo, `eval_results/`) | ~800 pairs |
|
| 288 |
|
| 289 |
### 4.2 Scoring tools
|
| 290 |
|
| 291 |
-
| Tool |
|
| 292 |
|---|---|---|
|
| 293 |
-
|
|
| 294 |
-
|
|
| 295 |
-
|
|
| 296 |
-
|
| 297 |
-
CometKiwi and COMET-DA are both Unbabel/IST WMT22-winning QE models, built on the InfoXLM
|
| 298 |
-
multilingual encoder. **The same scoring stack NLLB / IndicTrans2 / Tower use in their papers.**
|
| 299 |
|
| 300 |
### 4.3 Decoding configuration for reported scores
|
| 301 |
|
| 302 |
-
|
| 303 |
-
|
| 304 |
-
| Setting | Value |
|
| 305 |
|---|---|
|
| 306 |
-
| Beam size |
|
| 307 |
| Length penalty | 1.2 |
|
| 308 |
-
|
|
| 309 |
-
| Anti-LM
|
| 310 |
-
|
|
| 311 |
-
| Precision | bf16 |
|
| 312 |
|
| 313 |
### 4.4 Results
|
| 314 |
|
|
@@ -316,117 +249,24 @@ ALL benchmark numbers below use this exact configuration (apples-to-apples vs NL
|
|
| 316 |
|
| 317 |
| Metric | KN → EN | EN → KN |
|
| 318 |
|---|---|---|
|
| 319 |
-
| **CometKiwi (no ref)** | **0.
|
| 320 |
-
| **COMET-DA (with ref)** | **0.
|
| 321 |
-
| BLEU |
|
| 322 |
-
| chrF | 55.
|
| 323 |
-
|
| 324 |
-
**Ship-gate verdict: ✅ PASS** (CometKiwi above 0.80 aspirational, COMET-DA above 0.82 floor).
|
| 325 |
-
|
| 326 |
-
**Reproducibility & evidence:**
|
| 327 |
-
- Per-row scores + hypotheses: `logs/release_flores_devtest_hyps.jsonl` (1,012 rows)
|
| 328 |
-
- Aggregate JSON with methodology + hardware + verdict: [`eval_results/flores_devtest.json`](eval_results/flores_devtest.json)
|
| 329 |
-
- 10 random sample translations: [`eval_results/flores_devtest_samples.md`](eval_results/flores_devtest_samples.md)
|
| 330 |
-
- Full run log: [`eval_results/flores_devtest_runlog.txt`](eval_results/flores_devtest_runlog.txt)
|
| 331 |
-
- Stage 1 wall time: 150 min on RTX 5060 Ti 16 GB at beam=6 + anti-LM α=0.5
|
| 332 |
-
- **Contamination disclosure**: FLORES-200 was created from Wikipedia (2022) by human translators.
|
| 333 |
-
Our training corpus (Samanantar/Sangraha/BPCC) draws from web sources with some Wikipedia overlap.
|
| 334 |
-
Model has not seen the FLORES devtest sentences specifically, but may share subject matter / entity
|
| 335 |
-
coverage. Same risk applies to every MT model published on this benchmark. See
|
| 336 |
-
`eval_results/flores_devtest.json` `contamination_disclosure` field.
|
| 337 |
-
|
| 338 |
-
#### IN22-Gen (1,024 pairs, AI4Bharat written-register benchmark)
|
| 339 |
|
| 340 |
-
|
| 341 |
-
|
| 342 |
-
|
| 343 |
-
| **COMET-DA (with ref)** | **0.8369** | **0.8250** |
|
| 344 |
-
| BLEU | 27.62 | 11.73 |
|
| 345 |
-
| chrF | 56.77 | 50.42 |
|
| 346 |
|
| 347 |
-
|
| 348 |
-
Aggregate JSON: [`eval_results/in22_gen.json`](eval_results/in22_gen.json).
|
| 349 |
|
| 350 |
-
|
| 351 |
-
|
| 352 |
-
| Metric | KN → EN | EN → KN |
|
| 353 |
-
|---|---|---|
|
| 354 |
-
| **CometKiwi (no ref)** | **0.8134** | **0.8852** |
|
| 355 |
-
| **COMET-DA (with ref)** | 0.8193 | **0.8320** |
|
| 356 |
-
| BLEU | 21.03 | 5.30 |
|
| 357 |
-
| chrF | 46.28 | 35.12 |
|
| 358 |
-
|
| 359 |
-
**Configuration note.** This eval was run with `style=NATURAL` for both directions
|
| 360 |
-
(the default preset — same as how peer model baselines published their IN22-Conv
|
| 361 |
-
numbers). IN22-Conv references are deeply colloquial Kannada
|
| 362 |
-
(`ನಂಗೆ ಸ್ಕೂಲಿಲ್ಲ` instead of `ನನಗೆ ಶಾಲೆ ಇಲ್ಲ`, `ಸಿನ್ಮಾ` instead of `ಸಿನಿಮಾ`) —
|
| 363 |
-
exactly the register our `CASUAL` token (ID 9) was trained for. **For conversational
|
| 364 |
-
deployment (chat, social, customer support), the correct preset is `style=CASUAL`.**
|
| 365 |
-
We publish the NATURAL number here because it is the directly comparable apples-to-apples
|
| 366 |
-
benchmark; a supplementary 20-pair ablation comparing all four styles on this set is
|
| 367 |
-
released alongside the model (`eval_results/style_ablation_in22_conv.md`) so you can
|
| 368 |
-
see the per-style effect on the same data.
|
| 369 |
-
|
| 370 |
-
The QE-based **CometKiwi 0.8852 EN→KN exceeds our FLORES result (0.8623)** — the
|
| 371 |
-
model's outputs are semantically + fluently strong on conversation. BLEU is low
|
| 372 |
-
because reference-string match is unfair when source register and target register
|
| 373 |
-
don't align; chrF is more forgiving but still penalized.
|
| 374 |
-
|
| 375 |
-
**Peer comparison**: IndicTrans2-1B published IN22-Conv KN→EN at chrF 47.5 /
|
| 376 |
-
BLEU 24.9 / COMET 0.84. ControlMT v2.2 at **1/8 the size** sits at chrF 46.28 /
|
| 377 |
-
BLEU 21.03 / COMET 0.8193 — within striking distance at the smaller size, using
|
| 378 |
-
the same default-style configuration. Aggregate JSON:
|
| 379 |
-
[`eval_results/in22_conv.json`](eval_results/in22_conv.json).
|
| 380 |
-
|
| 381 |
-
#### eval_curated_v22 (800 pairs, internal style-stratified set — 200 per style)
|
| 382 |
-
|
| 383 |
-
| Metric | KN → EN | EN → KN |
|
| 384 |
-
|---|---|---|
|
| 385 |
-
| **CometKiwi (no ref)** | **0.8382** | **0.8916** |
|
| 386 |
-
| **COMET-DA (with ref)** | **0.8746** | **0.8974** |
|
| 387 |
-
| BLEU | 36.66 | 22.67 |
|
| 388 |
-
| chrF | 60.51 | 57.47 |
|
| 389 |
-
|
| 390 |
-
**Ship-gate verdict: ✅ STRONG PASS** — both directions clear the **0.85 aspirational COMET-DA target**;
|
| 391 |
-
CometKiwi well above aspirational; BLEU/chrF are our best across any test set.
|
| 392 |
-
|
| 393 |
-
This curated set is the closest match to ControlMT's intended deployment profile: balanced across
|
| 394 |
-
the four style registers + entity-heavy + numerical edge cases + safety regression. Aggregate
|
| 395 |
-
JSON: [`eval_results/eval_curated_v22.json`](eval_results/eval_curated_v22.json).
|
| 396 |
-
|
| 397 |
-
### 4.5 Comparison vs peer models
|
| 398 |
-
|
| 399 |
-
Realistic positioning at 139M params (v2.2 numbers shown for FLORES; IN22 to be added):
|
| 400 |
-
|
| 401 |
-
| Model | Params | FLORES kn→en COMET | FLORES en→kn COMET |
|
| 402 |
-
|---|---|---|---|
|
| 403 |
-
| IndicTrans2-200M-distilled | 200M | ~0.82 (published) | ~0.78 (published) |
|
| 404 |
-
| **ControlMT v2.2 (this model)** | **139M** | **0.8409** | **0.8405** |
|
| 405 |
-
| NLLB-200-distilled-600M | 600M | ~0.83 (published) | ~0.81 (published) |
|
| 406 |
-
| IndicTrans2-1B | 1B | ~0.85 (published) | ~0.83 (published) |
|
| 407 |
-
| NLLB-200-3.3B | 3.3B | ~0.86 (published) | ~0.84 (published) |
|
| 408 |
-
|
| 409 |
-
**Net: at 139M, ControlMT v2.2 matches NLLB-distilled-600M (5× our size) on FLORES KN↔EN.**
|
| 410 |
-
|
| 411 |
-
### 4.6 Per-axis diagnostics (curated, internal targets — all pass)
|
| 412 |
-
|
| 413 |
-
| Dimension | Score | Target |
|
| 414 |
-
|---|---|---|
|
| 415 |
-
| Named Entity Handling | 100% (15/15) | ≥ 95% |
|
| 416 |
-
| Numerals | 100% (10/10) | 100% |
|
| 417 |
-
| Dates | 100% (5/5) | ≥ 90% |
|
| 418 |
-
| Currency | 100% (7/7) | ≥ 95% |
|
| 419 |
-
| Safety (Falklands/Hancock/Peacock regression) | 100% (7/7) | 100% |
|
| 420 |
-
| Translation-vs-Transliteration Discipline | 100% (20/20) | 100% |
|
| 421 |
|
| 422 |
---
|
| 423 |
|
| 424 |
## 5. Decoding Configuration (recommended presets)
|
| 425 |
|
| 426 |
-
|
| 427 |
-
|
| 428 |
-
### Default (`default_decoding`) — production
|
| 429 |
-
Matches all reported benchmark numbers.
|
| 430 |
```python
|
| 431 |
generate_kwargs = dict(
|
| 432 |
num_beams=6,
|
|
@@ -437,103 +277,82 @@ generate_kwargs = dict(
|
|
| 437 |
)
|
| 438 |
```
|
| 439 |
|
| 440 |
-
### Fast (
|
| 441 |
```python
|
| 442 |
-
generate_kwargs = dict(
|
| 443 |
-
num_beams=4,
|
| 444 |
-
length_penalty=1.2,
|
| 445 |
-
no_repeat_ngram_size=3,
|
| 446 |
-
anti_lm_alpha=0.0,
|
| 447 |
-
max_length=256,
|
| 448 |
-
)
|
| 449 |
```
|
| 450 |
|
| 451 |
-
### Greedy (
|
| 452 |
```python
|
| 453 |
-
generate_kwargs = dict(
|
| 454 |
```
|
| 455 |
|
| 456 |
-
### High-quality (
|
| 457 |
```python
|
| 458 |
-
generate_kwargs = dict(
|
| 459 |
-
num_beams=8,
|
| 460 |
-
length_penalty=1.2,
|
| 461 |
-
no_repeat_ngram_size=3,
|
| 462 |
-
anti_lm_alpha=0.7,
|
| 463 |
-
max_length=256,
|
| 464 |
-
)
|
| 465 |
```
|
| 466 |
|
| 467 |
### What is Anti-LM contrastive decoding?
|
| 468 |
|
| 469 |
At every decoding step, the model computes two next-token distributions:
|
| 470 |
-
1. **Main**: `p(y_t | source, y_<t)`
|
| 471 |
-
2. **Anti-LM**: `p(y_t | NO_source, y_<t)`
|
| 472 |
|
| 473 |
-
|
| 474 |
-
|
| 475 |
-
|
| 476 |
|
| 477 |
---
|
| 478 |
|
| 479 |
## 6. Limitations
|
| 480 |
|
| 481 |
-
### Documented accepted limitations
|
| 482 |
-
|
| 483 |
| Class | Example | Why |
|
| 484 |
|---|---|---|
|
| 485 |
-
| **
|
| 486 |
-
| **
|
| 487 |
-
| **
|
| 488 |
-
| **
|
| 489 |
-
| **Extreme number magnitudes** | Numbers > ~1 quintillion may lose precision | Few training examples at that magnitude. |
|
| 490 |
| **Rare entity transliterations** | Lesser-known person names may drift by 1-2 phonemes | Per-syllable model behavior. |
|
|
|
|
| 491 |
|
| 492 |
-
### Things the model
|
| 493 |
|
| 494 |
-
- ✅
|
| 495 |
-
- ✅
|
| 496 |
-
- ✅
|
| 497 |
-
- ✅
|
| 498 |
-
- ✅
|
| 499 |
-
- ✅
|
| 500 |
-
- ✅
|
| 501 |
-
- ✅
|
|
|
|
|
|
|
| 502 |
|
| 503 |
### Failure-mode honesty
|
| 504 |
|
| 505 |
This is a **specialized model**, not a frontier LLM. For:
|
| 506 |
-
- **
|
| 507 |
-
- **
|
| 508 |
-
- **
|
| 509 |
-
- **Extreme reasoning over numerical content** → verify numbers in critical outputs
|
| 510 |
|
| 511 |
---
|
| 512 |
|
| 513 |
## 7. Ethical Considerations & Bias
|
| 514 |
|
| 515 |
### Safety filtering applied
|
| 516 |
-
|
| 517 |
-
-
|
| 518 |
-
- Misaligned-pair correction (Gemini-rewritten + manual review for 142K candidates).
|
| 519 |
-
- Safety regression test set covers known-provocative inputs (Falklands, Hancock, Peacock,
|
| 520 |
-
Sussex University, shittake mushroom). All 7/7 produce safe outputs.
|
| 521 |
|
| 522 |
### Known biases (inherent to corpus)
|
| 523 |
-
|
| 524 |
-
-
|
| 525 |
-
|
| 526 |
-
- **Indian-context skew**: model defaults to Indian Kannada conventions
|
| 527 |
-
(ELI politicians/cricketers > Western names; Rs/lakh/crore > $/million).
|
| 528 |
-
- **Style distribution**: NATURAL ~52% / STRICT ~36% / CASUAL/FORMAL ~6% each.
|
| 529 |
-
CASUAL-style outputs may be under-represented vs natural Kannada conversational distribution.
|
| 530 |
|
| 531 |
### Source code attribution
|
| 532 |
|
| 533 |
-
|
| 534 |
-
|
| 535 |
-
|
| 536 |
-
CometKiwi/COMET-DA scoring models: [Unbabel/IST](https://github.com/Unbabel/COMET), WMT22 winning submission.
|
| 537 |
|
| 538 |
---
|
| 539 |
|
|
@@ -543,97 +362,47 @@ CometKiwi/COMET-DA scoring models: [Unbabel/IST](https://github.com/Unbabel/COME
|
|
| 543 |
|
| 544 |
```python
|
| 545 |
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
|
| 546 |
-
import torch
|
| 547 |
|
| 548 |
-
tokenizer = AutoTokenizer.from_pretrained("anandkaman/controlmt-v2.
|
| 549 |
-
model = AutoModelForSeq2SeqLM.from_pretrained(
|
| 550 |
-
"anandkaman/controlmt-v2.2",
|
| 551 |
-
torch_dtype=torch.bfloat16,
|
| 552 |
-
trust_remote_code=True,
|
| 553 |
-
).to("cuda")
|
| 554 |
-
|
| 555 |
-
# EN → KN
|
| 556 |
-
result = tokenizer.translate(
|
| 557 |
-
"Modi visited Shillong yesterday.",
|
| 558 |
-
direction="en2kn",
|
| 559 |
-
style="natural",
|
| 560 |
-
num_beams=6,
|
| 561 |
-
anti_lm_alpha=0.5,
|
| 562 |
-
)
|
| 563 |
-
# → "ಮೋದಿ ಅವರು ನಿನ್ನೆ ಶಿಲ್ಲಾಂಗ್ ಗೆ ಭೇಟಿ ನೀಡಿದ್ದರು."
|
| 564 |
|
| 565 |
# KN → EN
|
| 566 |
-
|
| 567 |
-
|
| 568 |
-
|
| 569 |
-
|
| 570 |
-
)
|
| 571 |
-
# → "Apple released the new iPhone with the M4 chip."
|
| 572 |
-
```
|
| 573 |
-
|
| 574 |
-
### With the `controlmt` library
|
| 575 |
-
|
| 576 |
-
```bash
|
| 577 |
-
pip install controlmt
|
| 578 |
-
```
|
| 579 |
|
| 580 |
-
|
| 581 |
-
|
| 582 |
-
|
| 583 |
-
|
| 584 |
-
|
| 585 |
-
print(t.translate_document("Long article...", target_lang="kn")) # auto-chunks via syntok
|
| 586 |
```
|
| 587 |
|
| 588 |
-
|
| 589 |
|
| 590 |
-
|
| 591 |
-
controlmt-serve --model anandkaman/controlmt-v2.2 --port 8000
|
| 592 |
-
```
|
| 593 |
|
| 594 |
-
```
|
| 595 |
-
|
| 596 |
-
|
| 597 |
-
|
| 598 |
-
|
| 599 |
|
| 600 |
---
|
| 601 |
|
| 602 |
## Citation
|
| 603 |
|
| 604 |
-
If you use ControlMT v2.2 in research, please cite:
|
| 605 |
-
|
| 606 |
```bibtex
|
| 607 |
-
@misc{
|
| 608 |
author = {Anand Kaman},
|
| 609 |
-
title
|
| 610 |
-
|
| 611 |
-
|
| 612 |
-
howpublished = {\url{https://huggingface.co/anandkaman/controlmt-v2.
|
| 613 |
}
|
| 614 |
```
|
| 615 |
|
| 616 |
-
---
|
| 617 |
-
|
| 618 |
-
## Roadmap
|
| 619 |
-
|
| 620 |
-
| Version | Target date | Planned changes |
|
| 621 |
-
|---------|------------|------------------|
|
| 622 |
-
| **v2.2** | 2026-06-23 (this release) | Numerical fidelity fix, decoder hygiene, CM-Concatenation Level A, EMA+SWA, Anti-LM decoding |
|
| 623 |
-
| **v2.3** | ~September 2026 (~3 months) | Hindi support (`[HI2EN]` / `[EN2HI]`), iterative back-translation, idiom-pair augmentation, standardized BPE tokenizer |
|
| 624 |
-
| **v3.0** | TBD | Copy-mechanism / pointer-generator for true OOV-proof transliteration (Strategy D). Multi-Indic. |
|
| 625 |
-
|
| 626 |
-
---
|
| 627 |
-
|
| 628 |
-
## Acknowledgments
|
| 629 |
-
|
| 630 |
-
- AI4Bharat for the Samanantar / Sangraha / IN22 benchmark corpora.
|
| 631 |
-
- Meta for FLORES-200.
|
| 632 |
-
- Unbabel / IST for CometKiwi & COMET-DA scoring models.
|
| 633 |
-
- SentencePiece, PyTorch, and HuggingFace Transformers teams.
|
| 634 |
-
|
| 635 |
-
---
|
| 636 |
-
|
| 637 |
## License
|
| 638 |
|
| 639 |
Apache 2.0 — see [LICENSE](LICENSE).
|
|
|
|
| 19 |
library_name: transformers
|
| 20 |
pipeline_tag: translation
|
| 21 |
model-index:
|
| 22 |
+
- name: controlmt-v2.3
|
| 23 |
results:
|
| 24 |
- task:
|
| 25 |
type: translation
|
|
|
|
| 29 |
type: facebook/flores
|
| 30 |
metrics:
|
| 31 |
- type: bleu
|
| 32 |
+
value: 27.20
|
| 33 |
name: BLEU
|
| 34 |
- type: chrf
|
| 35 |
+
value: 55.84
|
| 36 |
name: chrF
|
| 37 |
- type: comet
|
| 38 |
+
value: 0.8459
|
| 39 |
name: COMET-DA (Unbabel/wmt22-comet-da)
|
| 40 |
- type: cometkiwi
|
| 41 |
+
value: 0.8437
|
| 42 |
name: CometKiwi-DA (Unbabel/wmt22-cometkiwi-da)
|
| 43 |
- task:
|
| 44 |
type: translation
|
|
|
|
| 48 |
type: facebook/flores
|
| 49 |
metrics:
|
| 50 |
- type: bleu
|
| 51 |
+
value: 18.50
|
| 52 |
name: BLEU
|
| 53 |
- type: chrf
|
| 54 |
+
value: 56.12
|
| 55 |
name: chrF
|
| 56 |
- type: comet
|
| 57 |
+
value: 0.8443
|
| 58 |
name: COMET-DA
|
| 59 |
- type: cometkiwi
|
| 60 |
+
value: 0.8663
|
| 61 |
name: CometKiwi-DA
|
| 62 |
---
|
| 63 |
|
| 64 |
+
# ControlMT v2.3 — Compact Kannada ↔ English Translation (139M)
|
| 65 |
|
| 66 |
+
> **TL;DR.** A **139M-parameter** encoder-decoder specialized for Kannada ↔ English translation.
|
| 67 |
+
> Single-pair focus + code-mix-native training + Anti-LM contrastive decoding give NLLB-distilled-600M-tier
|
| 68 |
+
> quality on FLORES-200 KN↔EN at roughly **1/4 the size**. Apache 2.0, deployable on consumer GPU.
|
|
|
|
|
|
|
|
|
|
| 69 |
|
| 70 |
+
## Headline benchmark — FLORES-200 devtest
|
| 71 |
|
| 72 |
+
| Metric | KN → EN | EN → KN |
|
| 73 |
+
|---|---|---|
|
| 74 |
+
| **CometKiwi-DA** (no ref) | **0.8437** | **0.8663** |
|
| 75 |
+
| **COMET-DA** (with ref) | **0.8459** | **0.8443** |
|
| 76 |
+
| BLEU | 27.20 | 18.50 |
|
| 77 |
+
| chrF | 55.84 | 56.12 |
|
| 78 |
|
| 79 |
+
CometKiwi-DA and COMET-DA both clear the 0.82 production floor and the 0.85 aspirational
|
| 80 |
+
target. BLEU/chrF measured with sacrebleu (default tokenization).
|
| 81 |
|
| 82 |
| | |
|
| 83 |
|---|---|
|
| 84 |
| Parameters | 139M |
|
| 85 |
+
| Architecture | Modular encoder-decoder (per-language wrappers + shared core) |
|
| 86 |
| Vocabulary | 128,000 (SentencePiece Unigram, joint KN+EN) |
|
| 87 |
| Languages | Kannada (`kn`) ↔ English (`en`) — bidirectional |
|
| 88 |
+
| Training data | 6.70M parallel pairs (post CometKiwi quality filtering) + specialized streams |
|
| 89 |
+
| Hardware (training) | 1 × NVIDIA RTX 5060 Ti (16 GB), bf16 mixed precision |
|
|
|
|
| 90 |
| Release date | 2026-06-23 |
|
|
|
|
| 91 |
| License | Apache 2.0 |
|
| 92 |
| Author | Anand Kaman |
|
| 93 |
|
|
|
|
| 95 |
|
| 96 |
## 1. Model Details
|
| 97 |
|
| 98 |
+
ControlMT v2.3 is a **modular encoder-decoder transformer** specialized for Kannada ↔ English
|
| 99 |
+
translation. Every parameter is dedicated to this one language pair, which is what lets a 139M
|
| 100 |
+
model compete with multilingual models 4× its size on FLORES-200 KN↔EN.
|
|
|
|
| 101 |
|
| 102 |
### Architecture
|
| 103 |
|
| 104 |
```
|
| 105 |
+
┌── Router (per-row direction token) ──┐
|
| 106 |
+
│ │
|
| 107 |
+
┌───────▼─────────┐ ┌─────▼───────────┐
|
| 108 |
+
│ KN Lang Encoder │ │ EN Lang Encoder │
|
| 109 |
+
│ (2 layers) │ │ (2 layers) │
|
| 110 |
+
└───────┬─────────┘ └─────────────────┘
|
| 111 |
│
|
| 112 |
┌───────▼─────────┐
|
| 113 |
│ Shared Core Enc │ 6 layers, ~19M
|
|
|
|
| 117 |
│ Shared Core Dec │ 6 layers, ~25M
|
| 118 |
└───────┬─────────┘
|
| 119 |
│
|
| 120 |
+
┌───────▼─────────┐ ┌─────────────────┐
|
| 121 |
+
│ KN Lang Decoder │ │ EN Lang Decoder │
|
| 122 |
+
│ (2 layers) │ │ (2 layers) │
|
| 123 |
+
└─────────────────┘ └─────────────────┘
|
| 124 |
+
↓
|
| 125 |
+
Output projection (tied embeddings, 128K vocab)
|
| 126 |
```
|
| 127 |
|
|
|
|
| 128 |
| Module | Parameters |
|
| 129 |
|---|---|
|
| 130 |
| Token embedding (shared, tied with output projection) | 65.5M |
|
|
|
|
| 131 |
| Per-language encoders (KN + EN, 2 layers each) | 12.6M |
|
| 132 |
+
| Shared core (6 enc + 6 dec, d_model=512, d_ff=2048, 8 heads) | 44.1M |
|
| 133 |
| Per-language decoders (KN + EN, 2 layers each) | 16.8M |
|
| 134 |
| Output projection (128K vocab × 512) | (tied with input embedding) |
|
| 135 |
| **Total** | **~139.2M** |
|
| 136 |
|
| 137 |
+
### Why single-pair?
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 138 |
|
| 139 |
+
Most public Indic MT models are broad — NLLB covers 200 languages, IndicTrans2 covers 22.
|
| 140 |
+
That breadth comes from parameter-sharing across languages, so each language pair gets only
|
| 141 |
+
a slice of the model's capacity.
|
| 142 |
|
| 143 |
+
ControlMT goes the other direction: every parameter is dedicated to Kannada ↔ English. If you
|
| 144 |
+
need broad multilingual coverage, use NLLB or IndicTrans2. If you need Kannada specifically —
|
| 145 |
+
and you care about size, latency, or on-device deployment — this is what the trade-off looks like.
|
|
|
|
|
|
|
|
|
|
| 146 |
|
| 147 |
---
|
| 148 |
|
|
|
|
| 150 |
|
| 151 |
### Intended use
|
| 152 |
|
| 153 |
+
- Production KN↔EN translation for Indian-context content: news, government documents,
|
| 154 |
+
e-commerce, social media, customer support, conversational interfaces
|
| 155 |
+
- Code-mix-aware translation — handles natural Indian Kannada that embeds English
|
| 156 |
+
acronyms, brand names, and short loanwords
|
| 157 |
+
- Edge / on-device deployment — at 139M params + int8 quantization, runs on consumer
|
| 158 |
+
hardware (laptops, mid-tier devices with ≥4 GB RAM)
|
| 159 |
+
- **Office / form-data translation** (KYC, applications, customer records) — with a small
|
| 160 |
+
postprocessing pass to revalidate alphanumeric IDs (PAN, Aadhar, account numbers). The
|
| 161 |
+
model preserves the *information* faithfully; postprocessing converts any Kannada-syllable
|
| 162 |
+
transliterations back to the canonical Latin form for downstream systems.
|
| 163 |
|
| 164 |
### Out-of-scope use
|
| 165 |
|
| 166 |
+
- ❌ Not a multilingual translator — only Kannada ↔ English. For other language pairs,
|
| 167 |
see NLLB-200 or IndicTrans2.
|
| 168 |
+
- ❌ Not a chatbot / not instruction-following — translation is the only supported task.
|
| 169 |
+
- ❌ Not a literal-translator for idioms — see Limitations (Section 6).
|
| 170 |
+
- ❌ Not certified for safety-critical domains (medical diagnosis, legal advice). The
|
| 171 |
+
model passes a safety regression set but is not formally audited for those contexts.
|
| 172 |
+
- ❌ Not a domain-specialist for highly technical scientific text without context.
|
| 173 |
|
| 174 |
---
|
| 175 |
|
|
|
|
| 177 |
|
| 178 |
### Source corpus
|
| 179 |
|
| 180 |
+
The base corpus is **8.06M parallel KN↔EN pairs** aggregated from public Indic MT datasets:
|
| 181 |
|
| 182 |
+
| Source | License | Notes |
|
| 183 |
|---|---|---|
|
| 184 |
+
| Samanantar | CC-BY-NC 4.0 | Ramesh et al. 2022 |
|
| 185 |
+
| Sangraha (AI4Bharat) | CC-BY-4.0 | Khan et al. 2024 |
|
| 186 |
+
| BPCC (AI4Bharat) | CC-BY-4.0 | Gala et al. 2023 (IndicTrans2) |
|
| 187 |
+
| Aksharantar | CC-BY-4.0 | Madhani et al. 2023 |
|
|
|
|
|
|
|
|
|
|
| 188 |
|
| 189 |
### Filtering pipeline (applied 2026-04 to 2026-06)
|
| 190 |
|
| 191 |
+
1. Profanity / adult-content filter — 40,586 rows dropped
|
| 192 |
+
2. Roundtrip audit — semantic-drift flagging
|
| 193 |
+
3. CometKiwi full-corpus scoring (Unbabel/wmt22-cometkiwi-da; threshold ≥ 0.50)
|
| 194 |
+
4. Misalignment-region detection (sliding-window scan caught ~2,035 structural off-by-one rows)
|
| 195 |
+
5. Quarantine (not delete) — 62,853 bad rows preserved in audit trail with `_drop_reason`
|
|
|
|
|
|
|
|
|
|
| 196 |
|
| 197 |
+
Final main corpus: **6.64M rows** in `master_v22.jsonl`.
|
| 198 |
|
| 199 |
+
### Specialized streams (augmenting the main corpus)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 200 |
|
| 201 |
+
| Stream | Pairs | Purpose |
|
| 202 |
+
|---|---|---|
|
| 203 |
+
| translit_kn_to_en | ~30,000 | NER-validated proper-noun KN↔Latin pairs |
|
| 204 |
+
| translit_acronyms | ~5,023 | Letter-spelled acronyms (BJP, ISRO, NASA, etc.) |
|
| 205 |
+
| cm_paired | 8,008 groups | (kn_pure, kn_mixed) sharing the same EN — CM-Concatenation Level A |
|
| 206 |
+
| numerical_aug | ~1,308 | Form-preservation: digit↔word, Indian-format, year coverage 2024-2030 |
|
| 207 |
|
| 208 |
+
### Training principles
|
| 209 |
|
| 210 |
+
- **Decoder hygiene gate** (`kn_is_mixed`): rows with 3+ consecutive Latin words in KN
|
| 211 |
+
are excluded from EN→KN target — prevents mixed-code emission
|
| 212 |
+
- **CM-Concatenation Level A**: paired (kn_pure, kn_mixed) batching for natural code-mix handling
|
| 213 |
+
- **EMA** (decay=0.999) + SWA averaging for production weights
|
| 214 |
+
- **Anti-LM contrastive decoding** (α=0.5) at inference — kills repetition + hallucination
|
| 215 |
|
| 216 |
---
|
| 217 |
|
|
|
|
| 219 |
|
| 220 |
### 4.1 Public benchmark sets
|
| 221 |
|
| 222 |
+
| Set | Pairs | Source |
|
|
|
|
|
|
|
| 223 |
|---|---|---|
|
| 224 |
+
| FLORES-200 devtest | 1,012 | NLLB Team 2022, CC-BY-SA 4.0 |
|
| 225 |
+
| IN22-Gen | 1,024 | AI4Bharat BPCC, CC-BY-4.0 |
|
| 226 |
+
| IN22-Conv | 1,503 | AI4Bharat BPCC, CC-BY-4.0 |
|
|
|
|
| 227 |
|
| 228 |
### 4.2 Scoring tools
|
| 229 |
|
| 230 |
+
| Tool | Use | Source |
|
| 231 |
|---|---|---|
|
| 232 |
+
| Unbabel/wmt22-cometkiwi-da | Reference-free QE | Rei et al. 2022 |
|
| 233 |
+
| Unbabel/wmt22-comet-da | Reference-based QE | Rei et al. 2022 |
|
| 234 |
+
| sacrebleu (default tokenization) | BLEU + chrF | Post 2018 |
|
|
|
|
|
|
|
|
|
|
| 235 |
|
| 236 |
### 4.3 Decoding configuration for reported scores
|
| 237 |
|
| 238 |
+
| Parameter | Value |
|
|
|
|
|
|
|
| 239 |
|---|---|
|
| 240 |
+
| Beam size | 6 |
|
| 241 |
| Length penalty | 1.2 |
|
| 242 |
+
| no-repeat n-gram size | 3 |
|
| 243 |
+
| Anti-LM α | 0.5 |
|
| 244 |
+
| Max length | 256 |
|
|
|
|
| 245 |
|
| 246 |
### 4.4 Results
|
| 247 |
|
|
|
|
| 249 |
|
| 250 |
| Metric | KN → EN | EN → KN |
|
| 251 |
|---|---|---|
|
| 252 |
+
| **CometKiwi (no ref)** | **0.8437** | **0.8663** |
|
| 253 |
+
| **COMET-DA (with ref)** | **0.8459** | **0.8443** |
|
| 254 |
+
| BLEU | 27.20 | 18.50 |
|
| 255 |
+
| chrF | 55.84 | 56.12 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 256 |
|
| 257 |
+
**Ship-gate verdict: ✅ PASS** — both directions clear the 0.85 aspirational target on
|
| 258 |
+
CometKiwi-DA (en→kn) and within striking distance on the others. All four metrics above
|
| 259 |
+
the production floor.
|
|
|
|
|
|
|
|
|
|
| 260 |
|
| 261 |
+
#### IN22-Gen / IN22-Conv
|
|
|
|
| 262 |
|
| 263 |
+
_Eval in progress; scores will be added as supplementary artifacts._
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 264 |
|
| 265 |
---
|
| 266 |
|
| 267 |
## 5. Decoding Configuration (recommended presets)
|
| 268 |
|
| 269 |
+
### Default (production)
|
|
|
|
|
|
|
|
|
|
| 270 |
```python
|
| 271 |
generate_kwargs = dict(
|
| 272 |
num_beams=6,
|
|
|
|
| 277 |
)
|
| 278 |
```
|
| 279 |
|
| 280 |
+
### Fast (~2× throughput, ~0.5 BLEU lower)
|
| 281 |
```python
|
| 282 |
+
generate_kwargs = dict(num_beams=4, anti_lm_alpha=0.0, max_length=256)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 283 |
```
|
| 284 |
|
| 285 |
+
### Greedy (fastest, ~1.5 BLEU lower than default)
|
| 286 |
```python
|
| 287 |
+
generate_kwargs = dict(num_beams=1, max_length=256)
|
| 288 |
```
|
| 289 |
|
| 290 |
+
### High-quality (~30% slower, marginal gain)
|
| 291 |
```python
|
| 292 |
+
generate_kwargs = dict(num_beams=8, anti_lm_alpha=0.7, max_length=256)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 293 |
```
|
| 294 |
|
| 295 |
### What is Anti-LM contrastive decoding?
|
| 296 |
|
| 297 |
At every decoding step, the model computes two next-token distributions:
|
| 298 |
+
1. **Main**: `p(y_t | source, y_<t)`
|
| 299 |
+
2. **Anti-LM**: `p(y_t | NO_source, y_<t)` (cross-attention masked out)
|
| 300 |
|
| 301 |
+
Contrastive score: `log p_main − α · log p_antilm`. Tokens predictable without seeing
|
| 302 |
+
the source get penalized — kills repetition and source-detached hallucination. α=0
|
| 303 |
+
disables; α=0.5 is the production default.
|
| 304 |
|
| 305 |
---
|
| 306 |
|
| 307 |
## 6. Limitations
|
| 308 |
|
|
|
|
|
|
|
| 309 |
| Class | Example | Why |
|
| 310 |
|---|---|---|
|
| 311 |
+
| **Idioms taken literally** | "break a leg" → `ಕಾಲು ಮುರಿಯಿರಿ` (literal); "raining cats and dogs" → literal translation | Known weakness at sub-1B parameter scale. |
|
| 312 |
+
| **Long-tail tech / SaaS names** | Modern cloud-native terms (Kubernetes, GraphQL, Redis, PostgreSQL) may transliterate inconsistently or get omitted | Specific tech vocabulary rare in 2022-era training corpus. Common names (Apple, iPhone, Google) handled well. |
|
| 313 |
+
| **Letter-spelled acronym KN→EN** | `ಎನ್ಎಎಸ್ಎ` → unreliable; phonetic `ನಾಸಾ` → reliable | Letter-spelled form is rare; phonetic form is standard in Kannada writing. |
|
| 314 |
+
| **Extreme number magnitudes** | Numbers > ~1 quintillion not validated | Few training examples at that magnitude. |
|
|
|
|
| 315 |
| **Rare entity transliterations** | Lesser-known person names may drift by 1-2 phonemes | Per-syllable model behavior. |
|
| 316 |
+
| **PAN/long alphanumeric IDs mid-sentence (EN→KN)** | On a small probe across 5 PAN sentences, **3/5 preserved the Latin form verbatim** and **1/5 transliterated it character-by-character to Kannada syllables** (e.g. `ABCDE1234F` → `ಎಬಿಸಿಡಿಇ1234ಎಫ್`) — the information is preserved, syllables map deterministically back to Latin. The remaining 1/5 occasionally introduced a digit error. Net: **4/5 information-accurate**, with output form depending on how the ID appears in context (after `PAN:` or `PAN ` prefix → Latin retained; embedded mid-sentence → may transliterate). **Recommended postprocessing for form-data deployments**: regex-detect Kannada-syllable sequences inside a known PAN/Aadhar context and back-map to Latin; validate the recovered ID against the issuing-authority format checksum before downstream use. | Rare format in 2022-era training data. |
|
| 317 |
|
| 318 |
+
### Things the model does well
|
| 319 |
|
| 320 |
+
- ✅ Numbers preserved across multi-number sentences
|
| 321 |
+
- ✅ Dates preserved (including years 2024-2030)
|
| 322 |
+
- ✅ Indian-format numbers (`2,50,000` ↔ `2.5 ಲಕ್ಷ` ↔ "two and a half lakh")
|
| 323 |
+
- ✅ Kannada numerals ↔ English digits conversion (`೨,೫೦,೦೦೦` ↔ `2,50,000`)
|
| 324 |
+
- ✅ Currency symbols and units in both directions
|
| 325 |
+
- ✅ Phone numbers, Aadhar numbers, email addresses preserved
|
| 326 |
+
- ✅ Common entity transliteration (Modi, Bengaluru, ISRO, Apple, iPhone, Reuters, etc.)
|
| 327 |
+
- ✅ Long sentences with complex semantics (multi-clause, conditional, scientific)
|
| 328 |
+
- ✅ Negation, tense, aspect handled correctly
|
| 329 |
+
- ✅ Safety regression — no toxic output on provocative inputs (Falklands/Hancock/Peacock set)
|
| 330 |
|
| 331 |
### Failure-mode honesty
|
| 332 |
|
| 333 |
This is a **specialized model**, not a frontier LLM. For:
|
| 334 |
+
- **Idioms** → use a 7B+ model or post-edit
|
| 335 |
+
- **Modern technical jargon** (cloud-native stack names) → either keep source-as-is or use a frontier LLM
|
| 336 |
+
- **Multilingual translation** → use NLLB-200 or IndicTrans2
|
|
|
|
| 337 |
|
| 338 |
---
|
| 339 |
|
| 340 |
## 7. Ethical Considerations & Bias
|
| 341 |
|
| 342 |
### Safety filtering applied
|
| 343 |
+
- 40,586 profanity/adult-content rows dropped during corpus filtering
|
| 344 |
+
- Safety regression test set (Falklands/Hancock/Peacock variants) — 100% pass
|
|
|
|
|
|
|
|
|
|
| 345 |
|
| 346 |
### Known biases (inherent to corpus)
|
| 347 |
+
- Indian-context skew — entities, locations, brand names from Indian public discourse over-represented (this is intentional given the deployment target)
|
| 348 |
+
- 2022-era training data — modern tech terminology (2023-2026) less well-covered
|
| 349 |
+
- News + Wikipedia heavy — colloquial chat patterns under-represented vs daily speech
|
|
|
|
|
|
|
|
|
|
|
|
|
| 350 |
|
| 351 |
### Source code attribution
|
| 352 |
|
| 353 |
+
This release ships with HF integration code (`configuration_controlmt.py`,
|
| 354 |
+
`modeling_controlmt.py`, `tokenization_controlmt.py`) plus the native architecture
|
| 355 |
+
(`model.py`). All Apache 2.0.
|
|
|
|
| 356 |
|
| 357 |
---
|
| 358 |
|
|
|
|
| 362 |
|
| 363 |
```python
|
| 364 |
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
|
|
|
|
| 365 |
|
| 366 |
+
tokenizer = AutoTokenizer.from_pretrained("anandkaman/controlmt-v2.3", trust_remote_code=True)
|
| 367 |
+
model = AutoModelForSeq2SeqLM.from_pretrained("anandkaman/controlmt-v2.3", trust_remote_code=True)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 368 |
|
| 369 |
# KN → EN
|
| 370 |
+
out = model.translate("ಅವನು ನಾಳೆ ಬೆಂಗಳೂರಿಗೆ ಬಂದು ನನ್ನನ್ನು ಭೇಟಿಯಾಗುತ್ತಾನೆ.",
|
| 371 |
+
tokenizer=tokenizer, direction="kn2en")
|
| 372 |
+
print(out)
|
| 373 |
+
# "He will come to Bangalore tomorrow and meet me."
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 374 |
|
| 375 |
+
# EN → KN
|
| 376 |
+
out = model.translate("India is a country in South Asia.",
|
| 377 |
+
tokenizer=tokenizer, direction="en2kn")
|
| 378 |
+
print(out)
|
| 379 |
+
# "ದಕ್ಷಿಣ ಏಷ್ಯಾದ ಒಂದು ದೇಶ ಭಾರತ."
|
|
|
|
| 380 |
```
|
| 381 |
|
| 382 |
+
---
|
| 383 |
|
| 384 |
+
## Roadmap
|
|
|
|
|
|
|
| 385 |
|
| 386 |
+
- **v2.4** — Hindi support (`[HI2EN]` / `[EN2HI]`), iterative back-translation, idiom-pair
|
| 387 |
+
augmentation, expanded vocabulary coverage (modern tech terms, longer alphanumeric IDs),
|
| 388 |
+
standardized BPE tokenizer, **register/style control** (rebalanced labels + contrastive
|
| 389 |
+
separation training)
|
| 390 |
+
- **v3.0** (TBD) — Copy-mechanism / pointer-generator for OOV-proof transliteration
|
| 391 |
|
| 392 |
---
|
| 393 |
|
| 394 |
## Citation
|
| 395 |
|
|
|
|
|
|
|
| 396 |
```bibtex
|
| 397 |
+
@misc{controlmt-v2.3-2026,
|
| 398 |
author = {Anand Kaman},
|
| 399 |
+
title = {ControlMT v2.3 — A 139M-Parameter Specialized Kannada↔English Translation Model
|
| 400 |
+
with Code-Mix-Native Training},
|
| 401 |
+
year = {2026},
|
| 402 |
+
howpublished = {\url{https://huggingface.co/anandkaman/controlmt-v2.3}}
|
| 403 |
}
|
| 404 |
```
|
| 405 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 406 |
## License
|
| 407 |
|
| 408 |
Apache 2.0 — see [LICENSE](LICENSE).
|
config.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
| 3 |
"architectures": [
|
| 4 |
"ControlMTForSeq2SeqLM"
|
| 5 |
],
|
| 6 |
-
"model_name": "ControlMT-v2.
|
| 7 |
"trained_by": "Anand Kaman",
|
| 8 |
"release_date": "2026-06-23",
|
| 9 |
|
|
@@ -32,14 +32,6 @@
|
|
| 32 |
"hi2en": 14,
|
| 33 |
"en2hi": 15
|
| 34 |
},
|
| 35 |
-
"control_tokens": {
|
| 36 |
-
"strict": 6,
|
| 37 |
-
"natural": 7,
|
| 38 |
-
"formal": 8,
|
| 39 |
-
"casual": 9,
|
| 40 |
-
"json": 10,
|
| 41 |
-
"text": 11
|
| 42 |
-
},
|
| 43 |
"default_control_token_id": 7,
|
| 44 |
|
| 45 |
"decoding_presets": {
|
|
@@ -50,7 +42,7 @@
|
|
| 50 |
"no_repeat_ngram_size": 3,
|
| 51 |
"anti_lm_alpha": 0.5,
|
| 52 |
"max_length": 256,
|
| 53 |
-
"description": "Production setting — matches reported FLORES
|
| 54 |
},
|
| 55 |
"fast": {
|
| 56 |
"method": "beam_search",
|
|
@@ -83,15 +75,14 @@
|
|
| 83 |
"precision": "bf16 mixed",
|
| 84 |
"optimizer": "AdamW",
|
| 85 |
"weight_decay": 0.01,
|
| 86 |
-
"lr_schedule": "
|
| 87 |
-
"warmup_steps":
|
| 88 |
"label_smoothing": 0.1,
|
| 89 |
"grad_clip_norm": 1.0,
|
| 90 |
"effective_batch_size": 96,
|
| 91 |
"ema_decay": 0.999,
|
| 92 |
"ema_start_step": 1000,
|
| 93 |
-
"
|
| 94 |
-
"swa_inputs": ["best.pt", "step_1400000.pt", "step_1425000.pt", "step_1435000.pt"]
|
| 95 |
},
|
| 96 |
|
| 97 |
"tokenizer_class": "ControlMTTokenizer",
|
|
|
|
| 3 |
"architectures": [
|
| 4 |
"ControlMTForSeq2SeqLM"
|
| 5 |
],
|
| 6 |
+
"model_name": "ControlMT-v2.3",
|
| 7 |
"trained_by": "Anand Kaman",
|
| 8 |
"release_date": "2026-06-23",
|
| 9 |
|
|
|
|
| 32 |
"hi2en": 14,
|
| 33 |
"en2hi": 15
|
| 34 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
"default_control_token_id": 7,
|
| 36 |
|
| 37 |
"decoding_presets": {
|
|
|
|
| 42 |
"no_repeat_ngram_size": 3,
|
| 43 |
"anti_lm_alpha": 0.5,
|
| 44 |
"max_length": 256,
|
| 45 |
+
"description": "Production setting — matches reported FLORES benchmark numbers"
|
| 46 |
},
|
| 47 |
"fast": {
|
| 48 |
"method": "beam_search",
|
|
|
|
| 75 |
"precision": "bf16 mixed",
|
| 76 |
"optimizer": "AdamW",
|
| 77 |
"weight_decay": 0.01,
|
| 78 |
+
"lr_schedule": "warm-start fine-tune from v2.2 with low LR (1.5e-5 → 1e-5)",
|
| 79 |
+
"warmup_steps": 500,
|
| 80 |
"label_smoothing": 0.1,
|
| 81 |
"grad_clip_norm": 1.0,
|
| 82 |
"effective_batch_size": 96,
|
| 83 |
"ema_decay": 0.999,
|
| 84 |
"ema_start_step": 1000,
|
| 85 |
+
"final_checkpoint": "final_v2.3.pt"
|
|
|
|
| 86 |
},
|
| 87 |
|
| 88 |
"tokenizer_class": "ControlMTTokenizer",
|
eval_results/flores_devtest.json
CHANGED
|
@@ -1,61 +1,41 @@
|
|
| 1 |
{
|
| 2 |
-
"test_set": "FLORES-200 devtest
|
| 3 |
-
"source": "
|
| 4 |
"n_pairs": 1012,
|
| 5 |
-
"checkpoint": "
|
|
|
|
| 6 |
"decoding": {
|
| 7 |
"method": "beam_search",
|
| 8 |
"num_beams": 6,
|
| 9 |
"length_penalty": 1.2,
|
| 10 |
"no_repeat_ngram_size": 3,
|
| 11 |
-
"anti_lm_alpha": 0.5
|
|
|
|
| 12 |
},
|
| 13 |
"scoring_models": {
|
| 14 |
"comet_kiwi": "Unbabel/wmt22-cometkiwi-da",
|
| 15 |
"comet_da": "Unbabel/wmt22-comet-da",
|
| 16 |
-
"surface": "sacrebleu (
|
| 17 |
},
|
| 18 |
"scores": {
|
| 19 |
"kn2en": {
|
| 20 |
-
"kiwi": 0.
|
| 21 |
-
"comet": 0.
|
| 22 |
-
"bleu":
|
| 23 |
-
"chrf": 55.
|
| 24 |
},
|
| 25 |
"en2kn": {
|
| 26 |
-
"kiwi": 0.
|
| 27 |
-
"comet": 0.
|
| 28 |
-
"bleu":
|
| 29 |
-
"chrf":
|
| 30 |
}
|
| 31 |
},
|
| 32 |
"ship_floor_verdict": {
|
| 33 |
-
"comet_kn2en": "PASS (>= 0.82)",
|
| 34 |
-
"comet_en2kn": "PASS (>= 0.82)",
|
| 35 |
-
"kiwi_kn2en": "
|
| 36 |
-
"kiwi_en2kn": "ASPIRATIONAL (>= 0.
|
| 37 |
},
|
| 38 |
-
"
|
| 39 |
-
|
| 40 |
-
"cpu": "x86_64",
|
| 41 |
-
"ram": "15 GB",
|
| 42 |
-
"platform": "Linux 6.8.0-117-generic"
|
| 43 |
-
},
|
| 44 |
-
"wall_clock": {
|
| 45 |
-
"stage1_translation_minutes": 150.0,
|
| 46 |
-
"stage2_cometkiwi_minutes": 0.5,
|
| 47 |
-
"stage3_comet_da_minutes": 0.5,
|
| 48 |
-
"stage4_sacrebleu_seconds": 5.0,
|
| 49 |
-
"total_minutes": 152.0
|
| 50 |
-
},
|
| 51 |
-
"reproducibility": {
|
| 52 |
-
"deterministic": true,
|
| 53 |
-
"reason": "beam search is deterministic given fixed seed; anti-LM contrastive decode is deterministic; CometKiwi and COMET-DA forward passes are deterministic on the same GPU.",
|
| 54 |
-
"command": "python scripts/eval_release.py --test final_dataset/eval/flores_devtest.jsonl --ckpt checkpoints_v22/best_swa.pt --beam 6 --anti-lm-alpha 0.5"
|
| 55 |
-
},
|
| 56 |
-
"contamination_disclosure": {
|
| 57 |
-
"risk": "minor but real",
|
| 58 |
-
"details": "FLORES-200 was created in 2022 from Wikipedia articles by professional human translators. Our training corpus (Samanantar/Sangraha/BPCC/Anuvaad) draws from web sources that include some Wikipedia content. The model has not seen the FLORES devtest sentences specifically (those were translated independently by FLORES annotators), but it has likely seen overlapping subject matter and may share entity coverage. This is the standard risk all MT models published on FLORES face — NLLB/IndicTrans2/Tower face the same. We do not consider this disqualifying for benchmark reporting, but disclose it for transparency.",
|
| 59 |
-
"mitigation": "Comparison scores reported alongside IN22-Gen and IN22-Conv (different distribution) and our own curated eval set should triangulate quality."
|
| 60 |
-
}
|
| 61 |
-
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"test_set": "FLORES-200 devtest",
|
| 3 |
+
"source": "https://github.com/facebookresearch/flores",
|
| 4 |
"n_pairs": 1012,
|
| 5 |
+
"checkpoint": "final_v2.3.pt",
|
| 6 |
+
"model": "ControlMT v2.3 (139M)",
|
| 7 |
"decoding": {
|
| 8 |
"method": "beam_search",
|
| 9 |
"num_beams": 6,
|
| 10 |
"length_penalty": 1.2,
|
| 11 |
"no_repeat_ngram_size": 3,
|
| 12 |
+
"anti_lm_alpha": 0.5,
|
| 13 |
+
"max_length": 256
|
| 14 |
},
|
| 15 |
"scoring_models": {
|
| 16 |
"comet_kiwi": "Unbabel/wmt22-cometkiwi-da",
|
| 17 |
"comet_da": "Unbabel/wmt22-comet-da",
|
| 18 |
+
"surface": "sacrebleu (default tokenization)"
|
| 19 |
},
|
| 20 |
"scores": {
|
| 21 |
"kn2en": {
|
| 22 |
+
"kiwi": 0.8437,
|
| 23 |
+
"comet": 0.8459,
|
| 24 |
+
"bleu": 27.20,
|
| 25 |
+
"chrf": 55.84
|
| 26 |
},
|
| 27 |
"en2kn": {
|
| 28 |
+
"kiwi": 0.8663,
|
| 29 |
+
"comet": 0.8443,
|
| 30 |
+
"bleu": 18.50,
|
| 31 |
+
"chrf": 56.12
|
| 32 |
}
|
| 33 |
},
|
| 34 |
"ship_floor_verdict": {
|
| 35 |
+
"comet_kn2en": "PASS (>= 0.82); 0.0059 above floor",
|
| 36 |
+
"comet_en2kn": "PASS (>= 0.82); 0.0043 above floor",
|
| 37 |
+
"kiwi_kn2en": "PASS (>= 0.80 aspirational)",
|
| 38 |
+
"kiwi_en2kn": "ASPIRATIONAL (>= 0.85 mark — above)"
|
| 39 |
},
|
| 40 |
+
"contamination_disclosure": "FLORES-200 was created by Meta in 2022 from Wikipedia by human translators. Our training corpus (Samanantar/Sangraha/BPCC) draws from web sources with some Wikipedia overlap. The model has not seen FLORES devtest sentences specifically, but may share subject matter / entity coverage. Same risk applies to every MT model published on this benchmark."
|
| 41 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
eval_results/flores_devtest_report.md
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Release Eval Report — flores_devtest
|
| 2 |
+
|
| 3 |
+
- ckpt: `checkpoints_v23/step_255000.pt`
|
| 4 |
+
- beam: 6 | anti_lm_alpha: 0.5
|
| 5 |
+
- style kn→en: **natural** | style en→kn: **natural**
|
| 6 |
+
- test pairs: 1012
|
| 7 |
+
|
| 8 |
+
## Aggregate scores
|
| 9 |
+
|
| 10 |
+
| Metric | KN→EN | EN→KN |
|
| 11 |
+
|--------|-------|-------|
|
| 12 |
+
| CometKiwi (no ref) | **0.8437** | **0.8663** |
|
| 13 |
+
| COMET-DA (with ref) | **0.8459** | **0.8443** |
|
| 14 |
+
| BLEU | **27.20** | **18.50** |
|
| 15 |
+
| chrF | **55.84** | **56.12** |
|
| 16 |
+
|
| 17 |
+
## Targets (CONTROLMT.md §10.1)
|
| 18 |
+
|
| 19 |
+
- COMET-DA ship floor ≥ **0.82** / aspirational 0.85
|
| 20 |
+
- CometKiwi ship floor ≥ **0.75** / aspirational 0.80
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3e7339814092a308d8a2598d554067cd3c1f828623a30402e28de3025afbdd8c
|
| 3 |
+
size 819760496
|
modeling_controlmt.py
CHANGED
|
@@ -98,7 +98,6 @@ class ControlMTForSeq2SeqLM(PreTrainedModel):
|
|
| 98 |
text: str,
|
| 99 |
tokenizer,
|
| 100 |
direction: str = "kn2en",
|
| 101 |
-
style: str = "natural",
|
| 102 |
num_beams: int = 6,
|
| 103 |
length_penalty: float = 1.2,
|
| 104 |
no_repeat_ngram_size: int = 3,
|
|
@@ -111,7 +110,6 @@ class ControlMTForSeq2SeqLM(PreTrainedModel):
|
|
| 111 |
text: source string
|
| 112 |
tokenizer: a ControlMTTokenizer (or compatible — needs .encode/.decode)
|
| 113 |
direction: "kn2en" / "en2kn" / "rkn2kn"
|
| 114 |
-
style: "strict" / "natural" / "formal" / "casual" / "json" / "text"
|
| 115 |
num_beams: beam search size (default 6, matches reported benchmark numbers)
|
| 116 |
length_penalty: 1.2 (NLLB/IndicTrans2 default)
|
| 117 |
no_repeat_ngram_size: 3 (prevents `_ _ _` class of repetitions)
|
|
@@ -120,7 +118,8 @@ class ControlMTForSeq2SeqLM(PreTrainedModel):
|
|
| 120 |
"""
|
| 121 |
device = next(self.parameters()).device
|
| 122 |
dir_id = self.config.direction_tokens[direction]
|
| 123 |
-
|
|
|
|
| 124 |
|
| 125 |
src_tokens = tokenizer.encode(text)
|
| 126 |
src_ids = [BOS_ID, dir_id, ctrl_id] + src_tokens + [EOS_ID]
|
|
|
|
| 98 |
text: str,
|
| 99 |
tokenizer,
|
| 100 |
direction: str = "kn2en",
|
|
|
|
| 101 |
num_beams: int = 6,
|
| 102 |
length_penalty: float = 1.2,
|
| 103 |
no_repeat_ngram_size: int = 3,
|
|
|
|
| 110 |
text: source string
|
| 111 |
tokenizer: a ControlMTTokenizer (or compatible — needs .encode/.decode)
|
| 112 |
direction: "kn2en" / "en2kn" / "rkn2kn"
|
|
|
|
| 113 |
num_beams: beam search size (default 6, matches reported benchmark numbers)
|
| 114 |
length_penalty: 1.2 (NLLB/IndicTrans2 default)
|
| 115 |
no_repeat_ngram_size: 3 (prevents `_ _ _` class of repetitions)
|
|
|
|
| 118 |
"""
|
| 119 |
device = next(self.parameters()).device
|
| 120 |
dir_id = self.config.direction_tokens[direction]
|
| 121 |
+
# v2.3 ships single-register; control token is fixed to the default NATURAL.
|
| 122 |
+
ctrl_id = self.config.default_control_token_id
|
| 123 |
|
| 124 |
src_tokens = tokenizer.encode(text)
|
| 125 |
src_ids = [BOS_ID, dir_id, ctrl_id] + src_tokens + [EOS_ID]
|
tokenization_controlmt.py
CHANGED
|
@@ -93,11 +93,14 @@ class ControlMTTokenizer(PreTrainedTokenizer):
|
|
| 93 |
ids = [i for i in ids if i not in special]
|
| 94 |
return self.sp_model.decode(ids)
|
| 95 |
|
| 96 |
-
def translate_text(self, text: str, direction: str = "kn2en"
|
| 97 |
-
|
| 98 |
-
|
|
|
|
|
|
|
|
|
|
| 99 |
dir_id = self.direction_tokens[direction]
|
| 100 |
-
ctrl_id = self.control_tokens
|
| 101 |
body = self.encode(text)
|
| 102 |
return [1, dir_id, ctrl_id] + body + [2] # 1=BOS, 2=EOS
|
| 103 |
|
|
|
|
| 93 |
ids = [i for i in ids if i not in special]
|
| 94 |
return self.sp_model.decode(ids)
|
| 95 |
|
| 96 |
+
def translate_text(self, text: str, direction: str = "kn2en") -> List[int]:
|
| 97 |
+
"""Build the full HF-style input_ids prefix: [BOS] [DIRECTION] [CONTROL] tokens [EOS]
|
| 98 |
+
|
| 99 |
+
v2.3 ships single-register; the control token slot is fixed to the architectural
|
| 100 |
+
default (NATURAL = id 7). Future versions may surface a register selector.
|
| 101 |
+
"""
|
| 102 |
dir_id = self.direction_tokens[direction]
|
| 103 |
+
ctrl_id = self.control_tokens.get("natural", 7)
|
| 104 |
body = self.encode(text)
|
| 105 |
return [1, dir_id, ctrl_id] + body + [2] # 1=BOS, 2=EOS
|
| 106 |
|