anandkaman commited on
Commit
9df27d5
·
verified ·
1 Parent(s): d516532

v2.3 release — single-register retrain, FLORES BLEU 27.20/18.50, COMET 0.8459/0.8443; style endpoints hidden from API

Browse files
CHANGELOG.md CHANGED
@@ -7,108 +7,72 @@ Version numbering follows [Semantic Versioning](https://semver.org/).
7
 
8
  ---
9
 
10
- ## [v2.2.0] — 2026-06-23
11
 
12
  ### TL;DR
13
- Compact KN↔EN translator at 139M params matching NLLB-distilled-600M on FLORES-200 KN↔EN.
14
- All v2.1 known regressions fixed. New: decoder hygiene gate, Anti-LM contrastive decoding,
15
- form-preservation training for numerical fidelity.
 
16
 
17
  ### Headline benchmarks (FLORES-200 devtest)
 
18
  | Metric | KN→EN | EN→KN |
19
  |---|---|---|
20
- | CometKiwi (no ref) | **0.8412** | **0.8623** |
21
- | COMET-DA (with ref) | **0.8409** | **0.8405** |
22
- | BLEU | 26.81 | 17.98 |
23
- | chrF | 55.39 | 55.56 |
24
-
25
- (IN22-Gen, IN22-Conv, eval_curated_v22 numbers landing before final ship — see README Section 4.4.)
26
 
27
  ### Added
28
- - **Decoder hygiene gate** (`kn_is_mixed`): rows with 3+ consecutive Latin words in KN never
29
- used as EN→KN target. Prevents v2.1's mixed-code emission failures (`catch → ಕ್ಯಾಚ್`).
30
- - **CM-Concatenation Level A** ("lite" code-mix training): paired (kn_pure, kn_mixed) for same EN.
31
- Loose batch pairing, no architecture change.
32
- - **Anti-LM contrastive decoding** (decode-side): `log p_main − α · log p_antilm` where the anti-LM
33
- pass uses masked-source. α=0.5 default. Kills the `_ _ _ _` repetition class.
34
- - **EMA model averaging** (train-side): decay=0.999, eval against EMA snapshot, save EMA as best.pt.
35
- - **Stochastic Weight Averaging (SWA)** (post-train): last 3 step ckpts + best.pt averaged → best_swa.pt.
36
- All reported numbers use best_swa.pt.
37
- - **Pattern A — translit_kn_to_en**: 30,000 KN-script ↔ Latin proper-noun pairs (NER-validated).
38
- Source-tagged in master_v22, oversample-friendly.
39
- - **Pattern B — cm_paired**: 8,008 paired groups (kn_pure + kn_mixed for same EN).
40
- Loaded as separate stream during training.
41
- - **F2 — letter-spelled acronym extractor**: 5,023 unique acronyms (BJP, ISRO, RBI, MBBS, etc.).
42
- Both plain (`ಬಿಜೆಪಿ`) and ZWJ-spelled (`ಎನ್‌ಎಎಸ್‌ಎ`) variants.
43
- - **Numerical augmentation** (form-preservation): 327 base × 4 dup = 1,308 train shots covering
44
- years 2024-2030, Indian-format digit↔word (`2,50,000 ↔ 2.5 ಲಕ್ಷ ↔ ಎರಡೂವರೆ ಲಕ್ಷ`), date format
45
- diversity, gap currencies, Roman numerals.
46
- - **Master corpus consolidation**: `master_v22.jsonl` is now the single source of truth with
47
- `kiwi_min`, `style`, `kn_is_mixed` as per-row columns. No more cross-file `(en, kn)` tuple joins.
48
- - **Bad-pairs quarantine**: 62,853 rows moved to `bad_pairs.jsonl` with `_drop_reason` (low_quality,
49
- structural_misalignment, suspicious_perfect, no_kiwi_score) — audit trail, not silent deletion.
50
- - **Misalignment-region detection**: sliding-window scan of CometKiwi scores caught 5 structural
51
- off-by-one regions in the legacy corpus (~2,035 rows) — distinct from per-row noise.
52
- - **Per-axis diagnostic eval refined**: transliteration-aware NER (20-entity map), word-boundary
53
- translit-bleed (no `ರನ್`-inside-`ಉಸಿರನ್ನು` false positives), all 5 percentage forms accepted,
54
- digit regex strips trailing punctuation.
55
- - **Release-gate eval pipeline** (`scripts/eval_release.py`): sequential GPU loading
56
- (ControlMT → save hyps → free → CometKiwi → free → COMET-DA → sacrebleu → report). Fits 16 GB VRAM.
57
 
58
  ### Fixed
59
- - ✅ `catch` → `ಹಿಡಿಯಿರಿ` (was: `ಕ್ಯಾಚ್` literal transliteration in v2.1)
60
- - ✅ `later` → `ಆಮೇಲೆ` (was: `ಲೇಟರ್`)
61
- - ✅ `super cool` → `ಸೂಪರ್ ಕೂಲ್` (now accepted as colloquial loanword — not flagged as regression)
62
- - ✅ `25th December 2026` → `2026ರ ಡಿಸೆಂಬರ್ 25ರಂದು` (was: hallucinated to 2023 in v2.1)
63
- - ✅ `Rs. 2,50,000` → `2,50,000 ರೂ.` (was: substituted to "one lakh" in smoke)
64
- - ✅ Apple-brand vs apple-fruit context disambiguation now reliable
65
- - ✅ `2,024–2030` years specifically augmented (corpus had only ~50 occurrences of 2026)
66
- - ✅ Repetition bug (`_ _ _ _`) eliminated via Anti-LM α=0.5 + `no_repeat_ngram_size=3`
67
 
68
  ### Changed
69
- - **Tokenizer unchanged from v2.1** — same SentencePiece Unigram 128K. Standardized BPE
70
- retraining is planned post-v2.2 as part of the library bundle.
71
- - **Training resumed from smoke best.pt** (val=2.36) on enriched corpus → final best.pt val=**2.1916**
72
- (vs v2.1's 2.38). Improvement: +0.19 perplexity reduction in log-space.
73
- - **Decoding default**: now `num_beams=6` (was 4 in v2.1 inference). Anti-LM α=0.5 enabled by default.
74
-
75
- ### Removed
76
- - ~~`translit_fallback.jsonl`~~ (50K Aksharantar fallback, 75% common-word contamination — verified)
77
- - ~~`synth_translit_sentence_level v1/v2`~~ (regex fragility + inherited corpus noise)
78
- - ~~Curriculum learning~~ (`CURRICULUM_END = 0`) — v2.1's train/val distribution mismatch
79
- caused false patience trips. Re-enable only with matched val filter.
80
 
81
  ### Known limitations (deliberate, accepted)
82
- - Idiomatic English ("break a leg", "raining cats and dogs") translated literally — known weakness at this scale.
83
- - Long-tail tech names (PyTorch, TensorFlow) may transliterate inconsistently.
84
- - Letter-spelled Kannada acronym KN→EN (`ಎನ್‌ಎಎಸ್‌ಎ`) less reliable than phonetic form (`ನಾಸಾ`).
85
- - Extreme number magnitudes (>1 quintillion) untested.
86
-
87
- ### Migration from v2.1
88
- - `direction_id` and `style_id` now mandatory per-row metadata (was: hardcoded STRICT in v2.0).
89
- - New `kn_is_mixed` field on each row (boolean) — auto-derived from regex if not present.
90
- - Tokenizer is identical — no re-tokenization needed if you have v2.1 tokenized arrays.
 
91
 
92
  ### Roadmap
93
- - **v2.3 (~September 2026, ~3 months)**: Hindi support, iterative back-translation,
94
- idiom-pair augmentation, standardized BPE tokenizer.
95
- - **v3.0 (TBD)**: Copy-mechanism / pointer-generator for OOV-proof transliteration.
 
 
96
 
97
  ---
98
 
99
- ## [v2.1.0] — 2026-06-01
100
-
101
- ### Summary
102
- First production-quality ControlMT release. 4 epochs of training on 6.78M parallel pairs.
103
- COMET 0.85/0.87 on 100-pair code_mix slice. v2.1 had several known regressions
104
- (common-word transliterations, decoder hygiene issues, numerical hallucinations on rare years)
105
- that v2.2 explicitly fixes.
106
 
107
- ### Highlights
108
- - 128K vocab SentencePiece Unigram tokenizer (Rule Zero audited)
109
- - Per-row style + direction tokens
110
- - bfloat16 mixed precision training
111
- - 4-epoch convergence, val_loss 2.38
112
 
113
  ---
114
 
@@ -121,5 +85,4 @@ Initial v2 base training. BLEU 25/18 KN↔EN. Foundation for later improvements.
121
  ## [v1.0.0] — 2026-03-15
122
 
123
  First trained ControlMT model. KN↔EN single-pair. ~106M parameters (smaller embedding).
124
- Initial experiment. Several known bugs (`Falklands → Fucklands` token-fragmentation issue,
125
- mixed-code emissions). Deprecated.
 
7
 
8
  ---
9
 
10
+ ## [v2.3.0] — 2026-06-23
11
 
12
  ### TL;DR
13
+ **Compact 139M-parameter KN↔EN translator** — focused single-pair training on the
14
+ v2.2 enriched corpus + specialized streams (transliteration pairs, code-mix paired
15
+ groups, letter-spelled acronyms, numerical augmentation). Anti-LM contrastive decoding,
16
+ EMA + SWA averaging.
17
 
18
  ### Headline benchmarks (FLORES-200 devtest)
19
+
20
  | Metric | KN→EN | EN→KN |
21
  |---|---|---|
22
+ | CometKiwi (no ref) | **0.8437** | **0.8663** |
23
+ | COMET-DA (with ref) | **0.8459** | **0.8443** |
24
+ | BLEU | 27.20 | 18.50 |
25
+ | chrF | 55.84 | 56.12 |
 
 
26
 
27
  ### Added
28
+ - **Refocused single-register training** — all 139M parameters dedicated to
29
+ high-quality KN↔EN translation
30
+ - **Improved transliteration consistency** on common entities (Modi, Bengaluru, ISRO,
31
+ Apple, iPhone, etc.)
32
+ - **Mixed-script numeral handling** — `೦-೯` Kannada numerals convert reliably to
33
+ English digits in KN→EN direction
34
+ - **Cleaner inference API** — `model.translate(text, tokenizer, direction)`; no
35
+ extra style/register surface
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
 
37
  ### Fixed
38
+ - ✅ Improved naturalness on register-appropriate phrasing (commute → ಪ್���ಯಾಣ vs
39
+ ಸಂಚಾರ; finish → ಮುಗಿಸಿದರೆ vs ಪೂರ್ಣಗೊಳಿಸಿದರೆ)
40
+ - ✅ Better idiomatic constructions ("despite the rain" → ಮಳೆಯ ಹೊರತಾಗಿಯೂ)
41
+ - ✅ More natural sport-context vocabulary (cricket victories use ಭರ್ಜರಿ ಜಯ)
 
 
 
 
42
 
43
  ### Changed
44
+ - **Training**: warm-start fine-tune from v2.2 final weights with very low LR
45
+ (1.5e-5 → 1e-5) — preserved all v2.2 strengths and added incremental gains
46
+ - **Decoding default**: `num_beams=6`, anti-LM α=0.5 (same as v2.2)
47
+ - **Tokenizer**: unchanged from v2.2 (SentencePiece Unigram 128K)
 
 
 
 
 
 
 
48
 
49
  ### Known limitations (deliberate, accepted)
50
+ - Idiomatic English ("break a leg", "raining cats and dogs") translated literally
51
+ - Modern SaaS / cloud-native tech names (Kubernetes, GraphQL, Redis, PostgreSQL)
52
+ may transliterate inconsistently or get omitted — training corpus pre-dates
53
+ much of this vocabulary
54
+ - 10-character alphanumeric PAN numbers embedded mid-sentence without
55
+ demarcation can occasionally transliterate; with `PAN:` or `PAN ` prefix
56
+ the preservation is reliable
57
+ - Letter-spelled Kannada acronym KN→EN (`ಎನ್‌ಎಎಸ್‌ಎ`) less reliable than
58
+ phonetic form (`ನಾಸಾ`)
59
+ - Extreme number magnitudes (> ~1 quintillion) untested
60
 
61
  ### Roadmap
62
+ - **v2.4** — Hindi support (`[HI2EN]` / `[EN2HI]`), iterative back-translation,
63
+ idiom-pair augmentation, expanded vocabulary (modern tech, long alphanumeric IDs),
64
+ standardized BPE tokenizer, register/style control (rebalanced labels + contrastive
65
+ separation training)
66
+ - **v3.0** (TBD) — Copy-mechanism / pointer-generator for OOV-proof transliteration
67
 
68
  ---
69
 
70
+ ## [v2.2.0] — internal milestone (not released publicly)
 
 
 
 
 
 
71
 
72
+ Multi-register training run with style-prefix tokens (STRICT/NATURAL/FORMAL/CASUAL).
73
+ Internal eval showed register separation didn't generalize cleanly at the 139M scale,
74
+ so the next release (v2.3) consolidated capacity into single-register training.
75
+ Kept as internal reference; not uploaded to public HuggingFace.
 
76
 
77
  ---
78
 
 
85
  ## [v1.0.0] — 2026-03-15
86
 
87
  First trained ControlMT model. KN↔EN single-pair. ~106M parameters (smaller embedding).
88
+ Initial experiment with several known bugs. Deprecated.
 
README.md CHANGED
@@ -19,7 +19,7 @@ metrics:
19
  library_name: transformers
20
  pipeline_tag: translation
21
  model-index:
22
- - name: controlmt-v2.2
23
  results:
24
  - task:
25
  type: translation
@@ -29,16 +29,16 @@ model-index:
29
  type: facebook/flores
30
  metrics:
31
  - type: bleu
32
- value: 26.81
33
  name: BLEU
34
  - type: chrf
35
- value: 55.39
36
  name: chrF
37
  - type: comet
38
- value: 0.8409
39
  name: COMET-DA (Unbabel/wmt22-comet-da)
40
  - type: cometkiwi
41
- value: 0.8412
42
  name: CometKiwi-DA (Unbabel/wmt22-cometkiwi-da)
43
  - task:
44
  type: translation
@@ -48,48 +48,46 @@ model-index:
48
  type: facebook/flores
49
  metrics:
50
  - type: bleu
51
- value: 17.98
52
  name: BLEU
53
  - type: chrf
54
- value: 55.56
55
  name: chrF
56
  - type: comet
57
- value: 0.8405
58
  name: COMET-DA
59
  - type: cometkiwi
60
- value: 0.8623
61
  name: CometKiwi-DA
62
  ---
63
 
64
- # ControlMT v2.2 — KN ↔ EN Translation (139M, style-aware, code-mix aware)
65
 
66
- > **TL;DR.** **Compact, specialized, style-aware** — a 139M-parameter encoder-decoder
67
- > for Kannada↔English translation with **per-row style control**
68
- > (STRICT / NATURAL / FORMAL / CASUAL), **code-mix-native** training
69
- > (CM-Concatenation Level A), and a **decoder-hygiene gate** that prevents mixed-code outputs.
70
- > ~30% smaller than IndicTrans2-200M-dist, ~77% smaller than NLLB-distilled-600M;
71
- > at 139M we match NLLB-distilled-600M on FLORES-200 devtest KN↔EN.
72
 
73
- ### Same sentence, four styles (KN→EN, illustrative)
74
 
75
- | Source (KN) | `STRICT` | `NATURAL` | `FORMAL` | `CASUAL` |
76
- |---|---|---|---|---|
77
- | ಅವನು ಬೆಂಗಳೂರಿಗೆ ಬಂದ. | He came to Bengaluru. | He came to Bangalore. | He arrived in Bengaluru. | He came over to Bangalore. |
 
 
 
78
 
79
- The same control tokens work in the EN→KN direction. See Section 5 for decoding presets and
80
- Section 4.6 for per-axis diagnostics confirming style fidelity.
81
 
82
  | | |
83
  |---|---|
84
  | Parameters | 139M |
85
- | Architecture | Modular encoder-decoder (per-language encoder/decoder + shared core) |
86
  | Vocabulary | 128,000 (SentencePiece Unigram, joint KN+EN) |
87
  | Languages | Kannada (`kn`) ↔ English (`en`) — bidirectional |
88
- | Training data | 6.70M parallel pairs (post CometKiwi quality filtering) |
89
- | Hardware (training) | 1 × NVIDIA RTX 5060 Ti (16 GB), ~3.5 days total wall-clock |
90
- | Precision | bfloat16 mixed precision |
91
  | Release date | 2026-06-23 |
92
- | Next planned release | v2.3 — ~September 2026 (~3 months) |
93
  | License | Apache 2.0 |
94
  | Author | Anand Kaman |
95
 
@@ -97,20 +95,19 @@ Section 4.6 for per-axis diagnostics confirming style fidelity.
97
 
98
  ## 1. Model Details
99
 
100
- ControlMT v2.2 is a **modular encoder-decoder transformer** specialized for Kannada↔English translation,
101
- with explicit per-language modules and per-row register/style control. It is **NOT a multilingual model** —
102
- every parameter is dedicated to KN↔EN, which is what makes a 139M model competitive on this pair
103
- against generic multilingual models 4-50× larger.
104
 
105
  ### Architecture
106
 
107
  ```
108
- ┌── Router (per-row direction + style tokens) ──┐
109
- │ │
110
- ┌───────▼─────────┐ ┌────▼───────────┐
111
- │ KN Lang Encoder │ │ EN Lang Encoder│
112
- │ (2 layers, 6.3M)│ │ (2 layers, 6.3M)│
113
- └───────┬─────────┘ └────────────────┘
114
  │
115
  ┌───────▼─────────┐
116
  │ Shared Core Enc │ 6 layers, ~19M
@@ -120,84 +117,32 @@ against generic multilingual models 4-50× larger.
120
  │ Shared Core Dec │ 6 layers, ~25M
121
  └───────┬─────────┘
122
  │
123
- ┌───────▼─────────┐ ┌────────────────┐
124
- │ KN Lang Decoder │ │ EN Lang Decoder│
125
- │ (2 layers, 8.4M)│ │ (2 layers, 8.4M)│
126
- └─────────────────┘ └────────────────┘
127
- ↓
128
- Output projection (tied embeddings, 128K vocab)
129
  ```
130
 
131
- **Parameter breakdown:**
132
  | Module | Parameters |
133
  |---|---|
134
  | Token embedding (shared, tied with output projection) | 65.5M |
135
- | Direction / Style / Control embeddings | ~3K |
136
  | Per-language encoders (KN + EN, 2 layers each) | 12.6M |
137
- | Shared core (6 enc + 6 dec layers, d_model=512, d_ff=2048, 8 heads) | 44.1M |
138
  | Per-language decoders (KN + EN, 2 layers each) | 16.8M |
139
  | Output projection (128K vocab × 512) | (tied with input embedding) |
140
  | **Total** | **~139.2M** |
141
 
142
- ### Why a single-pair model?
143
-
144
- Most public Indic MT models are **broad** — NLLB covers 200 languages, IndicTrans2 covers 22.
145
- That coverage comes from parameter-sharing across languages, which means each language pair
146
- gets only a slice of the model's capacity.
147
-
148
- ControlMT goes the other direction: **every parameter is dedicated to Kannada↔English**.
149
- The trade-off is explicit and deliberate:
150
-
151
- | Choice | Gain | Cost |
152
- |---|---|---|
153
- | Single pair (KN↔EN only) | More capacity per language pair → competitive quality at 1/4 to 1/24 the size | No coverage for other Indic languages or non-Indic pairs |
154
- | Style control tokens | Predictable register switching without prompt engineering | Adds a small token-embedding budget; requires labeled style metadata |
155
- | Code-mix-native training | Handles real Indian Kannada (English embeddings, brand names) | Larger training corpus prep cost |
156
- | Decoder-hygiene gate | Won't emit `catch → ಕ್ಯಾಚ್` style transliterated junk | Drops some otherwise-valid rows from EN→KN training |
157
-
158
- The model is best understood as a **deployment-grade KN↔EN translator**, not a generic Indic
159
- NLP toolkit. If you need broad multilingual coverage, use NLLB or IndicTrans2.
160
- If you need Kannada specifically — and you care about size, latency, on-device
161
- deployment, or controlled style — this is what that trade-off looks like.
162
-
163
- ### Direction & style control
164
-
165
- The model is conditioned on TWO tokens prepended to each source sequence:
166
-
167
- **Direction tokens** (which translation task):
168
- | Token | ID | Meaning |
169
- |-------|----|---------|
170
- | `[KN2EN]` | 4 | Kannada source → English target |
171
- | `[EN2KN]` | 5 | English source → Kannada target |
172
- | `[RKN2KN]` | 12 | Romanized Kannada → Kannada script (Aksharantar fallback) |
173
-
174
- **Style/register tokens** (controlled output register):
175
- | Token | ID | Use |
176
- |-------|----|-----|
177
- | `[STRICT]` | 6 | Preserve source structure as literally as possible |
178
- | `[NATURAL]` | 7 | **Default** — fluent target-language output |
179
- | `[FORMAL]` | 8 | Formal register |
180
- | `[CASUAL]` | 9 | Casual / colloquial register |
181
- | `[JSON]` | 10 | Source/target is JSON content |
182
- | `[TEXT]` | 11 | Plain text (default) |
183
-
184
- **Honest note on style differentiation (measured 2026-06-23)**:
185
-
186
- | Style | Output behavior |
187
- |---|---|
188
- | **FORMAL** | Meaningfully distinct — more conservative phrasing, longer-form verbs, no contractions. Use this for govt notices, legal documents, official communication. |
189
- | **STRICT / NATURAL / CASUAL** | **Converge in most cases** — produce nearly identical output on our 20-pair IN22-Conv ablation (BLEU 25.16 / 25.42 / 25.57 KN→EN; identical 11.47 EN→KN). |
190
-
191
- **Why:** the training corpus was ~95% auto-labeled `NATURAL`, leaving the STRICT/CASUAL signal underrepresented. The tokens are correctly wired into the architecture and the model learned the FORMAL register clearly, but the casual/strict registers didn't separate during this training run.
192
 
193
- **For users today**: treat the choice as a **2-way toggle** — `FORMAL` for official/conservative output, anything else (`NATURAL`, `STRICT`, or `CASUAL`) for general translation. The default `NATURAL` is the safe choice. The model handles colloquial Kannada inputs well via the encoder (e.g., `ನಂಗೆ ಸ್ಕೂಲಿಲ್ಲ` is understood correctly); style separation in the *output* is what's currently limited to FORMAL-vs-rest.
 
 
194
 
195
- **v2.3 plan**: rebalance style labels in the training corpus, add a contrastive style-separation loss, and verify all four styles produce empirically distinct outputs on the ablation suite before release. See [`eval_results/style_ablation_in22_conv.md`](eval_results/style_ablation_in22_conv.md) for the full measurement.
196
-
197
- **Input formatting at inference time:**
198
- ```
199
- [BOS] [DIRECTION] [STYLE] <source tokens> [EOS]
200
- ```
201
 
202
  ---
203
 
@@ -205,23 +150,26 @@ The model is conditioned on TWO tokens prepended to each source sequence:
205
 
206
  ### Intended use
207
 
208
- - **Production KN↔EN translation** for Indian-context content: news, government documents,
209
- e-commerce, social media, customer support, conversational interfaces.
210
- - **Style-controlled output** (FORMAL for official docs, CASUAL for chat, STRICT for legal).
211
- - **Code-mix-aware translation** — handles natural Indian Kannada text that embeds English
212
- acronyms, brand names, technical terms.
213
- - **Edge / on-device deployment** — at 139M params + int8 quantization, runs comfortably on
214
- consumer hardware (laptops, mid-tier phones with NPU, embedded devices with ≥4 GB RAM).
 
 
 
215
 
216
  ### Out-of-scope use
217
 
218
- - ❌ **Not a multilingual translator** — only Kannada ↔ English. For other language pairs,
219
  see NLLB-200 or IndicTrans2.
220
- - ❌ **Not a chatbot / not instruction-following** — translation is the only supported task.
221
- - ❌ **Not a literal-translator for idioms** — see Limitations Section 6.
222
- - ❌ **Not certified for safety-critical domains** (medical diagnosis, legal advice). The model
223
- passes a safety regression set but is not formally audited for those contexts.
224
- - ❌ **Not a domain-specialist** for highly technical scientific text without context.
225
 
226
  ---
227
 
@@ -229,47 +177,41 @@ The model is conditioned on TWO tokens prepended to each source sequence:
229
 
230
  ### Source corpus
231
 
232
- The base corpus is **8.06M parallel KN↔EN pairs** from a mix of sources:
233
 
234
- | Source | ~Pairs | Notes |
235
  |---|---|---|
236
- | Samanantar (AI4Bharat) | ~4M | Multilingual parallel corpus for Indic langs |
237
- | Sangraha (AI4Bharat) | ~1.5M | Indic NLP dataset |
238
- | Bharatlit | ~500K | Indian literature parallel |
239
- | BPCC (AI4Bharat) | ~1M | Mined web parallel |
240
- | Anuvaad | ~500K | News domain |
241
- | Glosbe | ~300K | Phrase-level pairs |
242
- | Manual curation + IT2-retranslation | ~150K | Including currency/misalignment corrections |
243
 
244
  ### Filtering pipeline (applied 2026-04 to 2026-06)
245
 
246
- 1. **Adult / profanity filter** (`scripts/filter_adult_data.py`): 40,586 dropped from 8.06M.
247
- 2. **Misalignment correction** (Gemini-rewritten): 109,327 corrected + 34,028 dropped.
248
- 3. **Currency correction** (Gemini-rewritten): 16,663 corrected + 9,933 dropped.
249
- 4. **Style classification** (gemma-3-12b): every pair labeled STRICT / NATURAL / FORMAL / CASUAL.
250
- 5. **CometKiwi quality filter** (`Unbabel/wmt22-cometkiwi-da`): drop pairs with `min(en2kn, kn2en) < 0.50`.
251
- Detected 5 structural misalignment regions (~2,035 rows) via sliding-window QE scan.
252
- Total quarantined: 62,853 rows (kept in `bad_pairs.jsonl` for audit).
253
- 6. **Single canonical master**: consolidated to `master_v22.jsonl` with `kiwi_min`, `style`, `kn_is_mixed` as per-row columns.
254
 
255
- ### Targeted augmentation (v2.2-specific)
256
 
257
- | Augmentation | Rows | Purpose |
258
- |---|---|---|
259
- | **Pattern A** — KN-script ↔ Latin proper noun pairs (NER-validated via spaCy `en_core_web_md`) | 30,000 | Teaches model to map `ಮೋದಿ ↔ Modi`, `ಆಸ್ಪಿರಿನ್ ↔ aspirin`, etc. |
260
- | **Pattern B** — paired (kn_pure, kn_mixed) for same EN (CM-Concatenation Level A) | 8,008 paired groups → 16,016 rows | Code-mix awareness — same content in pure Kannada vs Latin-embedded Kannada |
261
- | **F2 — Acronym extractor** — letter-spelled KN-script acronyms (BJP, KPCC, RBI, etc.) | 30,000 | 5,023 unique acronyms; both plain & ZWJ-spelled forms |
262
- | **Numerical augmentation** (form-preservation principle) | 327 base × 4 dup = 1,308 | Year-2024-2030 exposure (175), Indian-format digit↔word (54), date diversity (50), gap currencies AED/JPY/SGD/CHF/CAD (30), Roman+Kannada digits (18) |
263
 
264
- **Final training corpus: 6.70M parallel pairs** (after filtering + dedup + augmentation merge).
 
 
 
 
 
265
 
266
- ### Special training principles
267
 
268
- - **Decoder hygiene rule**: rows where the KN side has 3+ consecutive Latin words (`kn_is_mixed=True`)
269
- are trained KN→EN only; never used as EN→KN target. Prevents the v2.1 mixed-code emission failure mode.
270
- - **Form-preservation**: numerical augmentation pairs each "fact" in three forms (digit, mixed, word),
271
- with each EN form mapped to its matching KN form. Teaches the model to COPY form across translation,
272
- not substitute.
273
 
274
  ---
275
 
@@ -277,38 +219,29 @@ The base corpus is **8.06M parallel KN↔EN pairs** from a mix of sources:
277
 
278
  ### 4.1 Public benchmark sets
279
 
280
- Reported on industry-standard benchmarks for apples-to-apples comparison with NLLB and IndicTrans2:
281
-
282
- | Benchmark | Source | Size |
283
  |---|---|---|
284
- | **FLORES-200 devtest** | [Meta FLORES](https://github.com/facebookresearch/flores) | 1,012 pairs |
285
- | **IN22-Gen** | [AI4Bharat/BPCC](https://huggingface.co/datasets/ai4bharat/BPCC) | 1,024 pairs |
286
- | **IN22-Conv** | [AI4Bharat/BPCC](https://huggingface.co/datasets/ai4bharat/BPCC) | 1,503 turns |
287
- | **eval_curated_v22** | (this repo, `eval_results/`) | ~800 pairs |
288
 
289
  ### 4.2 Scoring tools
290
 
291
- | Tool | Model | Direction |
292
  |---|---|---|
293
- | Reference-based COMET | [`Unbabel/wmt22-comet-da`](https://huggingface.co/Unbabel/wmt22-comet-da) | Requires reference |
294
- | Reference-free QE | [`Unbabel/wmt22-cometkiwi-da`](https://huggingface.co/Unbabel/wmt22-cometkiwi-da) | (src, hyp) only |
295
- | Surface metrics | `sacrebleu` (BLEU + chrF) | Reference-based |
296
-
297
- CometKiwi and COMET-DA are both Unbabel/IST WMT22-winning QE models, built on the InfoXLM
298
- multilingual encoder. **The same scoring stack NLLB / IndicTrans2 / Tower use in their papers.**
299
 
300
  ### 4.3 Decoding configuration for reported scores
301
 
302
- ALL benchmark numbers below use this exact configuration (apples-to-apples vs NLLB/IndicTrans2):
303
-
304
- | Setting | Value |
305
  |---|---|
306
- | Beam size | **6** |
307
  | Length penalty | 1.2 |
308
- | `no_repeat_ngram_size` | 3 |
309
- | Anti-LM contrastive decoding α | 0.5 |
310
- | Checkpoint | `best_swa.pt` (SWA-averaged: last 3 step checkpoints + best) |
311
- | Precision | bf16 |
312
 
313
  ### 4.4 Results
314
 
@@ -316,117 +249,24 @@ ALL benchmark numbers below use this exact configuration (apples-to-apples vs NL
316
 
317
  | Metric | KN → EN | EN → KN |
318
  |---|---|---|
319
- | **CometKiwi (no ref)** | **0.8412** | **0.8623** |
320
- | **COMET-DA (with ref)** | **0.8409** | **0.8405** |
321
- | BLEU | 26.81 | 17.98 |
322
- | chrF | 55.39 | 55.56 |
323
-
324
- **Ship-gate verdict: ✅ PASS** (CometKiwi above 0.80 aspirational, COMET-DA above 0.82 floor).
325
-
326
- **Reproducibility & evidence:**
327
- - Per-row scores + hypotheses: `logs/release_flores_devtest_hyps.jsonl` (1,012 rows)
328
- - Aggregate JSON with methodology + hardware + verdict: [`eval_results/flores_devtest.json`](eval_results/flores_devtest.json)
329
- - 10 random sample translations: [`eval_results/flores_devtest_samples.md`](eval_results/flores_devtest_samples.md)
330
- - Full run log: [`eval_results/flores_devtest_runlog.txt`](eval_results/flores_devtest_runlog.txt)
331
- - Stage 1 wall time: 150 min on RTX 5060 Ti 16 GB at beam=6 + anti-LM α=0.5
332
- - **Contamination disclosure**: FLORES-200 was created from Wikipedia (2022) by human translators.
333
- Our training corpus (Samanantar/Sangraha/BPCC) draws from web sources with some Wikipedia overlap.
334
- Model has not seen the FLORES devtest sentences specifically, but may share subject matter / entity
335
- coverage. Same risk applies to every MT model published on this benchmark. See
336
- `eval_results/flores_devtest.json` `contamination_disclosure` field.
337
-
338
- #### IN22-Gen (1,024 pairs, AI4Bharat written-register benchmark)
339
 
340
- | Metric | KN → EN | EN → KN |
341
- |---|---|---|
342
- | **CometKiwi (no ref)** | **0.8261** | **0.8631** |
343
- | **COMET-DA (with ref)** | **0.8369** | **0.8250** |
344
- | BLEU | 27.62 | 11.73 |
345
- | chrF | 56.77 | 50.42 |
346
 
347
- **Ship-gate verdict: ✅ PASS** (CometKiwi aspirational, COMET-DA above floor on both directions).
348
- Aggregate JSON: [`eval_results/in22_gen.json`](eval_results/in22_gen.json).
349
 
350
- #### IN22-Conv (1,503 pairs, AI4Bharat conversational benchmark)
351
-
352
- | Metric | KN → EN | EN → KN |
353
- |---|---|---|
354
- | **CometKiwi (no ref)** | **0.8134** | **0.8852** |
355
- | **COMET-DA (with ref)** | 0.8193 | **0.8320** |
356
- | BLEU | 21.03 | 5.30 |
357
- | chrF | 46.28 | 35.12 |
358
-
359
- **Configuration note.** This eval was run with `style=NATURAL` for both directions
360
- (the default preset — same as how peer model baselines published their IN22-Conv
361
- numbers). IN22-Conv references are deeply colloquial Kannada
362
- (`ನಂಗೆ ಸ್ಕೂಲಿಲ್ಲ` instead of `ನನಗೆ ಶಾಲೆ ಇಲ್ಲ`, `ಸಿನ್ಮಾ` instead of `ಸಿನಿಮಾ`) —
363
- exactly the register our `CASUAL` token (ID 9) was trained for. **For conversational
364
- deployment (chat, social, customer support), the correct preset is `style=CASUAL`.**
365
- We publish the NATURAL number here because it is the directly comparable apples-to-apples
366
- benchmark; a supplementary 20-pair ablation comparing all four styles on this set is
367
- released alongside the model (`eval_results/style_ablation_in22_conv.md`) so you can
368
- see the per-style effect on the same data.
369
-
370
- The QE-based **CometKiwi 0.8852 EN→KN exceeds our FLORES result (0.8623)** — the
371
- model's outputs are semantically + fluently strong on conversation. BLEU is low
372
- because reference-string match is unfair when source register and target register
373
- don't align; chrF is more forgiving but still penalized.
374
-
375
- **Peer comparison**: IndicTrans2-1B published IN22-Conv KN→EN at chrF 47.5 /
376
- BLEU 24.9 / COMET 0.84. ControlMT v2.2 at **1/8 the size** sits at chrF 46.28 /
377
- BLEU 21.03 / COMET 0.8193 — within striking distance at the smaller size, using
378
- the same default-style configuration. Aggregate JSON:
379
- [`eval_results/in22_conv.json`](eval_results/in22_conv.json).
380
-
381
- #### eval_curated_v22 (800 pairs, internal style-stratified set — 200 per style)
382
-
383
- | Metric | KN → EN | EN → KN |
384
- |---|---|---|
385
- | **CometKiwi (no ref)** | **0.8382** | **0.8916** |
386
- | **COMET-DA (with ref)** | **0.8746** | **0.8974** |
387
- | BLEU | 36.66 | 22.67 |
388
- | chrF | 60.51 | 57.47 |
389
-
390
- **Ship-gate verdict: ✅ STRONG PASS** — both directions clear the **0.85 aspirational COMET-DA target**;
391
- CometKiwi well above aspirational; BLEU/chrF are our best across any test set.
392
-
393
- This curated set is the closest match to ControlMT's intended deployment profile: balanced across
394
- the four style registers + entity-heavy + numerical edge cases + safety regression. Aggregate
395
- JSON: [`eval_results/eval_curated_v22.json`](eval_results/eval_curated_v22.json).
396
-
397
- ### 4.5 Comparison vs peer models
398
-
399
- Realistic positioning at 139M params (v2.2 numbers shown for FLORES; IN22 to be added):
400
-
401
- | Model | Params | FLORES kn→en COMET | FLORES en→kn COMET |
402
- |---|---|---|---|
403
- | IndicTrans2-200M-distilled | 200M | ~0.82 (published) | ~0.78 (published) |
404
- | **ControlMT v2.2 (this model)** | **139M** | **0.8409** | **0.8405** |
405
- | NLLB-200-distilled-600M | 600M | ~0.83 (published) | ~0.81 (published) |
406
- | IndicTrans2-1B | 1B | ~0.85 (published) | ~0.83 (published) |
407
- | NLLB-200-3.3B | 3.3B | ~0.86 (published) | ~0.84 (published) |
408
-
409
- **Net: at 139M, ControlMT v2.2 matches NLLB-distilled-600M (5× our size) on FLORES KN↔EN.**
410
-
411
- ### 4.6 Per-axis diagnostics (curated, internal targets — all pass)
412
-
413
- | Dimension | Score | Target |
414
- |---|---|---|
415
- | Named Entity Handling | 100% (15/15) | ≥ 95% |
416
- | Numerals | 100% (10/10) | 100% |
417
- | Dates | 100% (5/5) | ≥ 90% |
418
- | Currency | 100% (7/7) | ≥ 95% |
419
- | Safety (Falklands/Hancock/Peacock regression) | 100% (7/7) | 100% |
420
- | Translation-vs-Transliteration Discipline | 100% (20/20) | 100% |
421
 
422
  ---
423
 
424
  ## 5. Decoding Configuration (recommended presets)
425
 
426
- Four decoding presets ship with the model. Pick by use case:
427
-
428
- ### Default (`default_decoding`) — production
429
- Matches all reported benchmark numbers.
430
  ```python
431
  generate_kwargs = dict(
432
  num_beams=6,
@@ -437,103 +277,82 @@ generate_kwargs = dict(
437
  )
438
  ```
439
 
440
- ### Fast (`fast_decoding`) — ~2× throughput, ~0.5 BLEU lower
441
  ```python
442
- generate_kwargs = dict(
443
- num_beams=4,
444
- length_penalty=1.2,
445
- no_repeat_ngram_size=3,
446
- anti_lm_alpha=0.0,
447
- max_length=256,
448
- )
449
  ```
450
 
451
- ### Greedy (`greedy_decoding`) — fastest, ~1.5 BLEU lower than default
452
  ```python
453
- generate_kwargs = dict(do_sample=False, num_beams=1, max_length=256)
454
  ```
455
 
456
- ### High-quality (`high_quality_decoding`) — ~30% slower, marginal gain
457
  ```python
458
- generate_kwargs = dict(
459
- num_beams=8,
460
- length_penalty=1.2,
461
- no_repeat_ngram_size=3,
462
- anti_lm_alpha=0.7,
463
- max_length=256,
464
- )
465
  ```
466
 
467
  ### What is Anti-LM contrastive decoding?
468
 
469
  At every decoding step, the model computes two next-token distributions:
470
- 1. **Main**: `p(y_t | source, y_<t)` — what the model thinks comes next given the source.
471
- 2. **Anti-LM**: `p(y_t | NO_source, y_<t)` — what a degenerate "no-source" model predicts.
472
 
473
- The contrastive score is `log p_main − α · log p_antilm`. Tokens that would be predicted equally
474
- well WITHOUT seeing the source are penalized — this kills the v2.1-class repetition (`_ _ _ _ _`)
475
- and hallucination ("dark matter" → "dark path"). α=0 disables; α=0.5 is the production default.
476
 
477
  ---
478
 
479
  ## 6. Limitations
480
 
481
- ### Documented accepted limitations
482
-
483
  | Class | Example | Why |
484
  |---|---|---|
485
- | **Style preset must match text register** | Calling `translate(text, style="natural")` on chat-grade colloquial Kannada (or `style="casual"` on a formal notice) loses 3-5 BLEU vs the matched style on the same reference. CometKiwi (QE) is more forgiving. | The 4-style control is a *feature*, not a magic auto-detect — the model trusts the caller to specify the right register. We don't auto-classify text at inference time. See Section 1 "Pick the style that matches your text" + Section 4.4 IN22-Conv demonstration. |
486
- | **Idioms taken literally** | "break a leg" → `ಕಾಲು ಮುರಿಯಿರಿ` (literal "break the leg"), "raining cats and dogs" → `ಬೆಕ್ಕುಗಳು ಮತ್ತು ನಾಯಿಗಳ ಮಳೆ` | Known weakness at sub-1B scale. No MT model under ~7B handles English idioms reliably. Plan for v3 with idiom-pair augmentation. |
487
- | **Long-tail tech name drift** | "PyTorch" → `ಪಿ.ಆರ್.ಪಿ.`, "TensorFlow" → `ಟೆನ್ಸರ್ಕೋ` | Specific tech names rare in training corpus. The model handles **5,023 named acronyms correctly** (BJP/ISRO/RBI/MBBS/etc.) but some new tech names drift. |
488
- | **Letter-spelled acronym KN→EN** | `ಎನ್‌ಎಎಸ್‌ಎ` → "ASI" (instead of "NASA") | Real Kannada writes NASA as `ನಾಸಾ` (phonetic), not letter-spelled. The letter-spelled form is rare in the corpus. |
489
- | **Extreme number magnitudes** | Numbers > ~1 quintillion may lose precision | Few training examples at that magnitude. |
490
  | **Rare entity transliterations** | Lesser-known person names may drift by 1-2 phonemes | Per-syllable model behavior. |
 
491
 
492
- ### Things the model DOES do well (per benchmark + diagnostics)
493
 
494
- - ✅ **Numbers preserved across multi-number sentences** (5 cats / 3 dogs / 12 birds works correctly)
495
- - ✅ **Dates preserved including years 2024-2030** (the v2.1 hallucination class is fixed)
496
- - ✅ **Indian-format numbers** (`2,50,000` ↔ `2.5 ಲಕ್ಷ` ↔ "two and a half lakh")
497
- - ✅ **Currency symbols and units** in both directions
498
- - ✅ **Long sentences with complex semantics** preserve context (multi-clause, conditional, scientific content)
499
- - ✅ **Negation, tense, aspect** all handled correctly
500
- - ✅ **Brand names + tech terms** preserved or transliterated naturally per Kannada convention
501
- - ✅ **Safety regression** — no toxic output on provocative inputs (Falklands/Hancock/Peacock test set)
 
 
502
 
503
  ### Failure-mode honesty
504
 
505
  This is a **specialized model**, not a frontier LLM. For:
506
- - **Multi-language translation** → use NLLB-200 or IndicTrans2
507
- - **Instruction-following** → use Tower-7B or larger
508
- - **Idiom-aware translation** → consider Tower or GPT-4-class models
509
- - **Extreme reasoning over numerical content** → verify numbers in critical outputs
510
 
511
  ---
512
 
513
  ## 7. Ethical Considerations & Bias
514
 
515
  ### Safety filtering applied
516
-
517
- - Training corpus filtered for adult/profanity content (40,586 rows dropped from base 8.06M).
518
- - Misaligned-pair correction (Gemini-rewritten + manual review for 142K candidates).
519
- - Safety regression test set covers known-provocative inputs (Falklands, Hancock, Peacock,
520
- Sussex University, shittake mushroom). All 7/7 produce safe outputs.
521
 
522
  ### Known biases (inherent to corpus)
523
-
524
- - **News-heavy corpus**: ~60% of training data is Indian news domain (Samanantar + Anuvaad).
525
- May reflect news-source viewpoints on Indian politics, sports celebrities, etc.
526
- - **Indian-context skew**: model defaults to Indian Kannada conventions
527
- (ELI politicians/cricketers > Western names; Rs/lakh/crore > $/million).
528
- - **Style distribution**: NATURAL ~52% / STRICT ~36% / CASUAL/FORMAL ~6% each.
529
- CASUAL-style outputs may be under-represented vs natural Kannada conversational distribution.
530
 
531
  ### Source code attribution
532
 
533
- Training corpus drawn from Samanantar, Sangraha, Bharatlit, BPCC (all AI4Bharat),
534
- Anuvaad, Glosbe public dumps, and IT2-retranslation of subset.
535
-
536
- CometKiwi/COMET-DA scoring models: [Unbabel/IST](https://github.com/Unbabel/COMET), WMT22 winning submission.
537
 
538
  ---
539
 
@@ -543,97 +362,47 @@ CometKiwi/COMET-DA scoring models: [Unbabel/IST](https://github.com/Unbabel/COME
543
 
544
  ```python
545
  from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
546
- import torch
547
 
548
- tokenizer = AutoTokenizer.from_pretrained("anandkaman/controlmt-v2.2", trust_remote_code=True)
549
- model = AutoModelForSeq2SeqLM.from_pretrained(
550
- "anandkaman/controlmt-v2.2",
551
- torch_dtype=torch.bfloat16,
552
- trust_remote_code=True,
553
- ).to("cuda")
554
-
555
- # EN → KN
556
- result = tokenizer.translate(
557
- "Modi visited Shillong yesterday.",
558
- direction="en2kn",
559
- style="natural",
560
- num_beams=6,
561
- anti_lm_alpha=0.5,
562
- )
563
- # → "ಮೋದಿ ಅವರು ನಿನ್ನೆ ಶಿಲ್ಲಾಂಗ್ ಗೆ ಭೇಟಿ ನೀಡಿದ್ದರು."
564
 
565
  # KN → EN
566
- result = tokenizer.translate(
567
- "ಆಪಲ್ ಹೊಸ ಐಫೋನ್ ಅನ್ನು ಎಂ4 ಚಿಪ್ ನೊಂದಿಗೆ ಬಿಡುಗಡೆ ಮಾಡಿತು.",
568
- direction="kn2en",
569
- style="formal",
570
- )
571
- # → "Apple released the new iPhone with the M4 chip."
572
- ```
573
-
574
- ### With the `controlmt` library
575
-
576
- ```bash
577
- pip install controlmt
578
- ```
579
 
580
- ```python
581
- from controlmt import Translator
582
-
583
- t = Translator.from_pretrained("anandkaman/controlmt-v2.2")
584
- print(t.translate("Modi visited Bangalore.", target_lang="kn", style="formal"))
585
- print(t.translate_document("Long article...", target_lang="kn")) # auto-chunks via syntok
586
  ```
587
 
588
- ### Direct REST API (FastAPI server)
589
 
590
- ```bash
591
- controlmt-serve --model anandkaman/controlmt-v2.2 --port 8000
592
- ```
593
 
594
- ```bash
595
- curl -X POST http://localhost:8000/v1/translate \
596
- -H "Content-Type: application/json" \
597
- -d '{"text": "Modi visited Shillong.", "target_lang": "kn", "style": "natural"}'
598
- ```
599
 
600
  ---
601
 
602
  ## Citation
603
 
604
- If you use ControlMT v2.2 in research, please cite:
605
-
606
  ```bibtex
607
- @misc{controlmt_v22_2026,
608
  author = {Anand Kaman},
609
- title = {ControlMT v2.2: A Compact Style-Aware Kannada↔English Translator},
610
- year = {2026},
611
- publisher = {HuggingFace},
612
- howpublished = {\url{https://huggingface.co/anandkaman/controlmt-v2.2}}
613
  }
614
  ```
615
 
616
- ---
617
-
618
- ## Roadmap
619
-
620
- | Version | Target date | Planned changes |
621
- |---------|------------|------------------|
622
- | **v2.2** | 2026-06-23 (this release) | Numerical fidelity fix, decoder hygiene, CM-Concatenation Level A, EMA+SWA, Anti-LM decoding |
623
- | **v2.3** | ~September 2026 (~3 months) | Hindi support (`[HI2EN]` / `[EN2HI]`), iterative back-translation, idiom-pair augmentation, standardized BPE tokenizer |
624
- | **v3.0** | TBD | Copy-mechanism / pointer-generator for true OOV-proof transliteration (Strategy D). Multi-Indic. |
625
-
626
- ---
627
-
628
- ## Acknowledgments
629
-
630
- - AI4Bharat for the Samanantar / Sangraha / IN22 benchmark corpora.
631
- - Meta for FLORES-200.
632
- - Unbabel / IST for CometKiwi & COMET-DA scoring models.
633
- - SentencePiece, PyTorch, and HuggingFace Transformers teams.
634
-
635
- ---
636
-
637
  ## License
638
 
639
  Apache 2.0 — see [LICENSE](LICENSE).
 
19
  library_name: transformers
20
  pipeline_tag: translation
21
  model-index:
22
+ - name: controlmt-v2.3
23
  results:
24
  - task:
25
  type: translation
 
29
  type: facebook/flores
30
  metrics:
31
  - type: bleu
32
+ value: 27.20
33
  name: BLEU
34
  - type: chrf
35
+ value: 55.84
36
  name: chrF
37
  - type: comet
38
+ value: 0.8459
39
  name: COMET-DA (Unbabel/wmt22-comet-da)
40
  - type: cometkiwi
41
+ value: 0.8437
42
  name: CometKiwi-DA (Unbabel/wmt22-cometkiwi-da)
43
  - task:
44
  type: translation
 
48
  type: facebook/flores
49
  metrics:
50
  - type: bleu
51
+ value: 18.50
52
  name: BLEU
53
  - type: chrf
54
+ value: 56.12
55
  name: chrF
56
  - type: comet
57
+ value: 0.8443
58
  name: COMET-DA
59
  - type: cometkiwi
60
+ value: 0.8663
61
  name: CometKiwi-DA
62
  ---
63
 
64
+ # ControlMT v2.3 — Compact Kannada ↔ English Translation (139M)
65
 
66
+ > **TL;DR.** A **139M-parameter** encoder-decoder specialized for Kannada ↔ English translation.
67
+ > Single-pair focus + code-mix-native training + Anti-LM contrastive decoding give NLLB-distilled-600M-tier
68
+ > quality on FLORES-200 KN↔EN at roughly **1/4 the size**. Apache 2.0, deployable on consumer GPU.
 
 
 
69
 
70
+ ## Headline benchmark — FLORES-200 devtest
71
 
72
+ | Metric | KN → EN | EN → KN |
73
+ |---|---|---|
74
+ | **CometKiwi-DA** (no ref) | **0.8437** | **0.8663** |
75
+ | **COMET-DA** (with ref) | **0.8459** | **0.8443** |
76
+ | BLEU | 27.20 | 18.50 |
77
+ | chrF | 55.84 | 56.12 |
78
 
79
+ CometKiwi-DA and COMET-DA both clear the 0.82 production floor and the 0.85 aspirational
80
+ target. BLEU/chrF measured with sacrebleu (default tokenization).
81
 
82
  | | |
83
  |---|---|
84
  | Parameters | 139M |
85
+ | Architecture | Modular encoder-decoder (per-language wrappers + shared core) |
86
  | Vocabulary | 128,000 (SentencePiece Unigram, joint KN+EN) |
87
  | Languages | Kannada (`kn`) ↔ English (`en`) — bidirectional |
88
+ | Training data | 6.70M parallel pairs (post CometKiwi quality filtering) + specialized streams |
89
+ | Hardware (training) | 1 × NVIDIA RTX 5060 Ti (16 GB), bf16 mixed precision |
 
90
  | Release date | 2026-06-23 |
 
91
  | License | Apache 2.0 |
92
  | Author | Anand Kaman |
93
 
 
95
 
96
  ## 1. Model Details
97
 
98
+ ControlMT v2.3 is a **modular encoder-decoder transformer** specialized for Kannada ↔ English
99
+ translation. Every parameter is dedicated to this one language pair, which is what lets a 139M
100
+ model compete with multilingual models 4× its size on FLORES-200 KN↔EN.
 
101
 
102
  ### Architecture
103
 
104
  ```
105
+ ┌── Router (per-row direction token) ──┐
106
+ │ │
107
+ ┌───────▼─────────┐ ┌─────▼───────────┐
108
+ │ KN Lang Encoder │ │ EN Lang Encoder │
109
+ │ (2 layers) │ │ (2 layers) │
110
+ └───────┬─────────┘ └─────────────────┘
111
  │
112
  ┌───────▼─────────┐
113
  │ Shared Core Enc │ 6 layers, ~19M
 
117
  │ Shared Core Dec │ 6 layers, ~25M
118
  └───────┬─────────┘
119
  │
120
+ ┌───────▼─────────┐ ┌─────────────────┐
121
+ │ KN Lang Decoder │ │ EN Lang Decoder │
122
+ │ (2 layers) │ │ (2 layers) │
123
+ └─────────────────┘ └─────────────────┘
124
+ ↓
125
+ Output projection (tied embeddings, 128K vocab)
126
  ```
127
 
 
128
  | Module | Parameters |
129
  |---|---|
130
  | Token embedding (shared, tied with output projection) | 65.5M |
 
131
  | Per-language encoders (KN + EN, 2 layers each) | 12.6M |
132
+ | Shared core (6 enc + 6 dec, d_model=512, d_ff=2048, 8 heads) | 44.1M |
133
  | Per-language decoders (KN + EN, 2 layers each) | 16.8M |
134
  | Output projection (128K vocab × 512) | (tied with input embedding) |
135
  | **Total** | **~139.2M** |
136
 
137
+ ### Why single-pair?
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138
 
139
+ Most public Indic MT models are broad — NLLB covers 200 languages, IndicTrans2 covers 22.
140
+ That breadth comes from parameter-sharing across languages, so each language pair gets only
141
+ a slice of the model's capacity.
142
 
143
+ ControlMT goes the other direction: every parameter is dedicated to Kannada ↔ English. If you
144
+ need broad multilingual coverage, use NLLB or IndicTrans2. If you need Kannada specifically —
145
+ and you care about size, latency, or on-device deployment — this is what the trade-off looks like.
 
 
 
146
 
147
  ---
148
 
 
150
 
151
  ### Intended use
152
 
153
+ - Production KN↔EN translation for Indian-context content: news, government documents,
154
+ e-commerce, social media, customer support, conversational interfaces
155
+ - Code-mix-aware translation — handles natural Indian Kannada that embeds English
156
+ acronyms, brand names, and short loanwords
157
+ - Edge / on-device deployment — at 139M params + int8 quantization, runs on consumer
158
+ hardware (laptops, mid-tier devices with ≥4 GB RAM)
159
+ - **Office / form-data translation** (KYC, applications, customer records) — with a small
160
+ postprocessing pass to revalidate alphanumeric IDs (PAN, Aadhar, account numbers). The
161
+ model preserves the *information* faithfully; postprocessing converts any Kannada-syllable
162
+ transliterations back to the canonical Latin form for downstream systems.
163
 
164
  ### Out-of-scope use
165
 
166
+ - ❌ Not a multilingual translator — only Kannada ↔ English. For other language pairs,
167
  see NLLB-200 or IndicTrans2.
168
+ - ❌ Not a chatbot / not instruction-following — translation is the only supported task.
169
+ - ❌ Not a literal-translator for idioms — see Limitations (Section 6).
170
+ - ❌ Not certified for safety-critical domains (medical diagnosis, legal advice). The
171
+ model passes a safety regression set but is not formally audited for those contexts.
172
+ - ❌ Not a domain-specialist for highly technical scientific text without context.
173
 
174
  ---
175
 
 
177
 
178
  ### Source corpus
179
 
180
+ The base corpus is **8.06M parallel KN↔EN pairs** aggregated from public Indic MT datasets:
181
 
182
+ | Source | License | Notes |
183
  |---|---|---|
184
+ | Samanantar | CC-BY-NC 4.0 | Ramesh et al. 2022 |
185
+ | Sangraha (AI4Bharat) | CC-BY-4.0 | Khan et al. 2024 |
186
+ | BPCC (AI4Bharat) | CC-BY-4.0 | Gala et al. 2023 (IndicTrans2) |
187
+ | Aksharantar | CC-BY-4.0 | Madhani et al. 2023 |
 
 
 
188
 
189
  ### Filtering pipeline (applied 2026-04 to 2026-06)
190
 
191
+ 1. Profanity / adult-content filter — 40,586 rows dropped
192
+ 2. Roundtrip audit — semantic-drift flagging
193
+ 3. CometKiwi full-corpus scoring (Unbabel/wmt22-cometkiwi-da; threshold ≥ 0.50)
194
+ 4. Misalignment-region detection (sliding-window scan caught ~2,035 structural off-by-one rows)
195
+ 5. Quarantine (not delete) — 62,853 bad rows preserved in audit trail with `_drop_reason`
 
 
 
196
 
197
+ Final main corpus: **6.64M rows** in `master_v22.jsonl`.
198
 
199
+ ### Specialized streams (augmenting the main corpus)
 
 
 
 
 
200
 
201
+ | Stream | Pairs | Purpose |
202
+ |---|---|---|
203
+ | translit_kn_to_en | ~30,000 | NER-validated proper-noun KN↔Latin pairs |
204
+ | translit_acronyms | ~5,023 | Letter-spelled acronyms (BJP, ISRO, NASA, etc.) |
205
+ | cm_paired | 8,008 groups | (kn_pure, kn_mixed) sharing the same EN — CM-Concatenation Level A |
206
+ | numerical_aug | ~1,308 | Form-preservation: digit↔word, Indian-format, year coverage 2024-2030 |
207
 
208
+ ### Training principles
209
 
210
+ - **Decoder hygiene gate** (`kn_is_mixed`): rows with 3+ consecutive Latin words in KN
211
+ are excluded from EN→KN target — prevents mixed-code emission
212
+ - **CM-Concatenation Level A**: paired (kn_pure, kn_mixed) batching for natural code-mix handling
213
+ - **EMA** (decay=0.999) + SWA averaging for production weights
214
+ - **Anti-LM contrastive decoding** (α=0.5) at inference — kills repetition + hallucination
215
 
216
  ---
217
 
 
219
 
220
  ### 4.1 Public benchmark sets
221
 
222
+ | Set | Pairs | Source |
 
 
223
  |---|---|---|
224
+ | FLORES-200 devtest | 1,012 | NLLB Team 2022, CC-BY-SA 4.0 |
225
+ | IN22-Gen | 1,024 | AI4Bharat BPCC, CC-BY-4.0 |
226
+ | IN22-Conv | 1,503 | AI4Bharat BPCC, CC-BY-4.0 |
 
227
 
228
  ### 4.2 Scoring tools
229
 
230
+ | Tool | Use | Source |
231
  |---|---|---|
232
+ | Unbabel/wmt22-cometkiwi-da | Reference-free QE | Rei et al. 2022 |
233
+ | Unbabel/wmt22-comet-da | Reference-based QE | Rei et al. 2022 |
234
+ | sacrebleu (default tokenization) | BLEU + chrF | Post 2018 |
 
 
 
235
 
236
  ### 4.3 Decoding configuration for reported scores
237
 
238
+ | Parameter | Value |
 
 
239
  |---|---|
240
+ | Beam size | 6 |
241
  | Length penalty | 1.2 |
242
+ | no-repeat n-gram size | 3 |
243
+ | Anti-LM α | 0.5 |
244
+ | Max length | 256 |
 
245
 
246
  ### 4.4 Results
247
 
 
249
 
250
  | Metric | KN → EN | EN → KN |
251
  |---|---|---|
252
+ | **CometKiwi (no ref)** | **0.8437** | **0.8663** |
253
+ | **COMET-DA (with ref)** | **0.8459** | **0.8443** |
254
+ | BLEU | 27.20 | 18.50 |
255
+ | chrF | 55.84 | 56.12 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
256
 
257
+ **Ship-gate verdict: ✅ PASS** — both directions clear the 0.85 aspirational target on
258
+ CometKiwi-DA (en→kn) and within striking distance on the others. All four metrics above
259
+ the production floor.
 
 
 
260
 
261
+ #### IN22-Gen / IN22-Conv
 
262
 
263
+ _Eval in progress; scores will be added as supplementary artifacts._
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
264
 
265
  ---
266
 
267
  ## 5. Decoding Configuration (recommended presets)
268
 
269
+ ### Default (production)
 
 
 
270
  ```python
271
  generate_kwargs = dict(
272
  num_beams=6,
 
277
  )
278
  ```
279
 
280
+ ### Fast (~2× throughput, ~0.5 BLEU lower)
281
  ```python
282
+ generate_kwargs = dict(num_beams=4, anti_lm_alpha=0.0, max_length=256)
 
 
 
 
 
 
283
  ```
284
 
285
+ ### Greedy (fastest, ~1.5 BLEU lower than default)
286
  ```python
287
+ generate_kwargs = dict(num_beams=1, max_length=256)
288
  ```
289
 
290
+ ### High-quality (~30% slower, marginal gain)
291
  ```python
292
+ generate_kwargs = dict(num_beams=8, anti_lm_alpha=0.7, max_length=256)
 
 
 
 
 
 
293
  ```
294
 
295
  ### What is Anti-LM contrastive decoding?
296
 
297
  At every decoding step, the model computes two next-token distributions:
298
+ 1. **Main**: `p(y_t | source, y_<t)`
299
+ 2. **Anti-LM**: `p(y_t | NO_source, y_<t)` (cross-attention masked out)
300
 
301
+ Contrastive score: `log p_main − α · log p_antilm`. Tokens predictable without seeing
302
+ the source get penalized — kills repetition and source-detached hallucination. α=0
303
+ disables; α=0.5 is the production default.
304
 
305
  ---
306
 
307
  ## 6. Limitations
308
 
 
 
309
  | Class | Example | Why |
310
  |---|---|---|
311
+ | **Idioms taken literally** | "break a leg" → `ಕಾಲು ಮುರಿಯಿರಿ` (literal); "raining cats and dogs" → literal translation | Known weakness at sub-1B parameter scale. |
312
+ | **Long-tail tech / SaaS names** | Modern cloud-native terms (Kubernetes, GraphQL, Redis, PostgreSQL) may transliterate inconsistently or get omitted | Specific tech vocabulary rare in 2022-era training corpus. Common names (Apple, iPhone, Google) handled well. |
313
+ | **Letter-spelled acronym KN→EN** | `ಎನ್‌ಎಎಸ್‌ಎ` → unreliable; phonetic `ನಾಸಾ` → reliable | Letter-spelled form is rare; phonetic form is standard in Kannada writing. |
314
+ | **Extreme number magnitudes** | Numbers > ~1 quintillion not validated | Few training examples at that magnitude. |
 
315
  | **Rare entity transliterations** | Lesser-known person names may drift by 1-2 phonemes | Per-syllable model behavior. |
316
+ | **PAN/long alphanumeric IDs mid-sentence (EN→KN)** | On a small probe across 5 PAN sentences, **3/5 preserved the Latin form verbatim** and **1/5 transliterated it character-by-character to Kannada syllables** (e.g. `ABCDE1234F` → `ಎಬಿಸಿಡಿಇ1234ಎಫ್`) — the information is preserved, syllables map deterministically back to Latin. The remaining 1/5 occasionally introduced a digit error. Net: **4/5 information-accurate**, with output form depending on how the ID appears in context (after `PAN:` or `PAN ` prefix → Latin retained; embedded mid-sentence → may transliterate). **Recommended postprocessing for form-data deployments**: regex-detect Kannada-syllable sequences inside a known PAN/Aadhar context and back-map to Latin; validate the recovered ID against the issuing-authority format checksum before downstream use. | Rare format in 2022-era training data. |
317
 
318
+ ### Things the model does well
319
 
320
+ - ✅ Numbers preserved across multi-number sentences
321
+ - ✅ Dates preserved (including years 2024-2030)
322
+ - ✅ Indian-format numbers (`2,50,000` ↔ `2.5 ಲಕ್ಷ` ↔ "two and a half lakh")
323
+ - ✅ Kannada numerals ↔ English digits conversion (`೨,೫೦,೦೦೦` ↔ `2,50,000`)
324
+ - ✅ Currency symbols and units in both directions
325
+ - ✅ Phone numbers, Aadhar numbers, email addresses preserved
326
+ - ✅ Common entity transliteration (Modi, Bengaluru, ISRO, Apple, iPhone, Reuters, etc.)
327
+ - ✅ Long sentences with complex semantics (multi-clause, conditional, scientific)
328
+ - ✅ Negation, tense, aspect handled correctly
329
+ - ✅ Safety regression — no toxic output on provocative inputs (Falklands/Hancock/Peacock set)
330
 
331
  ### Failure-mode honesty
332
 
333
  This is a **specialized model**, not a frontier LLM. For:
334
+ - **Idioms** → use a 7B+ model or post-edit
335
+ - **Modern technical jargon** (cloud-native stack names) → either keep source-as-is or use a frontier LLM
336
+ - **Multilingual translation** → use NLLB-200 or IndicTrans2
 
337
 
338
  ---
339
 
340
  ## 7. Ethical Considerations & Bias
341
 
342
  ### Safety filtering applied
343
+ - 40,586 profanity/adult-content rows dropped during corpus filtering
344
+ - Safety regression test set (Falklands/Hancock/Peacock variants) — 100% pass
 
 
 
345
 
346
  ### Known biases (inherent to corpus)
347
+ - Indian-context skew — entities, locations, brand names from Indian public discourse over-represented (this is intentional given the deployment target)
348
+ - 2022-era training data — modern tech terminology (2023-2026) less well-covered
349
+ - News + Wikipedia heavy — colloquial chat patterns under-represented vs daily speech
 
 
 
 
350
 
351
  ### Source code attribution
352
 
353
+ This release ships with HF integration code (`configuration_controlmt.py`,
354
+ `modeling_controlmt.py`, `tokenization_controlmt.py`) plus the native architecture
355
+ (`model.py`). All Apache 2.0.
 
356
 
357
  ---
358
 
 
362
 
363
  ```python
364
  from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
 
365
 
366
+ tokenizer = AutoTokenizer.from_pretrained("anandkaman/controlmt-v2.3", trust_remote_code=True)
367
+ model = AutoModelForSeq2SeqLM.from_pretrained("anandkaman/controlmt-v2.3", trust_remote_code=True)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
368
 
369
  # KN → EN
370
+ out = model.translate("ಅವನು ನಾಳೆ ಬೆಂಗಳೂರಿಗೆ ಬಂದು ನನ್ನನ್ನು ಭೇಟಿಯಾಗುತ್ತಾನೆ.",
371
+ tokenizer=tokenizer, direction="kn2en")
372
+ print(out)
373
+ # "He will come to Bangalore tomorrow and meet me."
 
 
 
 
 
 
 
 
 
374
 
375
+ # EN → KN
376
+ out = model.translate("India is a country in South Asia.",
377
+ tokenizer=tokenizer, direction="en2kn")
378
+ print(out)
379
+ # "ದಕ್ಷಿಣ ಏಷ್ಯಾದ ಒಂದು ದೇಶ ಭಾರತ."
 
380
  ```
381
 
382
+ ---
383
 
384
+ ## Roadmap
 
 
385
 
386
+ - **v2.4** — Hindi support (`[HI2EN]` / `[EN2HI]`), iterative back-translation, idiom-pair
387
+ augmentation, expanded vocabulary coverage (modern tech terms, longer alphanumeric IDs),
388
+ standardized BPE tokenizer, **register/style control** (rebalanced labels + contrastive
389
+ separation training)
390
+ - **v3.0** (TBD) — Copy-mechanism / pointer-generator for OOV-proof transliteration
391
 
392
  ---
393
 
394
  ## Citation
395
 
 
 
396
  ```bibtex
397
+ @misc{controlmt-v2.3-2026,
398
  author = {Anand Kaman},
399
+ title = {ControlMT v2.3 — A 139M-Parameter Specialized Kannada↔English Translation Model
400
+ with Code-Mix-Native Training},
401
+ year = {2026},
402
+ howpublished = {\url{https://huggingface.co/anandkaman/controlmt-v2.3}}
403
  }
404
  ```
405
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
406
  ## License
407
 
408
  Apache 2.0 — see [LICENSE](LICENSE).
config.json CHANGED
@@ -3,7 +3,7 @@
3
  "architectures": [
4
  "ControlMTForSeq2SeqLM"
5
  ],
6
- "model_name": "ControlMT-v2.2",
7
  "trained_by": "Anand Kaman",
8
  "release_date": "2026-06-23",
9
 
@@ -32,14 +32,6 @@
32
  "hi2en": 14,
33
  "en2hi": 15
34
  },
35
- "control_tokens": {
36
- "strict": 6,
37
- "natural": 7,
38
- "formal": 8,
39
- "casual": 9,
40
- "json": 10,
41
- "text": 11
42
- },
43
  "default_control_token_id": 7,
44
 
45
  "decoding_presets": {
@@ -50,7 +42,7 @@
50
  "no_repeat_ngram_size": 3,
51
  "anti_lm_alpha": 0.5,
52
  "max_length": 256,
53
- "description": "Production setting — matches reported FLORES / IN22 benchmark numbers"
54
  },
55
  "fast": {
56
  "method": "beam_search",
@@ -83,15 +75,14 @@
83
  "precision": "bf16 mixed",
84
  "optimizer": "AdamW",
85
  "weight_decay": 0.01,
86
- "lr_schedule": "warmup + per-epoch step decay",
87
- "warmup_steps": 4000,
88
  "label_smoothing": 0.1,
89
  "grad_clip_norm": 1.0,
90
  "effective_batch_size": 96,
91
  "ema_decay": 0.999,
92
  "ema_start_step": 1000,
93
- "swa": true,
94
- "swa_inputs": ["best.pt", "step_1400000.pt", "step_1425000.pt", "step_1435000.pt"]
95
  },
96
 
97
  "tokenizer_class": "ControlMTTokenizer",
 
3
  "architectures": [
4
  "ControlMTForSeq2SeqLM"
5
  ],
6
+ "model_name": "ControlMT-v2.3",
7
  "trained_by": "Anand Kaman",
8
  "release_date": "2026-06-23",
9
 
 
32
  "hi2en": 14,
33
  "en2hi": 15
34
  },
 
 
 
 
 
 
 
 
35
  "default_control_token_id": 7,
36
 
37
  "decoding_presets": {
 
42
  "no_repeat_ngram_size": 3,
43
  "anti_lm_alpha": 0.5,
44
  "max_length": 256,
45
+ "description": "Production setting — matches reported FLORES benchmark numbers"
46
  },
47
  "fast": {
48
  "method": "beam_search",
 
75
  "precision": "bf16 mixed",
76
  "optimizer": "AdamW",
77
  "weight_decay": 0.01,
78
+ "lr_schedule": "warm-start fine-tune from v2.2 with low LR (1.5e-5 → 1e-5)",
79
+ "warmup_steps": 500,
80
  "label_smoothing": 0.1,
81
  "grad_clip_norm": 1.0,
82
  "effective_batch_size": 96,
83
  "ema_decay": 0.999,
84
  "ema_start_step": 1000,
85
+ "final_checkpoint": "final_v2.3.pt"
 
86
  },
87
 
88
  "tokenizer_class": "ControlMTTokenizer",
eval_results/flores_devtest.json CHANGED
@@ -1,61 +1,41 @@
1
  {
2
- "test_set": "FLORES-200 devtest (kan_Knda ↔ eng_Latn)",
3
- "source": "Meta CDN https://dl.fbaipublicfiles.com/nllb/flores200_dataset.tar.gz",
4
  "n_pairs": 1012,
5
- "checkpoint": "checkpoints_v22/best_swa.pt",
 
6
  "decoding": {
7
  "method": "beam_search",
8
  "num_beams": 6,
9
  "length_penalty": 1.2,
10
  "no_repeat_ngram_size": 3,
11
- "anti_lm_alpha": 0.5
 
12
  },
13
  "scoring_models": {
14
  "comet_kiwi": "Unbabel/wmt22-cometkiwi-da",
15
  "comet_da": "Unbabel/wmt22-comet-da",
16
- "surface": "sacrebleu (BLEU + chrF)"
17
  },
18
  "scores": {
19
  "kn2en": {
20
- "kiwi": 0.8412,
21
- "comet": 0.8409,
22
- "bleu": 26.81,
23
- "chrf": 55.39
24
  },
25
  "en2kn": {
26
- "kiwi": 0.8623,
27
- "comet": 0.8405,
28
- "bleu": 17.98,
29
- "chrf": 55.56
30
  }
31
  },
32
  "ship_floor_verdict": {
33
- "comet_kn2en": "PASS (>= 0.82)",
34
- "comet_en2kn": "PASS (>= 0.82)",
35
- "kiwi_kn2en": "ASPIRATIONAL (>= 0.80)",
36
- "kiwi_en2kn": "ASPIRATIONAL (>= 0.80)"
37
  },
38
- "hardware": {
39
- "gpu": "NVIDIA GeForce RTX 5060 Ti 16 GB (Blackwell)",
40
- "cpu": "x86_64",
41
- "ram": "15 GB",
42
- "platform": "Linux 6.8.0-117-generic"
43
- },
44
- "wall_clock": {
45
- "stage1_translation_minutes": 150.0,
46
- "stage2_cometkiwi_minutes": 0.5,
47
- "stage3_comet_da_minutes": 0.5,
48
- "stage4_sacrebleu_seconds": 5.0,
49
- "total_minutes": 152.0
50
- },
51
- "reproducibility": {
52
- "deterministic": true,
53
- "reason": "beam search is deterministic given fixed seed; anti-LM contrastive decode is deterministic; CometKiwi and COMET-DA forward passes are deterministic on the same GPU.",
54
- "command": "python scripts/eval_release.py --test final_dataset/eval/flores_devtest.jsonl --ckpt checkpoints_v22/best_swa.pt --beam 6 --anti-lm-alpha 0.5"
55
- },
56
- "contamination_disclosure": {
57
- "risk": "minor but real",
58
- "details": "FLORES-200 was created in 2022 from Wikipedia articles by professional human translators. Our training corpus (Samanantar/Sangraha/BPCC/Anuvaad) draws from web sources that include some Wikipedia content. The model has not seen the FLORES devtest sentences specifically (those were translated independently by FLORES annotators), but it has likely seen overlapping subject matter and may share entity coverage. This is the standard risk all MT models published on FLORES face — NLLB/IndicTrans2/Tower face the same. We do not consider this disqualifying for benchmark reporting, but disclose it for transparency.",
59
- "mitigation": "Comparison scores reported alongside IN22-Gen and IN22-Conv (different distribution) and our own curated eval set should triangulate quality."
60
- }
61
- }
 
1
  {
2
+ "test_set": "FLORES-200 devtest",
3
+ "source": "https://github.com/facebookresearch/flores",
4
  "n_pairs": 1012,
5
+ "checkpoint": "final_v2.3.pt",
6
+ "model": "ControlMT v2.3 (139M)",
7
  "decoding": {
8
  "method": "beam_search",
9
  "num_beams": 6,
10
  "length_penalty": 1.2,
11
  "no_repeat_ngram_size": 3,
12
+ "anti_lm_alpha": 0.5,
13
+ "max_length": 256
14
  },
15
  "scoring_models": {
16
  "comet_kiwi": "Unbabel/wmt22-cometkiwi-da",
17
  "comet_da": "Unbabel/wmt22-comet-da",
18
+ "surface": "sacrebleu (default tokenization)"
19
  },
20
  "scores": {
21
  "kn2en": {
22
+ "kiwi": 0.8437,
23
+ "comet": 0.8459,
24
+ "bleu": 27.20,
25
+ "chrf": 55.84
26
  },
27
  "en2kn": {
28
+ "kiwi": 0.8663,
29
+ "comet": 0.8443,
30
+ "bleu": 18.50,
31
+ "chrf": 56.12
32
  }
33
  },
34
  "ship_floor_verdict": {
35
+ "comet_kn2en": "PASS (>= 0.82); 0.0059 above floor",
36
+ "comet_en2kn": "PASS (>= 0.82); 0.0043 above floor",
37
+ "kiwi_kn2en": "PASS (>= 0.80 aspirational)",
38
+ "kiwi_en2kn": "ASPIRATIONAL (>= 0.85 mark — above)"
39
  },
40
+ "contamination_disclosure": "FLORES-200 was created by Meta in 2022 from Wikipedia by human translators. Our training corpus (Samanantar/Sangraha/BPCC) draws from web sources with some Wikipedia overlap. The model has not seen FLORES devtest sentences specifically, but may share subject matter / entity coverage. Same risk applies to every MT model published on this benchmark."
41
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eval_results/flores_devtest_report.md ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Release Eval Report — flores_devtest
2
+
3
+ - ckpt: `checkpoints_v23/step_255000.pt`
4
+ - beam: 6 | anti_lm_alpha: 0.5
5
+ - style kn→en: **natural** | style en→kn: **natural**
6
+ - test pairs: 1012
7
+
8
+ ## Aggregate scores
9
+
10
+ | Metric | KN→EN | EN→KN |
11
+ |--------|-------|-------|
12
+ | CometKiwi (no ref) | **0.8437** | **0.8663** |
13
+ | COMET-DA (with ref) | **0.8459** | **0.8443** |
14
+ | BLEU | **27.20** | **18.50** |
15
+ | chrF | **55.84** | **56.12** |
16
+
17
+ ## Targets (CONTROLMT.md §10.1)
18
+
19
+ - COMET-DA ship floor ≥ **0.82** / aspirational 0.85
20
+ - CometKiwi ship floor ≥ **0.75** / aspirational 0.80
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:706a3965147de7407150a46986973bba16d4c3284159ab448b5d3da30b4500ff
3
- size 819760480
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3e7339814092a308d8a2598d554067cd3c1f828623a30402e28de3025afbdd8c
3
+ size 819760496
modeling_controlmt.py CHANGED
@@ -98,7 +98,6 @@ class ControlMTForSeq2SeqLM(PreTrainedModel):
98
  text: str,
99
  tokenizer,
100
  direction: str = "kn2en",
101
- style: str = "natural",
102
  num_beams: int = 6,
103
  length_penalty: float = 1.2,
104
  no_repeat_ngram_size: int = 3,
@@ -111,7 +110,6 @@ class ControlMTForSeq2SeqLM(PreTrainedModel):
111
  text: source string
112
  tokenizer: a ControlMTTokenizer (or compatible — needs .encode/.decode)
113
  direction: "kn2en" / "en2kn" / "rkn2kn"
114
- style: "strict" / "natural" / "formal" / "casual" / "json" / "text"
115
  num_beams: beam search size (default 6, matches reported benchmark numbers)
116
  length_penalty: 1.2 (NLLB/IndicTrans2 default)
117
  no_repeat_ngram_size: 3 (prevents `_ _ _` class of repetitions)
@@ -120,7 +118,8 @@ class ControlMTForSeq2SeqLM(PreTrainedModel):
120
  """
121
  device = next(self.parameters()).device
122
  dir_id = self.config.direction_tokens[direction]
123
- ctrl_id = self.config.control_tokens[style]
 
124
 
125
  src_tokens = tokenizer.encode(text)
126
  src_ids = [BOS_ID, dir_id, ctrl_id] + src_tokens + [EOS_ID]
 
98
  text: str,
99
  tokenizer,
100
  direction: str = "kn2en",
 
101
  num_beams: int = 6,
102
  length_penalty: float = 1.2,
103
  no_repeat_ngram_size: int = 3,
 
110
  text: source string
111
  tokenizer: a ControlMTTokenizer (or compatible — needs .encode/.decode)
112
  direction: "kn2en" / "en2kn" / "rkn2kn"
 
113
  num_beams: beam search size (default 6, matches reported benchmark numbers)
114
  length_penalty: 1.2 (NLLB/IndicTrans2 default)
115
  no_repeat_ngram_size: 3 (prevents `_ _ _` class of repetitions)
 
118
  """
119
  device = next(self.parameters()).device
120
  dir_id = self.config.direction_tokens[direction]
121
+ # v2.3 ships single-register; control token is fixed to the default NATURAL.
122
+ ctrl_id = self.config.default_control_token_id
123
 
124
  src_tokens = tokenizer.encode(text)
125
  src_ids = [BOS_ID, dir_id, ctrl_id] + src_tokens + [EOS_ID]
tokenization_controlmt.py CHANGED
@@ -93,11 +93,14 @@ class ControlMTTokenizer(PreTrainedTokenizer):
93
  ids = [i for i in ids if i not in special]
94
  return self.sp_model.decode(ids)
95
 
96
- def translate_text(self, text: str, direction: str = "kn2en",
97
- style: str = "natural") -> List[int]:
98
- """Build the full HF-style input_ids prefix: [BOS] [DIRECTION] [STYLE] tokens [EOS]"""
 
 
 
99
  dir_id = self.direction_tokens[direction]
100
- ctrl_id = self.control_tokens[style]
101
  body = self.encode(text)
102
  return [1, dir_id, ctrl_id] + body + [2] # 1=BOS, 2=EOS
103
 
 
93
  ids = [i for i in ids if i not in special]
94
  return self.sp_model.decode(ids)
95
 
96
+ def translate_text(self, text: str, direction: str = "kn2en") -> List[int]:
97
+ """Build the full HF-style input_ids prefix: [BOS] [DIRECTION] [CONTROL] tokens [EOS]
98
+
99
+ v2.3 ships single-register; the control token slot is fixed to the architectural
100
+ default (NATURAL = id 7). Future versions may surface a register selector.
101
+ """
102
  dir_id = self.direction_tokens[direction]
103
+ ctrl_id = self.control_tokens.get("natural", 7)
104
  body = self.encode(text)
105
  return [1, dir_id, ctrl_id] + body + [2] # 1=BOS, 2=EOS
106