File size: 15,613 Bytes
52975b9
 
df6b9d1
 
 
52975b9
df6b9d1
 
 
d46b80d
52975b9
df6b9d1
52975b9
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
 
 
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
df6b9d1
52975b9
 
 
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
 
df6b9d1
 
52975b9
 
 
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
df6b9d1
 
 
52975b9
 
 
df6b9d1
52975b9
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
 
 
df6b9d1
52975b9
df6b9d1
 
52975b9
 
df6b9d1
 
52975b9
df6b9d1
52975b9
 
df6b9d1
52975b9
df6b9d1
 
 
 
 
 
 
52975b9
 
df6b9d1
52975b9
df6b9d1
52975b9
 
 
df6b9d1
 
 
 
 
 
 
 
 
52975b9
 
df6b9d1
 
52975b9
 
 
df6b9d1
52975b9
df6b9d1
 
52975b9
 
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
52975b9
df6b9d1
 
 
 
52975b9
 
 
 
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
 
 
df6b9d1
 
 
 
 
 
52975b9
df6b9d1
 
d46b80d
 
 
 
 
df6b9d1
 
 
 
52975b9
fde63ef
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
 
df6b9d1
52975b9
df6b9d1
 
 
52975b9
df6b9d1
52975b9
 
df6b9d1
 
 
 
 
 
 
52975b9
 
df6b9d1
 
52975b9
 
 
df6b9d1
52975b9
df6b9d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52975b9
 
 
df6b9d1
 
d46b80d
df6b9d1
 
 
52975b9
df6b9d1
 
d46b80d
 
52975b9
 
 
df6b9d1
52975b9
df6b9d1
 
d46b80d
df6b9d1
 
 
 
52975b9
 
 
 
 
df6b9d1
 
 
52975b9
df6b9d1
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
# ControlMT v2.3 — Deployment Guide

ControlMT v2.3 is a small (139M params, ~280 MB bf16 safetensors) **encoder-decoder**
seq2seq translator. This guide gives you verified recipes for every realistic deployment
target, pinned to the exact versions we tested with.

> **Architecture note** — ControlMT is a **bidirectional encoder + causal decoder + cross-attention**
> seq2seq model, in the family of T5/mBART. It is **not** a decoder-only autoregressive LM.
> That distinction matters: it determines what inference frameworks can run it. See
> [Section 9 — Not directly supported](#9-not-directly-supported) below before reaching for vLLM / Ollama / GGUF.

---

## 1. Verified deployment matrix

All numbers below are measured on the same machine (Linux, RTX 5060 Ti, Python 3.12.3,
torch 2.10.0+cu128, transformers 4.57.6), `--num_beams=2`, 6 KN↔EN test pairs. Median
of per-pair latencies. **Load time excludes first-time HF download (~30s for 280 MB).**

| Recipe                     | Hardware            | Latency / pair | Memory      | Verified |
|----------------------------|---------------------|----------------|-------------|----------|
| **CPU bf16** (recommended) | Any x86/ARM, ≥1 GB  | 0.51 s         | 280 MB RAM  | ✓        |
| CPU fp32                   | Any x86/ARM, ≥1 GB  | 1.44 s         | 560 MB RAM  | ✓        |
| **CPU int8-dynamic** (fastest CPU) | Any x86/ARM ≥1 GB | 0.28 s   | ~140 MB RAM | ✓        |
| **GPU fp16** (recommended) | ≥ 1 GB VRAM (T4/3050+/M-series) | 0.19 s | 404 MB VRAM | ✓ |
| GPU bf16                   | ≥ 1 GB VRAM         | 0.19 s         | 404 MB VRAM | ✓        |
| GPU fp32                   | ≥ 1 GB VRAM         | 0.20 s         | 793 MB VRAM | ✓        |
| HuggingFace Space (Docker) | HF free-tier CPU    | 3 – 15 s       | shared      | ✓ live   |
| HF Inference Endpoints     | T4 / A10            | 0.2 – 0.5 s    | managed     | ⚠ untested but pattern works |
| FastAPI self-host          | any                 | matches device | matches dev | ✓        |
| Docker                     | any with Docker     | matches device | matches dev | ✓        |
| ONNX Runtime               | x86/ARM/GPU         | ~0.5–2 s       | varies      | ⚠ experimental (manual export) |

---

## 2. Pinned versions (the exact stack we verified)

```
python              3.12.3   (3.10–3.12 supported, 3.10/3.11 also tested informally)
torch               2.10.0   (CUDA build: 2.10.0+cu128 for GPU; CPU build: 2.10.0+cpu)
transformers        4.57.6   (4.40+ should work; uses trust_remote_code)
sentencepiece       0.2.1
safetensors         0.7.0
huggingface_hub     0.36.2   (>= 0.27 fine; 0.34+ recommended)
# Optional:
fastapi             0.129.0  (only if wrapping in an HTTP API)
uvicorn             0.40.0
pydantic            2.12.5
```

Minimum `requirements.txt` for the model alone:

```text
torch>=2.0,<3
transformers>=4.40,<5
sentencepiece>=0.1.99
safetensors>=0.4
huggingface_hub>=0.27
```

---

## 3. Quick start — Python (CPU or GPU)

```python
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

MODEL_ID = "anandkaman/controlmt-v2.3"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    dtype=torch.bfloat16,                 # bf16 on CPU is 2.8× faster than fp32 (verified)
)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device).eval()

# Translate
out = model.translate(
    "ನಾನು ಕನ್ನಡ ಮಾತನಾಡುತ್ತೇನೆ.",
    tokenizer=tokenizer,
    direction="kn2en",       # or "en2kn"
    num_beams=2,             # 1=fastest, 4 is a reasonable CPU cap, 6 ≈ no gain
    anti_lm_alpha=0.5,       # contrastive decoding; 0 = vanilla beam
    max_length=200,
)
print(out)   # → "I speak Kannada."
```

**On GPU**, use `dtype=torch.float16` for slightly broader hardware compatibility
(Volta/Pascal generation GPUs don't have native bf16). bf16 is preferred on Ampere+ (3000-series and newer).

---

## 4. Maximum-throughput CPU recipe — int8 dynamic quantization

Best for ≤ 4 GB RAM devices, Raspberry Pi 5, or CPU-only containerized deployments.
**~1.8× faster than CPU bf16, ~50% memory reduction, identical output on our 6-pair test.**

```python
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

MODEL_ID = "anandkaman/controlmt-v2.3"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(MODEL_ID, trust_remote_code=True)
model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)
model.eval()

out = model.translate("I speak Kannada.", tokenizer=tokenizer, direction="en2kn",
                      num_beams=2, max_length=200)
```

No quality regression observed in our test set. For production deployment, you should
re-validate on your own representative sentences — int8 dynamic occasionally drops 0.5–1 BLEU
on long-tail outputs.

---

## 5. HuggingFace Space (Docker, FastAPI + static HTML)

A live demo is at **[anandkaman/controlmt-demo](https://huggingface.co/spaces/anandkaman/controlmt-demo)**.
The Space is a **Docker SDK** Space running a FastAPI backend + vanilla HTML/CSS/JS
frontend. Full source ships in this release under `assets/space/` — clone-and-modify
to build your own demo.

Structure:
```
controlmt-demo/
├── Dockerfile
├── main.py                # FastAPI app: /api/translate, /api/rate, /api/health
├── pipeline.py            # Opt-in logging (PII redaction, batched dataset upload)
├── requirements.txt
├── README.md              # HF Space frontmatter (sdk: docker, app_port: 7860)
└── static/
    ├── index.html         # Hero, settings, source/translation, ratings, examples
    ├── style.css          # Mobile-first, 2-col on desktop ≥900px
    └── app.js             # Vanilla JS — fetches /api/* from same origin
```

---

## 6. FastAPI / Flask REST API (self-hosted)

For a production API on your own infrastructure, the Space's `main.py` is the canonical
reference. Minimal version:

```python
# requirements: fastapi==0.129.0 uvicorn[standard]==0.40.0 pydantic==2.12.5 + the model stack
import torch
from fastapi import FastAPI
from pydantic import BaseModel, Field
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

MODEL_ID = "anandkaman/controlmt-v2.3"
app = FastAPI()

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(
    MODEL_ID, trust_remote_code=True,
    dtype=torch.bfloat16 if not torch.cuda.is_available() else torch.float16,
)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()

class Req(BaseModel):
    text: str = Field(..., min_length=1, max_length=5000)
    direction: str = "kn2en"
    num_beams: int = 2

@app.post("/translate")
def translate(req: Req):
    out = model.translate(
        req.text, tokenizer=tokenizer, direction=req.direction,
        num_beams=req.num_beams, anti_lm_alpha=0.5, max_length=200,
    )
    return {"translation": out}
```

```bash
uvicorn main:app --host 0.0.0.0 --port 8000 --workers 1
```

> Use `--workers 1` for a single-GPU box: each worker loads its own model copy. For
> CPU deployments you can scale to `--workers 2` if you have ≥1 GB RAM per worker.

---

## 7. Docker (self-hosted)

The Space's `Dockerfile` is the verified reference. For non-HF deployment
(your own server / Kubernetes / Cloud Run / Fly.io):

```dockerfile
FROM python:3.12-slim

ENV DEBIAN_FRONTEND=noninteractive \
    PYTHONDONTWRITEBYTECODE=1 \
    PYTHONUNBUFFERED=1 \
    PIP_NO_CACHE_DIR=1

RUN useradd -m -u 1000 user
USER user
ENV PATH=/home/user/.local/bin:$PATH \
    HF_HOME=/home/user/.cache/huggingface

WORKDIR /app
COPY --chown=user:user requirements.txt ./
RUN pip install --user --no-cache-dir -r requirements.txt
COPY --chown=user:user . ./

EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
```

For GPU containers, swap base image to `nvidia/cuda:12.8.0-cudnn8-runtime-ubuntu22.04`
and install Python + pip on top.

---

## 8. HuggingFace Inference Endpoints (managed cloud)

For zero-DevOps managed deployment:

1. Open [HF Inference Endpoints](https://ui.endpoints.huggingface.co/)
2. Point at `anandkaman/controlmt-v2.3`
3. In **Advanced configuration**: enable `trust_remote_code=true`
4. Recommended: **GPU Small (T4)** — ~$0.40/hr, ~0.3 s per pair after warmup

**Not yet end-to-end verified by us** — recipe follows the standard HF custom-code pattern;
if you hit an issue, file at the GitHub repo.

---

## 9. Not directly supported

These platforms are popular for LLM serving, but ControlMT's encoder-decoder architecture
is not compatible without significant adapter work. If you find yourself reaching for one
of these, the right answer is **use Python+Transformers or FastAPI instead** — they give
the same throughput on this model size.

| Platform        | Reason                                                                     | Recommended alternative |
|-----------------|----------------------------------------------------------------------------|-------------------------|
| **vLLM**        | Supports BART/T5 via known class names only; custom `trust_remote_code` models fall outside its optimized seq2seq path | FastAPI wrapper (Section 6) |
| **Ollama**      | Uses GGUF format + llama.cpp — both are decoder-only architectures, no encoder pass or cross-attention | FastAPI wrapper (Section 6) |
| **llama.cpp / GGUF** | Same architectural reason as Ollama                                  | FastAPI wrapper (Section 6)    |
| **HF TGI**      | Same `trust_remote_code` limitation as vLLM                                | HF Inference Endpoints (Section 8) or FastAPI (Section 6) |
| **bitsandbytes int8** | Our model code doesn't route int8↔fp16 dtype conversions through linear layers in the way bnb expects. Verified failure: `RuntimeError: self and mat2 must have the same dtype, but got Half and Char` | CPU int8 dynamic (Section 4) — same memory savings, no patching required |

At v2.3's size (139M params, 0.19 s/pair on a $300 GPU), the optimizations these tools provide
(KV-cache reuse, paged attention, continuous batching) are dominated by request overhead and
not worth the integration cost.

### If you absolutely need GGUF / Ollama / vLLM compatibility

The realistic path is **knowledge distillation**: train a small decoder-only LLM
(~200–600M parameters, Llama-style or Phi-style architecture) on ControlMT's outputs as targets.
The student is decoder-only, so it's GGUF-compatible, runs natively in Ollama, can be quantized
with AWQ/GPTQ, and integrates with the entire LLM serving ecosystem.

- **Quality trade-off**: decoder-only is less parameter-efficient than seq2seq for translation.
  A ~400–600M decoder-only student is the realistic minimum to match v2.3's quality on common cases.
- **Cost**: ~1 week to set up distillation + 2–3 days of student training + re-eval. Comparable to a normal model retrain.
- **Win**: a sibling repo (planned `anandkaman/controlmt-v3.0-llama`) that ships as native GGUF
  + works with Ollama directly. Same KN↔EN coverage, ~3–5× larger model file, but 10× simpler deployment for end users.
- **Prior art**: AI4Bharat does this exact pattern (IndicTrans2 → IndicTrans2-distilled).
  The architecture-shift loss is real but recoverable.

This is a **v3.0+ roadmap item**, not currently available. If you're building infrastructure
around ControlMT today and absolutely need an Ollama-deployable variant, watch the v3 changelog
or open an issue at the GitHub repo to signal demand.

---

## 10. ONNX Runtime (experimental)

ONNX export is possible but **requires manual encoder/decoder splitting** and a re-implementation
of beam search in your target runtime — `optimum-onnx` doesn't know about ControlMT's custom
architecture and won't auto-trace it.

Sketch:
```python
import torch
# 1. Export encoder: takes (input_ids, attn_mask, direction_id, control_id) → encoder_states
torch.onnx.export(model.controlmt.encoder, (x, mask, dir_id, ctrl_id),
                  "encoder.onnx", opset_version=17, ...)
# 2. Export decoder step: takes (encoder_states, decoder_input_ids, cache) → next_token_logits
torch.onnx.export(decoder_step_fn, (...), "decoder_step.onnx", ...)
# 3. Reimplement beam search in your runtime (Python+onnxruntime, JS, C++) by calling
#    encoder once and decoder_step N×beams times.
```

Contributions of a working ONNX recipe + an `assets/onnx_export.py` script welcome at
[github.com/anandkaman/ControlMT](https://github.com/anandkaman/ControlMT).

---

## 11. Batched inference (high-throughput document translation)

For document-level workloads (1000+ sentences) on GPU:

```python
# Sketch — beam search runs per-sentence, but you can dispatch in parallel by chunking.
from concurrent.futures import ThreadPoolExecutor

def translate_batch(texts, direction="kn2en", num_beams=2):
    with ThreadPoolExecutor(max_workers=4) as ex:
        return list(ex.map(
            lambda t: model.translate(t, tokenizer=tokenizer, direction=direction,
                                       num_beams=num_beams, max_length=200),
            texts,
        ))
```

Typical RTX 3060+ throughput: **~30–50 sentences/second at beam=2** in this configuration.
For genuinely batched (cross-sentence) beam search you'd need to extend `model.translate`
to handle variable EOS-completion, which our v2.3 weights support but the wrapper exposes
only one sentence at a time.

---

## 12. Windows / cross-platform notes

The Python recipe (Section 3) works on Windows, macOS, and Linux without changes:
- **Windows native** — install Python 3.10–3.12, `pip install` the pinned versions, run the script
- **macOS (Apple Silicon)** — works on CPU. For GPU, `device="mps"` with `dtype=torch.float16`
- **WSL2** — use the Linux recipe, identical performance to native Linux

For a Windows desktop app that bundles ControlMT:
- **PyInstaller / Briefcase**: package the Python script + cached HF model files
- **Docker Desktop**: run the Docker recipe (Section 7) on Windows hosts
- **ONNX**: cross-platform but see Section 10 — experimental

---

## 13. Pre-launch checklist

- [ ] Smoke-test on your hardware with `python verify_deployment.py --device <cpu|cuda>`
  (script lives at `assets/scripts/verify_deployment.py`)
- [ ] **For form-data / KYC use**: add the PAN-postprocessing regex (model card Section 6 Limitations)
- [ ] Rate-limit your endpoint — model accepts ≤ 5000 chars per request; truncate or chunk
- [ ] Health endpoint that round-trips a known pair (e.g. `"hello" → "ಹಲೋ"`)
- [ ] Memory headroom: fp32 ≥ 700 MB, bf16/fp16 ≥ 400 MB, int8 dynamic ≥ 200 MB
- [ ] Idempotent restarts — model is stateless; safe to redeploy without warmup

---

## Related docs

- **Model card**: [README.md](README.md) — FLORES/IN22 scores, limitations, intended use
- **Training methodology**: [TRAINING_GUIDE.md](TRAINING_GUIDE.md) — corpus + filtering + training principles
- **Privacy policy** (hosted demo): [PRIVACY.md](PRIVACY.md)
- **License**: [LICENSE](LICENSE) — Apache 2.0
- **Live demo**: [huggingface.co/spaces/anandkaman/controlmt-demo](https://huggingface.co/spaces/anandkaman/controlmt-demo)
- **GitHub**: [github.com/anandkaman/ControlMT](https://github.com/anandkaman/ControlMT)