# ControlMT v2.3 — Deployment Guide ControlMT v2.3 is a small (139M params, ~280 MB bf16 safetensors) **encoder-decoder** seq2seq translator. This guide gives you verified recipes for every realistic deployment target, pinned to the exact versions we tested with. > **Architecture note** — ControlMT is a **bidirectional encoder + causal decoder + cross-attention** > seq2seq model, in the family of T5/mBART. It is **not** a decoder-only autoregressive LM. > That distinction matters: it determines what inference frameworks can run it. See > [Section 9 — Not directly supported](#9-not-directly-supported) below before reaching for vLLM / Ollama / GGUF. --- ## 1. Verified deployment matrix All numbers below are measured on the same machine (Linux, RTX 5060 Ti, Python 3.12.3, torch 2.10.0+cu128, transformers 4.57.6), `--num_beams=2`, 6 KN↔EN test pairs. Median of per-pair latencies. **Load time excludes first-time HF download (~30s for 280 MB).** | Recipe | Hardware | Latency / pair | Memory | Verified | |----------------------------|---------------------|----------------|-------------|----------| | **CPU bf16** (recommended) | Any x86/ARM, ≥1 GB | 0.51 s | 280 MB RAM | ✓ | | CPU fp32 | Any x86/ARM, ≥1 GB | 1.44 s | 560 MB RAM | ✓ | | **CPU int8-dynamic** (fastest CPU) | Any x86/ARM ≥1 GB | 0.28 s | ~140 MB RAM | ✓ | | **GPU fp16** (recommended) | ≥ 1 GB VRAM (T4/3050+/M-series) | 0.19 s | 404 MB VRAM | ✓ | | GPU bf16 | ≥ 1 GB VRAM | 0.19 s | 404 MB VRAM | ✓ | | GPU fp32 | ≥ 1 GB VRAM | 0.20 s | 793 MB VRAM | ✓ | | HuggingFace Space (Docker) | HF free-tier CPU | 3 – 15 s | shared | ✓ live | | HF Inference Endpoints | T4 / A10 | 0.2 – 0.5 s | managed | ⚠ untested but pattern works | | FastAPI self-host | any | matches device | matches dev | ✓ | | Docker | any with Docker | matches device | matches dev | ✓ | | ONNX Runtime | x86/ARM/GPU | ~0.5–2 s | varies | ⚠ experimental (manual export) | --- ## 2. Pinned versions (the exact stack we verified) ``` python 3.12.3 (3.10–3.12 supported, 3.10/3.11 also tested informally) torch 2.10.0 (CUDA build: 2.10.0+cu128 for GPU; CPU build: 2.10.0+cpu) transformers 4.57.6 (4.40+ should work; uses trust_remote_code) sentencepiece 0.2.1 safetensors 0.7.0 huggingface_hub 0.36.2 (>= 0.27 fine; 0.34+ recommended) # Optional: fastapi 0.129.0 (only if wrapping in an HTTP API) uvicorn 0.40.0 pydantic 2.12.5 ``` Minimum `requirements.txt` for the model alone: ```text torch>=2.0,<3 transformers>=4.40,<5 sentencepiece>=0.1.99 safetensors>=0.4 huggingface_hub>=0.27 ``` --- ## 3. Quick start — Python (CPU or GPU) ```python import torch from transformers import AutoModelForSeq2SeqLM, AutoTokenizer MODEL_ID = "anandkaman/controlmt-v2.3" tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True) model = AutoModelForSeq2SeqLM.from_pretrained( MODEL_ID, trust_remote_code=True, dtype=torch.bfloat16, # bf16 on CPU is 2.8× faster than fp32 (verified) ) device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model = model.to(device).eval() # Translate out = model.translate( "ನಾನು ಕನ್ನಡ ಮಾತನಾಡುತ್ತೇನೆ.", tokenizer=tokenizer, direction="kn2en", # or "en2kn" num_beams=2, # 1=fastest, 4 is a reasonable CPU cap, 6 ≈ no gain anti_lm_alpha=0.5, # contrastive decoding; 0 = vanilla beam max_length=200, ) print(out) # → "I speak Kannada." ``` **On GPU**, use `dtype=torch.float16` for slightly broader hardware compatibility (Volta/Pascal generation GPUs don't have native bf16). bf16 is preferred on Ampere+ (3000-series and newer). --- ## 4. Maximum-throughput CPU recipe — int8 dynamic quantization Best for ≤ 4 GB RAM devices, Raspberry Pi 5, or CPU-only containerized deployments. **~1.8× faster than CPU bf16, ~50% memory reduction, identical output on our 6-pair test.** ```python import torch from transformers import AutoModelForSeq2SeqLM, AutoTokenizer MODEL_ID = "anandkaman/controlmt-v2.3" tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True) model = AutoModelForSeq2SeqLM.from_pretrained(MODEL_ID, trust_remote_code=True) model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8) model.eval() out = model.translate("I speak Kannada.", tokenizer=tokenizer, direction="en2kn", num_beams=2, max_length=200) ``` No quality regression observed in our test set. For production deployment, you should re-validate on your own representative sentences — int8 dynamic occasionally drops 0.5–1 BLEU on long-tail outputs. --- ## 5. HuggingFace Space (Docker, FastAPI + static HTML) A live demo is at **[anandkaman/controlmt-demo](https://huggingface.co/spaces/anandkaman/controlmt-demo)**. The Space is a **Docker SDK** Space running a FastAPI backend + vanilla HTML/CSS/JS frontend. Full source ships in this release under `assets/space/` — clone-and-modify to build your own demo. Structure: ``` controlmt-demo/ ├── Dockerfile ├── main.py # FastAPI app: /api/translate, /api/rate, /api/health ├── pipeline.py # Opt-in logging (PII redaction, batched dataset upload) ├── requirements.txt ├── README.md # HF Space frontmatter (sdk: docker, app_port: 7860) └── static/ ├── index.html # Hero, settings, source/translation, ratings, examples ├── style.css # Mobile-first, 2-col on desktop ≥900px └── app.js # Vanilla JS — fetches /api/* from same origin ``` --- ## 6. FastAPI / Flask REST API (self-hosted) For a production API on your own infrastructure, the Space's `main.py` is the canonical reference. Minimal version: ```python # requirements: fastapi==0.129.0 uvicorn[standard]==0.40.0 pydantic==2.12.5 + the model stack import torch from fastapi import FastAPI from pydantic import BaseModel, Field from transformers import AutoModelForSeq2SeqLM, AutoTokenizer MODEL_ID = "anandkaman/controlmt-v2.3" app = FastAPI() tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True) model = AutoModelForSeq2SeqLM.from_pretrained( MODEL_ID, trust_remote_code=True, dtype=torch.bfloat16 if not torch.cuda.is_available() else torch.float16, ) model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval() class Req(BaseModel): text: str = Field(..., min_length=1, max_length=5000) direction: str = "kn2en" num_beams: int = 2 @app.post("/translate") def translate(req: Req): out = model.translate( req.text, tokenizer=tokenizer, direction=req.direction, num_beams=req.num_beams, anti_lm_alpha=0.5, max_length=200, ) return {"translation": out} ``` ```bash uvicorn main:app --host 0.0.0.0 --port 8000 --workers 1 ``` > Use `--workers 1` for a single-GPU box: each worker loads its own model copy. For > CPU deployments you can scale to `--workers 2` if you have ≥1 GB RAM per worker. --- ## 7. Docker (self-hosted) The Space's `Dockerfile` is the verified reference. For non-HF deployment (your own server / Kubernetes / Cloud Run / Fly.io): ```dockerfile FROM python:3.12-slim ENV DEBIAN_FRONTEND=noninteractive \ PYTHONDONTWRITEBYTECODE=1 \ PYTHONUNBUFFERED=1 \ PIP_NO_CACHE_DIR=1 RUN useradd -m -u 1000 user USER user ENV PATH=/home/user/.local/bin:$PATH \ HF_HOME=/home/user/.cache/huggingface WORKDIR /app COPY --chown=user:user requirements.txt ./ RUN pip install --user --no-cache-dir -r requirements.txt COPY --chown=user:user . ./ EXPOSE 8000 CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"] ``` For GPU containers, swap base image to `nvidia/cuda:12.8.0-cudnn8-runtime-ubuntu22.04` and install Python + pip on top. --- ## 8. HuggingFace Inference Endpoints (managed cloud) For zero-DevOps managed deployment: 1. Open [HF Inference Endpoints](https://ui.endpoints.huggingface.co/) 2. Point at `anandkaman/controlmt-v2.3` 3. In **Advanced configuration**: enable `trust_remote_code=true` 4. Recommended: **GPU Small (T4)** — ~$0.40/hr, ~0.3 s per pair after warmup **Not yet end-to-end verified by us** — recipe follows the standard HF custom-code pattern; if you hit an issue, file at the GitHub repo. --- ## 9. Not directly supported These platforms are popular for LLM serving, but ControlMT's encoder-decoder architecture is not compatible without significant adapter work. If you find yourself reaching for one of these, the right answer is **use Python+Transformers or FastAPI instead** — they give the same throughput on this model size. | Platform | Reason | Recommended alternative | |-----------------|----------------------------------------------------------------------------|-------------------------| | **vLLM** | Supports BART/T5 via known class names only; custom `trust_remote_code` models fall outside its optimized seq2seq path | FastAPI wrapper (Section 6) | | **Ollama** | Uses GGUF format + llama.cpp — both are decoder-only architectures, no encoder pass or cross-attention | FastAPI wrapper (Section 6) | | **llama.cpp / GGUF** | Same architectural reason as Ollama | FastAPI wrapper (Section 6) | | **HF TGI** | Same `trust_remote_code` limitation as vLLM | HF Inference Endpoints (Section 8) or FastAPI (Section 6) | | **bitsandbytes int8** | Our model code doesn't route int8↔fp16 dtype conversions through linear layers in the way bnb expects. Verified failure: `RuntimeError: self and mat2 must have the same dtype, but got Half and Char` | CPU int8 dynamic (Section 4) — same memory savings, no patching required | At v2.3's size (139M params, 0.19 s/pair on a $300 GPU), the optimizations these tools provide (KV-cache reuse, paged attention, continuous batching) are dominated by request overhead and not worth the integration cost. ### If you absolutely need GGUF / Ollama / vLLM compatibility The realistic path is **knowledge distillation**: train a small decoder-only LLM (~200–600M parameters, Llama-style or Phi-style architecture) on ControlMT's outputs as targets. The student is decoder-only, so it's GGUF-compatible, runs natively in Ollama, can be quantized with AWQ/GPTQ, and integrates with the entire LLM serving ecosystem. - **Quality trade-off**: decoder-only is less parameter-efficient than seq2seq for translation. A ~400–600M decoder-only student is the realistic minimum to match v2.3's quality on common cases. - **Cost**: ~1 week to set up distillation + 2–3 days of student training + re-eval. Comparable to a normal model retrain. - **Win**: a sibling repo (planned `anandkaman/controlmt-v3.0-llama`) that ships as native GGUF + works with Ollama directly. Same KN↔EN coverage, ~3–5× larger model file, but 10× simpler deployment for end users. - **Prior art**: AI4Bharat does this exact pattern (IndicTrans2 → IndicTrans2-distilled). The architecture-shift loss is real but recoverable. This is a **v3.0+ roadmap item**, not currently available. If you're building infrastructure around ControlMT today and absolutely need an Ollama-deployable variant, watch the v3 changelog or open an issue at the GitHub repo to signal demand. --- ## 10. ONNX Runtime (experimental) ONNX export is possible but **requires manual encoder/decoder splitting** and a re-implementation of beam search in your target runtime — `optimum-onnx` doesn't know about ControlMT's custom architecture and won't auto-trace it. Sketch: ```python import torch # 1. Export encoder: takes (input_ids, attn_mask, direction_id, control_id) → encoder_states torch.onnx.export(model.controlmt.encoder, (x, mask, dir_id, ctrl_id), "encoder.onnx", opset_version=17, ...) # 2. Export decoder step: takes (encoder_states, decoder_input_ids, cache) → next_token_logits torch.onnx.export(decoder_step_fn, (...), "decoder_step.onnx", ...) # 3. Reimplement beam search in your runtime (Python+onnxruntime, JS, C++) by calling # encoder once and decoder_step N×beams times. ``` Contributions of a working ONNX recipe + an `assets/onnx_export.py` script welcome at [github.com/anandkaman/ControlMT](https://github.com/anandkaman/ControlMT). --- ## 11. Batched inference (high-throughput document translation) For document-level workloads (1000+ sentences) on GPU: ```python # Sketch — beam search runs per-sentence, but you can dispatch in parallel by chunking. from concurrent.futures import ThreadPoolExecutor def translate_batch(texts, direction="kn2en", num_beams=2): with ThreadPoolExecutor(max_workers=4) as ex: return list(ex.map( lambda t: model.translate(t, tokenizer=tokenizer, direction=direction, num_beams=num_beams, max_length=200), texts, )) ``` Typical RTX 3060+ throughput: **~30–50 sentences/second at beam=2** in this configuration. For genuinely batched (cross-sentence) beam search you'd need to extend `model.translate` to handle variable EOS-completion, which our v2.3 weights support but the wrapper exposes only one sentence at a time. --- ## 12. Windows / cross-platform notes The Python recipe (Section 3) works on Windows, macOS, and Linux without changes: - **Windows native** — install Python 3.10–3.12, `pip install` the pinned versions, run the script - **macOS (Apple Silicon)** — works on CPU. For GPU, `device="mps"` with `dtype=torch.float16` - **WSL2** — use the Linux recipe, identical performance to native Linux For a Windows desktop app that bundles ControlMT: - **PyInstaller / Briefcase**: package the Python script + cached HF model files - **Docker Desktop**: run the Docker recipe (Section 7) on Windows hosts - **ONNX**: cross-platform but see Section 10 — experimental --- ## 13. Pre-launch checklist - [ ] Smoke-test on your hardware with `python verify_deployment.py --device ` (script lives at `assets/scripts/verify_deployment.py`) - [ ] **For form-data / KYC use**: add the PAN-postprocessing regex (model card Section 6 Limitations) - [ ] Rate-limit your endpoint — model accepts ≤ 5000 chars per request; truncate or chunk - [ ] Health endpoint that round-trips a known pair (e.g. `"hello" → "ಹಲೋ"`) - [ ] Memory headroom: fp32 ≥ 700 MB, bf16/fp16 ≥ 400 MB, int8 dynamic ≥ 200 MB - [ ] Idempotent restarts — model is stateless; safe to redeploy without warmup --- ## Related docs - **Model card**: [README.md](README.md) — FLORES/IN22 scores, limitations, intended use - **Training methodology**: [TRAINING_GUIDE.md](TRAINING_GUIDE.md) — corpus + filtering + training principles - **Privacy policy** (hosted demo): [PRIVACY.md](PRIVACY.md) - **License**: [LICENSE](LICENSE) — Apache 2.0 - **Live demo**: [huggingface.co/spaces/anandkaman/controlmt-demo](https://huggingface.co/spaces/anandkaman/controlmt-demo) - **GitHub**: [github.com/anandkaman/ControlMT](https://github.com/anandkaman/ControlMT)