Instructions to use anandkaman/controlmt-v2.3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anandkaman/controlmt-v2.3 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="anandkaman/controlmt-v2.3", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("anandkaman/controlmt-v2.3", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download DEPLOYMENT.md from anandkaman/controlmt-v2.3: direct link, hf CLI and curl.
- Browser
- Download file 15.6 kB
-
https://huggingface.co/anandkaman/controlmt-v2.3/resolve/main/DEPLOYMENT.md
- Command line
-
hf download hf://anandkaman/controlmt-v2.3/DEPLOYMENT.md
-
curl -L -o DEPLOYMENT.md https://huggingface.co/anandkaman/controlmt-v2.3/resolve/main/DEPLOYMENT.md
ControlMT v2.3 — Deployment Guide
ControlMT v2.3 is a small (139M params, ~280 MB bf16 safetensors) encoder-decoder seq2seq translator. This guide gives you verified recipes for every realistic deployment target, pinned to the exact versions we tested with.
Architecture note — ControlMT is a bidirectional encoder + causal decoder + cross-attention seq2seq model, in the family of T5/mBART. It is not a decoder-only autoregressive LM. That distinction matters: it determines what inference frameworks can run it. See Section 9 — Not directly supported below before reaching for vLLM / Ollama / GGUF.
1. Verified deployment matrix
All numbers below are measured on the same machine (Linux, RTX 5060 Ti, Python 3.12.3,
torch 2.10.0+cu128, transformers 4.57.6), --num_beams=2, 6 KN↔EN test pairs. Median
of per-pair latencies. Load time excludes first-time HF download (~30s for 280 MB).
| Recipe | Hardware | Latency / pair | Memory | Verified |
|---|---|---|---|---|
| CPU bf16 (recommended) | Any x86/ARM, ≥1 GB | 0.51 s | 280 MB RAM | ✓ |
| CPU fp32 | Any x86/ARM, ≥1 GB | 1.44 s | 560 MB RAM | ✓ |
| CPU int8-dynamic (fastest CPU) | Any x86/ARM ≥1 GB | 0.28 s | ~140 MB RAM | ✓ |
| GPU fp16 (recommended) | ≥ 1 GB VRAM (T4/3050+/M-series) | 0.19 s | 404 MB VRAM | ✓ |
| GPU bf16 | ≥ 1 GB VRAM | 0.19 s | 404 MB VRAM | ✓ |
| GPU fp32 | ≥ 1 GB VRAM | 0.20 s | 793 MB VRAM | ✓ |
| HuggingFace Space (Docker) | HF free-tier CPU | 3 – 15 s | shared | ✓ live |
| HF Inference Endpoints | T4 / A10 | 0.2 – 0.5 s | managed | ⚠ untested but pattern works |
| FastAPI self-host | any | matches device | matches dev | ✓ |
| Docker | any with Docker | matches device | matches dev | ✓ |
| ONNX Runtime | x86/ARM/GPU | ~0.5–2 s | varies | ⚠ experimental (manual export) |
2. Pinned versions (the exact stack we verified)
python 3.12.3 (3.10–3.12 supported, 3.10/3.11 also tested informally)
torch 2.10.0 (CUDA build: 2.10.0+cu128 for GPU; CPU build: 2.10.0+cpu)
transformers 4.57.6 (4.40+ should work; uses trust_remote_code)
sentencepiece 0.2.1
safetensors 0.7.0
huggingface_hub 0.36.2 (>= 0.27 fine; 0.34+ recommended)
# Optional:
fastapi 0.129.0 (only if wrapping in an HTTP API)
uvicorn 0.40.0
pydantic 2.12.5
Minimum requirements.txt for the model alone:
torch>=2.0,<3
transformers>=4.40,<5
sentencepiece>=0.1.99
safetensors>=0.4
huggingface_hub>=0.27
3. Quick start — Python (CPU or GPU)
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
MODEL_ID = "anandkaman/controlmt-v2.3"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16, # bf16 on CPU is 2.8× faster than fp32 (verified)
)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device).eval()
# Translate
out = model.translate(
"ನಾನು ಕನ್ನಡ ಮಾತನಾಡುತ್ತೇನೆ.",
tokenizer=tokenizer,
direction="kn2en", # or "en2kn"
num_beams=2, # 1=fastest, 4 is a reasonable CPU cap, 6 ≈ no gain
anti_lm_alpha=0.5, # contrastive decoding; 0 = vanilla beam
max_length=200,
)
print(out) # → "I speak Kannada."
On GPU, use dtype=torch.float16 for slightly broader hardware compatibility
(Volta/Pascal generation GPUs don't have native bf16). bf16 is preferred on Ampere+ (3000-series and newer).
4. Maximum-throughput CPU recipe — int8 dynamic quantization
Best for ≤ 4 GB RAM devices, Raspberry Pi 5, or CPU-only containerized deployments. ~1.8× faster than CPU bf16, ~50% memory reduction, identical output on our 6-pair test.
import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
MODEL_ID = "anandkaman/controlmt-v2.3"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(MODEL_ID, trust_remote_code=True)
model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)
model.eval()
out = model.translate("I speak Kannada.", tokenizer=tokenizer, direction="en2kn",
num_beams=2, max_length=200)
No quality regression observed in our test set. For production deployment, you should re-validate on your own representative sentences — int8 dynamic occasionally drops 0.5–1 BLEU on long-tail outputs.
5. HuggingFace Space (Docker, FastAPI + static HTML)
A live demo is at anandkaman/controlmt-demo.
The Space is a Docker SDK Space running a FastAPI backend + vanilla HTML/CSS/JS
frontend. Full source ships in this release under assets/space/ — clone-and-modify
to build your own demo.
Structure:
controlmt-demo/
├── Dockerfile
├── main.py # FastAPI app: /api/translate, /api/rate, /api/health
├── pipeline.py # Opt-in logging (PII redaction, batched dataset upload)
├── requirements.txt
├── README.md # HF Space frontmatter (sdk: docker, app_port: 7860)
└── static/
├── index.html # Hero, settings, source/translation, ratings, examples
├── style.css # Mobile-first, 2-col on desktop ≥900px
└── app.js # Vanilla JS — fetches /api/* from same origin
6. FastAPI / Flask REST API (self-hosted)
For a production API on your own infrastructure, the Space's main.py is the canonical
reference. Minimal version:
# requirements: fastapi==0.129.0 uvicorn[standard]==0.40.0 pydantic==2.12.5 + the model stack
import torch
from fastapi import FastAPI
from pydantic import BaseModel, Field
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
MODEL_ID = "anandkaman/controlmt-v2.3"
app = FastAPI()
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(
MODEL_ID, trust_remote_code=True,
dtype=torch.bfloat16 if not torch.cuda.is_available() else torch.float16,
)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()
class Req(BaseModel):
text: str = Field(..., min_length=1, max_length=5000)
direction: str = "kn2en"
num_beams: int = 2
@app.post("/translate")
def translate(req: Req):
out = model.translate(
req.text, tokenizer=tokenizer, direction=req.direction,
num_beams=req.num_beams, anti_lm_alpha=0.5, max_length=200,
)
return {"translation": out}
uvicorn main:app --host 0.0.0.0 --port 8000 --workers 1
Use
--workers 1for a single-GPU box: each worker loads its own model copy. For CPU deployments you can scale to--workers 2if you have ≥1 GB RAM per worker.
7. Docker (self-hosted)
The Space's Dockerfile is the verified reference. For non-HF deployment
(your own server / Kubernetes / Cloud Run / Fly.io):
FROM python:3.12-slim
ENV DEBIAN_FRONTEND=noninteractive \
PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1 \
PIP_NO_CACHE_DIR=1
RUN useradd -m -u 1000 user
USER user
ENV PATH=/home/user/.local/bin:$PATH \
HF_HOME=/home/user/.cache/huggingface
WORKDIR /app
COPY --chown=user:user requirements.txt ./
RUN pip install --user --no-cache-dir -r requirements.txt
COPY --chown=user:user . ./
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
For GPU containers, swap base image to nvidia/cuda:12.8.0-cudnn8-runtime-ubuntu22.04
and install Python + pip on top.
8. HuggingFace Inference Endpoints (managed cloud)
For zero-DevOps managed deployment:
- Open HF Inference Endpoints
- Point at
anandkaman/controlmt-v2.3 - In Advanced configuration: enable
trust_remote_code=true - Recommended: GPU Small (T4) — ~$0.40/hr, ~0.3 s per pair after warmup
Not yet end-to-end verified by us — recipe follows the standard HF custom-code pattern; if you hit an issue, file at the GitHub repo.
9. Not directly supported
These platforms are popular for LLM serving, but ControlMT's encoder-decoder architecture is not compatible without significant adapter work. If you find yourself reaching for one of these, the right answer is use Python+Transformers or FastAPI instead — they give the same throughput on this model size.
| Platform | Reason | Recommended alternative |
|---|---|---|
| vLLM | Supports BART/T5 via known class names only; custom trust_remote_code models fall outside its optimized seq2seq path |
FastAPI wrapper (Section 6) |
| Ollama | Uses GGUF format + llama.cpp — both are decoder-only architectures, no encoder pass or cross-attention | FastAPI wrapper (Section 6) |
| llama.cpp / GGUF | Same architectural reason as Ollama | FastAPI wrapper (Section 6) |
| HF TGI | Same trust_remote_code limitation as vLLM |
HF Inference Endpoints (Section 8) or FastAPI (Section 6) |
| bitsandbytes int8 | Our model code doesn't route int8↔fp16 dtype conversions through linear layers in the way bnb expects. Verified failure: RuntimeError: self and mat2 must have the same dtype, but got Half and Char |
CPU int8 dynamic (Section 4) — same memory savings, no patching required |
At v2.3's size (139M params, 0.19 s/pair on a $300 GPU), the optimizations these tools provide (KV-cache reuse, paged attention, continuous batching) are dominated by request overhead and not worth the integration cost.
If you absolutely need GGUF / Ollama / vLLM compatibility
The realistic path is knowledge distillation: train a small decoder-only LLM (~200–600M parameters, Llama-style or Phi-style architecture) on ControlMT's outputs as targets. The student is decoder-only, so it's GGUF-compatible, runs natively in Ollama, can be quantized with AWQ/GPTQ, and integrates with the entire LLM serving ecosystem.
- Quality trade-off: decoder-only is less parameter-efficient than seq2seq for translation. A ~400–600M decoder-only student is the realistic minimum to match v2.3's quality on common cases.
- Cost: ~1 week to set up distillation + 2–3 days of student training + re-eval. Comparable to a normal model retrain.
- Win: a sibling repo (planned
anandkaman/controlmt-v3.0-llama) that ships as native GGUF- works with Ollama directly. Same KN↔EN coverage, ~3–5× larger model file, but 10× simpler deployment for end users.
- Prior art: AI4Bharat does this exact pattern (IndicTrans2 → IndicTrans2-distilled). The architecture-shift loss is real but recoverable.
This is a v3.0+ roadmap item, not currently available. If you're building infrastructure around ControlMT today and absolutely need an Ollama-deployable variant, watch the v3 changelog or open an issue at the GitHub repo to signal demand.
10. ONNX Runtime (experimental)
ONNX export is possible but requires manual encoder/decoder splitting and a re-implementation
of beam search in your target runtime — optimum-onnx doesn't know about ControlMT's custom
architecture and won't auto-trace it.
Sketch:
import torch
# 1. Export encoder: takes (input_ids, attn_mask, direction_id, control_id) → encoder_states
torch.onnx.export(model.controlmt.encoder, (x, mask, dir_id, ctrl_id),
"encoder.onnx", opset_version=17, ...)
# 2. Export decoder step: takes (encoder_states, decoder_input_ids, cache) → next_token_logits
torch.onnx.export(decoder_step_fn, (...), "decoder_step.onnx", ...)
# 3. Reimplement beam search in your runtime (Python+onnxruntime, JS, C++) by calling
# encoder once and decoder_step N×beams times.
Contributions of a working ONNX recipe + an assets/onnx_export.py script welcome at
github.com/anandkaman/ControlMT.
11. Batched inference (high-throughput document translation)
For document-level workloads (1000+ sentences) on GPU:
# Sketch — beam search runs per-sentence, but you can dispatch in parallel by chunking.
from concurrent.futures import ThreadPoolExecutor
def translate_batch(texts, direction="kn2en", num_beams=2):
with ThreadPoolExecutor(max_workers=4) as ex:
return list(ex.map(
lambda t: model.translate(t, tokenizer=tokenizer, direction=direction,
num_beams=num_beams, max_length=200),
texts,
))
Typical RTX 3060+ throughput: ~30–50 sentences/second at beam=2 in this configuration.
For genuinely batched (cross-sentence) beam search you'd need to extend model.translate
to handle variable EOS-completion, which our v2.3 weights support but the wrapper exposes
only one sentence at a time.
12. Windows / cross-platform notes
The Python recipe (Section 3) works on Windows, macOS, and Linux without changes:
- Windows native — install Python 3.10–3.12,
pip installthe pinned versions, run the script - macOS (Apple Silicon) — works on CPU. For GPU,
device="mps"withdtype=torch.float16 - WSL2 — use the Linux recipe, identical performance to native Linux
For a Windows desktop app that bundles ControlMT:
- PyInstaller / Briefcase: package the Python script + cached HF model files
- Docker Desktop: run the Docker recipe (Section 7) on Windows hosts
- ONNX: cross-platform but see Section 10 — experimental
13. Pre-launch checklist
- Smoke-test on your hardware with
python verify_deployment.py --device <cpu|cuda>(script lives atassets/scripts/verify_deployment.py) - For form-data / KYC use: add the PAN-postprocessing regex (model card Section 6 Limitations)
- Rate-limit your endpoint — model accepts ≤ 5000 chars per request; truncate or chunk
- Health endpoint that round-trips a known pair (e.g.
"hello" → "ಹಲೋ") - Memory headroom: fp32 ≥ 700 MB, bf16/fp16 ≥ 400 MB, int8 dynamic ≥ 200 MB
- Idempotent restarts — model is stateless; safe to redeploy without warmup
Related docs
- Model card: README.md — FLORES/IN22 scores, limitations, intended use
- Training methodology: TRAINING_GUIDE.md — corpus + filtering + training principles
- Privacy policy (hosted demo): PRIVACY.md
- License: LICENSE — Apache 2.0
- Live demo: huggingface.co/spaces/anandkaman/controlmt-demo
- GitHub: github.com/anandkaman/ControlMT