controlmt-v2.3 / DEPLOYMENT.md
anandkaman's picture
DEPLOYMENT.md §9: add distillation-to-decoder-only path as v3.0+ option for GGUF/Ollama/vLLM compatibility
fde63ef verified
|
Raw History Blame Contribute Delete
15.6 kB

ControlMT v2.3 — Deployment Guide

ControlMT v2.3 is a small (139M params, ~280 MB bf16 safetensors) encoder-decoder seq2seq translator. This guide gives you verified recipes for every realistic deployment target, pinned to the exact versions we tested with.

Architecture note — ControlMT is a bidirectional encoder + causal decoder + cross-attention seq2seq model, in the family of T5/mBART. It is not a decoder-only autoregressive LM. That distinction matters: it determines what inference frameworks can run it. See Section 9 — Not directly supported below before reaching for vLLM / Ollama / GGUF.


1. Verified deployment matrix

All numbers below are measured on the same machine (Linux, RTX 5060 Ti, Python 3.12.3, torch 2.10.0+cu128, transformers 4.57.6), --num_beams=2, 6 KN↔EN test pairs. Median of per-pair latencies. Load time excludes first-time HF download (~30s for 280 MB).

Recipe Hardware Latency / pair Memory Verified
CPU bf16 (recommended) Any x86/ARM, ≥1 GB 0.51 s 280 MB RAM ✓
CPU fp32 Any x86/ARM, ≥1 GB 1.44 s 560 MB RAM ✓
CPU int8-dynamic (fastest CPU) Any x86/ARM ≥1 GB 0.28 s ~140 MB RAM ✓
GPU fp16 (recommended) ≥ 1 GB VRAM (T4/3050+/M-series) 0.19 s 404 MB VRAM ✓
GPU bf16 ≥ 1 GB VRAM 0.19 s 404 MB VRAM ✓
GPU fp32 ≥ 1 GB VRAM 0.20 s 793 MB VRAM ✓
HuggingFace Space (Docker) HF free-tier CPU 3 – 15 s shared ✓ live
HF Inference Endpoints T4 / A10 0.2 – 0.5 s managed ⚠ untested but pattern works
FastAPI self-host any matches device matches dev ✓
Docker any with Docker matches device matches dev ✓
ONNX Runtime x86/ARM/GPU ~0.5–2 s varies ⚠ experimental (manual export)

2. Pinned versions (the exact stack we verified)

python              3.12.3   (3.10–3.12 supported, 3.10/3.11 also tested informally)
torch               2.10.0   (CUDA build: 2.10.0+cu128 for GPU; CPU build: 2.10.0+cpu)
transformers        4.57.6   (4.40+ should work; uses trust_remote_code)
sentencepiece       0.2.1
safetensors         0.7.0
huggingface_hub     0.36.2   (>= 0.27 fine; 0.34+ recommended)
# Optional:
fastapi             0.129.0  (only if wrapping in an HTTP API)
uvicorn             0.40.0
pydantic            2.12.5

Minimum requirements.txt for the model alone:

torch>=2.0,<3
transformers>=4.40,<5
sentencepiece>=0.1.99
safetensors>=0.4
huggingface_hub>=0.27

3. Quick start — Python (CPU or GPU)

import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

MODEL_ID = "anandkaman/controlmt-v2.3"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    dtype=torch.bfloat16,                 # bf16 on CPU is 2.8× faster than fp32 (verified)
)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device).eval()

# Translate
out = model.translate(
    "ನಾನು ಕನ್ನಡ ಮಾತನಾಡುತ್ತೇನೆ.",
    tokenizer=tokenizer,
    direction="kn2en",       # or "en2kn"
    num_beams=2,             # 1=fastest, 4 is a reasonable CPU cap, 6 ≈ no gain
    anti_lm_alpha=0.5,       # contrastive decoding; 0 = vanilla beam
    max_length=200,
)
print(out)   # → "I speak Kannada."

On GPU, use dtype=torch.float16 for slightly broader hardware compatibility (Volta/Pascal generation GPUs don't have native bf16). bf16 is preferred on Ampere+ (3000-series and newer).


4. Maximum-throughput CPU recipe — int8 dynamic quantization

Best for ≤ 4 GB RAM devices, Raspberry Pi 5, or CPU-only containerized deployments. ~1.8× faster than CPU bf16, ~50% memory reduction, identical output on our 6-pair test.

import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

MODEL_ID = "anandkaman/controlmt-v2.3"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(MODEL_ID, trust_remote_code=True)
model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)
model.eval()

out = model.translate("I speak Kannada.", tokenizer=tokenizer, direction="en2kn",
                      num_beams=2, max_length=200)

No quality regression observed in our test set. For production deployment, you should re-validate on your own representative sentences — int8 dynamic occasionally drops 0.5–1 BLEU on long-tail outputs.


5. HuggingFace Space (Docker, FastAPI + static HTML)

A live demo is at anandkaman/controlmt-demo. The Space is a Docker SDK Space running a FastAPI backend + vanilla HTML/CSS/JS frontend. Full source ships in this release under assets/space/ — clone-and-modify to build your own demo.

Structure:

controlmt-demo/
├── Dockerfile
├── main.py                # FastAPI app: /api/translate, /api/rate, /api/health
├── pipeline.py            # Opt-in logging (PII redaction, batched dataset upload)
├── requirements.txt
├── README.md              # HF Space frontmatter (sdk: docker, app_port: 7860)
└── static/
    ├── index.html         # Hero, settings, source/translation, ratings, examples
    ├── style.css          # Mobile-first, 2-col on desktop ≥900px
    └── app.js             # Vanilla JS — fetches /api/* from same origin

6. FastAPI / Flask REST API (self-hosted)

For a production API on your own infrastructure, the Space's main.py is the canonical reference. Minimal version:

# requirements: fastapi==0.129.0 uvicorn[standard]==0.40.0 pydantic==2.12.5 + the model stack
import torch
from fastapi import FastAPI
from pydantic import BaseModel, Field
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

MODEL_ID = "anandkaman/controlmt-v2.3"
app = FastAPI()

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(
    MODEL_ID, trust_remote_code=True,
    dtype=torch.bfloat16 if not torch.cuda.is_available() else torch.float16,
)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()

class Req(BaseModel):
    text: str = Field(..., min_length=1, max_length=5000)
    direction: str = "kn2en"
    num_beams: int = 2

@app.post("/translate")
def translate(req: Req):
    out = model.translate(
        req.text, tokenizer=tokenizer, direction=req.direction,
        num_beams=req.num_beams, anti_lm_alpha=0.5, max_length=200,
    )
    return {"translation": out}
uvicorn main:app --host 0.0.0.0 --port 8000 --workers 1

Use --workers 1 for a single-GPU box: each worker loads its own model copy. For CPU deployments you can scale to --workers 2 if you have ≥1 GB RAM per worker.


7. Docker (self-hosted)

The Space's Dockerfile is the verified reference. For non-HF deployment (your own server / Kubernetes / Cloud Run / Fly.io):

FROM python:3.12-slim

ENV DEBIAN_FRONTEND=noninteractive \
    PYTHONDONTWRITEBYTECODE=1 \
    PYTHONUNBUFFERED=1 \
    PIP_NO_CACHE_DIR=1

RUN useradd -m -u 1000 user
USER user
ENV PATH=/home/user/.local/bin:$PATH \
    HF_HOME=/home/user/.cache/huggingface

WORKDIR /app
COPY --chown=user:user requirements.txt ./
RUN pip install --user --no-cache-dir -r requirements.txt
COPY --chown=user:user . ./

EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]

For GPU containers, swap base image to nvidia/cuda:12.8.0-cudnn8-runtime-ubuntu22.04 and install Python + pip on top.


8. HuggingFace Inference Endpoints (managed cloud)

For zero-DevOps managed deployment:

  1. Open HF Inference Endpoints
  2. Point at anandkaman/controlmt-v2.3
  3. In Advanced configuration: enable trust_remote_code=true
  4. Recommended: GPU Small (T4) — ~$0.40/hr, ~0.3 s per pair after warmup

Not yet end-to-end verified by us — recipe follows the standard HF custom-code pattern; if you hit an issue, file at the GitHub repo.


9. Not directly supported

These platforms are popular for LLM serving, but ControlMT's encoder-decoder architecture is not compatible without significant adapter work. If you find yourself reaching for one of these, the right answer is use Python+Transformers or FastAPI instead — they give the same throughput on this model size.

Platform Reason Recommended alternative
vLLM Supports BART/T5 via known class names only; custom trust_remote_code models fall outside its optimized seq2seq path FastAPI wrapper (Section 6)
Ollama Uses GGUF format + llama.cpp — both are decoder-only architectures, no encoder pass or cross-attention FastAPI wrapper (Section 6)
llama.cpp / GGUF Same architectural reason as Ollama FastAPI wrapper (Section 6)
HF TGI Same trust_remote_code limitation as vLLM HF Inference Endpoints (Section 8) or FastAPI (Section 6)
bitsandbytes int8 Our model code doesn't route int8↔fp16 dtype conversions through linear layers in the way bnb expects. Verified failure: RuntimeError: self and mat2 must have the same dtype, but got Half and Char CPU int8 dynamic (Section 4) — same memory savings, no patching required

At v2.3's size (139M params, 0.19 s/pair on a $300 GPU), the optimizations these tools provide (KV-cache reuse, paged attention, continuous batching) are dominated by request overhead and not worth the integration cost.

If you absolutely need GGUF / Ollama / vLLM compatibility

The realistic path is knowledge distillation: train a small decoder-only LLM (~200–600M parameters, Llama-style or Phi-style architecture) on ControlMT's outputs as targets. The student is decoder-only, so it's GGUF-compatible, runs natively in Ollama, can be quantized with AWQ/GPTQ, and integrates with the entire LLM serving ecosystem.

  • Quality trade-off: decoder-only is less parameter-efficient than seq2seq for translation. A ~400–600M decoder-only student is the realistic minimum to match v2.3's quality on common cases.
  • Cost: ~1 week to set up distillation + 2–3 days of student training + re-eval. Comparable to a normal model retrain.
  • Win: a sibling repo (planned anandkaman/controlmt-v3.0-llama) that ships as native GGUF
    • works with Ollama directly. Same KN↔EN coverage, ~3–5× larger model file, but 10× simpler deployment for end users.
  • Prior art: AI4Bharat does this exact pattern (IndicTrans2 → IndicTrans2-distilled). The architecture-shift loss is real but recoverable.

This is a v3.0+ roadmap item, not currently available. If you're building infrastructure around ControlMT today and absolutely need an Ollama-deployable variant, watch the v3 changelog or open an issue at the GitHub repo to signal demand.


10. ONNX Runtime (experimental)

ONNX export is possible but requires manual encoder/decoder splitting and a re-implementation of beam search in your target runtime — optimum-onnx doesn't know about ControlMT's custom architecture and won't auto-trace it.

Sketch:

import torch
# 1. Export encoder: takes (input_ids, attn_mask, direction_id, control_id) → encoder_states
torch.onnx.export(model.controlmt.encoder, (x, mask, dir_id, ctrl_id),
                  "encoder.onnx", opset_version=17, ...)
# 2. Export decoder step: takes (encoder_states, decoder_input_ids, cache) → next_token_logits
torch.onnx.export(decoder_step_fn, (...), "decoder_step.onnx", ...)
# 3. Reimplement beam search in your runtime (Python+onnxruntime, JS, C++) by calling
#    encoder once and decoder_step N×beams times.

Contributions of a working ONNX recipe + an assets/onnx_export.py script welcome at github.com/anandkaman/ControlMT.


11. Batched inference (high-throughput document translation)

For document-level workloads (1000+ sentences) on GPU:

# Sketch — beam search runs per-sentence, but you can dispatch in parallel by chunking.
from concurrent.futures import ThreadPoolExecutor

def translate_batch(texts, direction="kn2en", num_beams=2):
    with ThreadPoolExecutor(max_workers=4) as ex:
        return list(ex.map(
            lambda t: model.translate(t, tokenizer=tokenizer, direction=direction,
                                       num_beams=num_beams, max_length=200),
            texts,
        ))

Typical RTX 3060+ throughput: ~30–50 sentences/second at beam=2 in this configuration. For genuinely batched (cross-sentence) beam search you'd need to extend model.translate to handle variable EOS-completion, which our v2.3 weights support but the wrapper exposes only one sentence at a time.


12. Windows / cross-platform notes

The Python recipe (Section 3) works on Windows, macOS, and Linux without changes:

  • Windows native — install Python 3.10–3.12, pip install the pinned versions, run the script
  • macOS (Apple Silicon) — works on CPU. For GPU, device="mps" with dtype=torch.float16
  • WSL2 — use the Linux recipe, identical performance to native Linux

For a Windows desktop app that bundles ControlMT:

  • PyInstaller / Briefcase: package the Python script + cached HF model files
  • Docker Desktop: run the Docker recipe (Section 7) on Windows hosts
  • ONNX: cross-platform but see Section 10 — experimental

13. Pre-launch checklist

  • Smoke-test on your hardware with python verify_deployment.py --device <cpu|cuda> (script lives at assets/scripts/verify_deployment.py)
  • For form-data / KYC use: add the PAN-postprocessing regex (model card Section 6 Limitations)
  • Rate-limit your endpoint — model accepts ≤ 5000 chars per request; truncate or chunk
  • Health endpoint that round-trips a known pair (e.g. "hello" → "ಹಲೋ")
  • Memory headroom: fp32 ≥ 700 MB, bf16/fp16 ≥ 400 MB, int8 dynamic ≥ 200 MB
  • Idempotent restarts — model is stateless; safe to redeploy without warmup

Related docs