limite-1b-violetto-mlx-4bit-hq

paradigma-inc/limite-1b-violetto converted for MLX on Apple Silicon. 4bit-hq - 4-bit body with the tied output head and value embeddings at 6-bit. Same speed as plain 4bit, 3.5x lower KL, for 55 MB more.

size on disk 0.61 GiB
effective bits/weight 5.04
KL divergence from bf16 0.07134 nats
top-1 agreement with bf16 88.8%
generation speed 72 tok/s
peak memory 0.66 GiB

Speed and memory are measured on an M2 (8 GB) with mlx-lm 0.31.3. Fidelity is measured on the deployed distribution rather than as raw perplexity, because this model's head squashes logits into [0, 23] via 23 * sigmoid((raw + 5) / 7.5) before sampling ever sees them.

Setup

limite is not one of mlx-lm's built-in architectures, so stock mlx-lm needs one extra package to know how to build this model. It registers itself through a sys.meta_path hook, so mlx-lm's own files are never modified and an mlx-lm upgrade cannot break it:

pip install mlx-lm "limite-mlx @ https://huggingface.co/pierjoe/limite-mlx/resolve/main/limite_mlx-0.1.0-py3-none-any.whl"

Nothing to import and nothing to configure after that. The same file is also in this repo as limite.py if you would rather drop it into mlx_lm/models/ yourself.

Usage

mlx_lm.generate --model pierjoe/limite-1b-violetto-mlx-4bit-hq \
  --prompt "What is the remainder when 7^2026 is divided by 13?" \
  --max-tokens 3000 --temp 0.6 --top-p 0.95
from mlx_lm import generate, load

model, tokenizer = load("pierjoe/limite-1b-violetto-mlx-4bit-hq")
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "How many positive divisors does 360 have?"}],
    add_generation_prompt=True,
    tokenize=False,
)
print(generate(model, tokenizer, prompt, max_tokens=3000))

The chat template is required. Without it the model emits <|endoftext|> immediately. temperature=0.6 / top_p=0.95 are the upstream recommendation. This is a reasoning model that writes long <think> passages, so budget 500-3000 tokens per answer.

Changes made to the upstream artifact

Weights are untouched. Two things were fixed so the model is usable:

  1. Stop tokens. Upstream disagrees with itself: config.json and generation_config.json say eos_token_id = 151645 (<|im_end|>), while tokenizer_config.json says <|endoftext|> (151643). mlx-lm takes the id from generation_config.json only, so generation never stopped and padded output with <|endoftext|> until the token limit. Both ids are now set.
  2. Quantization of the tied head. tie_word_embeddings is required by this architecture, so embed_tokens is also the output head and there is no lm_head module. mlx-lm's stock mixed recipes key on the string lm_head, so they silently push the output head to the lowest bit width. The -hq variants protect it explicitly; see the 3bit vs 3bit-hq gap.

How the port was validated

The architecture was reimplemented from the reference vLLM plugin and checked against an independent NumPy oracle with bit-level bfloat16 emulation:

  • logits vs oracle - argmax identical at every position; max deviation smaller than the architecture's own bfloat16 rounding noise
  • 1025-key sliding window - positions straddling the window boundary match the oracle exactly (the window is inclusive of the query token, an easy off-by-one)
  • prefill vs incremental decode - argmax agreement at 64 and 1100 tokens
  • bf16 conversion vs original checkpoint - bit-identical logits

This architecture is unusually sensitive to rounding: its residual stream reaches |h| ~ 1.5e5 by layer 47, and a single bfloat16 ULP perturbation at layer 1 shifts the final logits by 1.26. That is why quantization KL values here look higher than for a typical 1B model.

License and attribution

Apache-2.0, inherited from the original model. All credit for the model itself goes to Paradigma. The MLX port and quantization are mechanical work on top.

@misc{paradigma2026limite,
  title  = {{Limite 1B - Violetto}},
  author = {Prignano, Mario and Cirillo, Gabriele and Morosini, Alessio and
            Cerovaz, Luca and Bartolocci, Alessandro and Rodol\`a, Emanuele and
            Starace, Giulio and Pappone, Francesco},
  year   = {2026},
  howpublished = {\url{https://paradigma.inc/blog/limite-1b-violetto/}}
}
Downloads last month
47
Safetensors
Model size
1B params
Tensor type
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pierjoe/limite-1b-violetto-mlx-4bit-hq

Quantized
(8)
this model

Collection including pierjoe/limite-1b-violetto-mlx-4bit-hq