Instructions to use pierjoe/limite-1b-violetto-mlx-4bit-hq with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use pierjoe/limite-1b-violetto-mlx-4bit-hq with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("pierjoe/limite-1b-violetto-mlx-4bit-hq") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use pierjoe/limite-1b-violetto-mlx-4bit-hq with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "pierjoe/limite-1b-violetto-mlx-4bit-hq"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "pierjoe/limite-1b-violetto-mlx-4bit-hq" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use pierjoe/limite-1b-violetto-mlx-4bit-hq with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "pierjoe/limite-1b-violetto-mlx-4bit-hq"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "pierjoe/limite-1b-violetto-mlx-4bit-hq" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pierjoe/limite-1b-violetto-mlx-4bit-hq", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use pierjoe/limite-1b-violetto-mlx-4bit-hq with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "pierjoe/limite-1b-violetto-mlx-4bit-hq"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default pierjoe/limite-1b-violetto-mlx-4bit-hq
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use pierjoe/limite-1b-violetto-mlx-4bit-hq with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "pierjoe/limite-1b-violetto-mlx-4bit-hq"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "pierjoe/limite-1b-violetto-mlx-4bit-hq" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
limite-1b-violetto-mlx-4bit-hq
paradigma-inc/limite-1b-violetto converted for MLX
on Apple Silicon. 4bit-hq - 4-bit body with the tied output head and value embeddings at 6-bit. Same speed as plain 4bit, 3.5x lower KL, for 55 MB more.
| size on disk | 0.61 GiB |
| effective bits/weight | 5.04 |
| KL divergence from bf16 | 0.07134 nats |
| top-1 agreement with bf16 | 88.8% |
| generation speed | 72 tok/s |
| peak memory | 0.66 GiB |
Speed and memory are measured on an M2 (8 GB) with mlx-lm 0.31.3. Fidelity is
measured on the deployed distribution rather than as raw perplexity, because this
model's head squashes logits into [0, 23] via 23 * sigmoid((raw + 5) / 7.5)
before sampling ever sees them.
Setup
limite is not one of mlx-lm's built-in architectures, so stock mlx-lm needs one
extra package to know how to build this model. It registers itself through a
sys.meta_path hook, so mlx-lm's own files are never modified and an mlx-lm
upgrade cannot break it:
pip install mlx-lm "limite-mlx @ https://huggingface.co/pierjoe/limite-mlx/resolve/main/limite_mlx-0.1.0-py3-none-any.whl"
Nothing to import and nothing to configure after that. The same file is also in
this repo as limite.py if you would rather drop it into
mlx_lm/models/ yourself.
Usage
mlx_lm.generate --model pierjoe/limite-1b-violetto-mlx-4bit-hq \
--prompt "What is the remainder when 7^2026 is divided by 13?" \
--max-tokens 3000 --temp 0.6 --top-p 0.95
from mlx_lm import generate, load
model, tokenizer = load("pierjoe/limite-1b-violetto-mlx-4bit-hq")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "How many positive divisors does 360 have?"}],
add_generation_prompt=True,
tokenize=False,
)
print(generate(model, tokenizer, prompt, max_tokens=3000))
The chat template is required. Without it the model emits <|endoftext|>
immediately. temperature=0.6 / top_p=0.95 are the upstream recommendation.
This is a reasoning model that writes long <think> passages, so budget
500-3000 tokens per answer.
Changes made to the upstream artifact
Weights are untouched. Two things were fixed so the model is usable:
- Stop tokens. Upstream disagrees with itself:
config.jsonandgeneration_config.jsonsayeos_token_id = 151645(<|im_end|>), whiletokenizer_config.jsonsays<|endoftext|>(151643). mlx-lm takes the id fromgeneration_config.jsononly, so generation never stopped and padded output with<|endoftext|>until the token limit. Both ids are now set. - Quantization of the tied head.
tie_word_embeddingsis required by this architecture, soembed_tokensis also the output head and there is nolm_headmodule. mlx-lm's stock mixed recipes key on the stringlm_head, so they silently push the output head to the lowest bit width. The-hqvariants protect it explicitly; see the3bitvs3bit-hqgap.
How the port was validated
The architecture was reimplemented from the reference vLLM plugin and checked against an independent NumPy oracle with bit-level bfloat16 emulation:
- logits vs oracle - argmax identical at every position; max deviation smaller than the architecture's own bfloat16 rounding noise
- 1025-key sliding window - positions straddling the window boundary match the oracle exactly (the window is inclusive of the query token, an easy off-by-one)
- prefill vs incremental decode - argmax agreement at 64 and 1100 tokens
- bf16 conversion vs original checkpoint - bit-identical logits
This architecture is unusually sensitive to rounding: its residual stream
reaches |h| ~ 1.5e5 by layer 47, and a single bfloat16 ULP perturbation at
layer 1 shifts the final logits by 1.26. That is why quantization KL values here
look higher than for a typical 1B model.
License and attribution
Apache-2.0, inherited from the original model. All credit for the model itself goes to Paradigma. The MLX port and quantization are mechanical work on top.
@misc{paradigma2026limite,
title = {{Limite 1B - Violetto}},
author = {Prignano, Mario and Cirillo, Gabriele and Morosini, Alessio and
Cerovaz, Luca and Bartolocci, Alessandro and Rodol\`a, Emanuele and
Starace, Giulio and Pappone, Francesco},
year = {2026},
howpublished = {\url{https://paradigma.inc/blog/limite-1b-violetto/}}
}
- Downloads last month
- 47
4-bit
Model tree for pierjoe/limite-1b-violetto-mlx-4bit-hq
Base model
paradigma-inc/limite-1b-violetto