MiniGPT2-22M (GGUF, F16)

A 22M-parameter language model trained from scratch by SmallAICreator / UltraLabs, small enough to chat on a phone.

  • File: MiniGPT2-22M-F16.gguf (44.6 MB), for llama.cpp, PocketPal, LM Studio, Ollama, etc.
  • Chat template is embedded, so apps pick up the right format automatically.

Model

Parameters 22.03M (15.7M non-embedding)
Architecture Llama-style decoder: 10 layers, d=384, 6 heads / 2 KV heads (GQA), SwiGLU 1024, RoPE, RMSNorm, tied embeddings
Context 1024 tokens
Tokenizer 16,384-token byte-level BPE (trained for this model)
Training tokens ~26.5B (single pass, no repetition)
Data FineWeb-Edu, Cosmopedia-v2, SlimPajama, FineMath, Python code, HTML, plus a chat/instruction blend distilled from a larger model, mixed in during the final part of training
Optimizer Muon (hidden matrices) + AdamW (embeddings/norms)
Hardware TPU v5e-8, 7 hours

Prompt format

<|system|>You are a helpful assistant.<|end|><|user|>What is the capital of France?<|end|><|assistant|>

No BOS token. The model ends its turn with <|end|>.

llama-cli -m MiniGPT2-22M-F16.gguf -cnv

Tool calling

The embedded chat template supports tools (OpenAI-style tools / tool_calls), so apps with a tool picker enable it automatically. The model emits:

<tool_call>
{"name": "web_search", "arguments": {"query": "Tokyo weather today"}}
</tool_call>

Tool results go back as a tool message, which the template renders as a <tool_response> turn. The model then answers from the result. It was trained on calculator, web_search, fetch_url, search_knowledge_base and execute_javascript calls. Remember it's 22M, so double-check arguments.

Examples

Who are you? I'm MiniGPT2, a tiny 22M-parameter model from UltraLabs. I'm small, so I get facts wrong and I'm weak at multi-step arithmetic.

What is the capital of France? The capital of France is Paris.

Write a haiku about the ocean. Waves whisper softly / rolling onto quiet sand / the ocean breathes and rests.

Benchmarks

0-shot, lm-evaluation-harness 0.4.13, fp32. Every model was run through the same harness. Our Supra-50M run reproduces that model's published numbers exactly.

Task MiniGPT2 (22M) Pythia-70M Supra-50M
ARC-Easy (acc) 42.2 37.5 52.2
ARC-Easy (acc_norm) 40.0 35.2 46.0
ARC-Challenge (acc_norm) 23.1 21.9 25.0
HellaSwag (acc_norm) 28.2 27.4 31.8
PIQA (acc_norm) 58.1 59.2 62.1
SciQ (acc) 66.4 64.0 77.2
Winogrande (acc) 50.0 53.0 51.0
BLiMP (acc) 74.2 74.2 76.3
LAMBADA OpenAI (acc) 23.1 22.7 25.9
LAMBADA standard (acc) 13.5 15.5 17.2
WikiText (bits/byte, lower is better) 1.176 1.080 1.027

MiniGPT2 beats Pythia-70M (3.2x its size) on 6 of 11 tasks and ties BLiMP. It trails Supra-50M, a model 2.3x larger.

Conversion

Converted from the original JAX checkpoint to HF Llama, then to GGUF. Verified:

  • The HF model matches the training forward pass (max logit diff 4e-5).
  • llama.cpp tokenization matches the HF tokenizer exactly.
  • The F16 GGUF's logits match HF (top-1 agreement 100%, mean KL 7.6e-6).

Limitations

This is a 22M model. It makes factual mistakes, is weak at arithmetic and multi-step reasoning, and can be verbose. Do not rely on it for anything important.

Downloads last month
133
GGUF
Model size
22M params
Architecture
llama
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support