Instructions to use SmallAICreator/MiniGPT2-22M-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SmallAICreator/MiniGPT2-22M-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SmallAICreator/MiniGPT2-22M-GGUF:F16 # Run inference directly in the terminal: llama cli -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SmallAICreator/MiniGPT2-22M-GGUF:F16 # Run inference directly in the terminal: llama cli -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SmallAICreator/MiniGPT2-22M-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SmallAICreator/MiniGPT2-22M-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Use Docker
docker model run hf.co/SmallAICreator/MiniGPT2-22M-GGUF:F16
- LM Studio
- Jan
- vLLM
How to use SmallAICreator/MiniGPT2-22M-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SmallAICreator/MiniGPT2-22M-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SmallAICreator/MiniGPT2-22M-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SmallAICreator/MiniGPT2-22M-GGUF:F16
- Ollama
How to use SmallAICreator/MiniGPT2-22M-GGUF with Ollama:
ollama run hf.co/SmallAICreator/MiniGPT2-22M-GGUF:F16
- Unsloth Desktop
- Pi
How to use SmallAICreator/MiniGPT2-22M-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SmallAICreator/MiniGPT2-22M-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use SmallAICreator/MiniGPT2-22M-GGUF with Docker Model Runner:
docker model run hf.co/SmallAICreator/MiniGPT2-22M-GGUF:F16
- Lemonade
How to use SmallAICreator/MiniGPT2-22M-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SmallAICreator/MiniGPT2-22M-GGUF:F16
Run and chat with the model
lemonade run user.MiniGPT2-22M-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use SmallAICreator/MiniGPT2-22M-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SmallAICreator/MiniGPT2-22M-GGUF:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use SmallAICreator/MiniGPT2-22M-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SmallAICreator/MiniGPT2-22M-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SmallAICreator/MiniGPT2-22M-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
MiniGPT2-22M (GGUF, F16)
A 22M-parameter language model trained from scratch by SmallAICreator / UltraLabs, small enough to chat on a phone.
- File:
MiniGPT2-22M-F16.gguf(44.6 MB), for llama.cpp, PocketPal, LM Studio, Ollama, etc. - Chat template is embedded, so apps pick up the right format automatically.
Model
| Parameters | 22.03M (15.7M non-embedding) |
| Architecture | Llama-style decoder: 10 layers, d=384, 6 heads / 2 KV heads (GQA), SwiGLU 1024, RoPE, RMSNorm, tied embeddings |
| Context | 1024 tokens |
| Tokenizer | 16,384-token byte-level BPE (trained for this model) |
| Training tokens | ~26.5B (single pass, no repetition) |
| Data | FineWeb-Edu, Cosmopedia-v2, SlimPajama, FineMath, Python code, HTML, plus a chat/instruction blend distilled from a larger model, mixed in during the final part of training |
| Optimizer | Muon (hidden matrices) + AdamW (embeddings/norms) |
| Hardware | TPU v5e-8, 7 hours |
Prompt format
<|system|>You are a helpful assistant.<|end|><|user|>What is the capital of France?<|end|><|assistant|>
No BOS token. The model ends its turn with <|end|>.
llama-cli -m MiniGPT2-22M-F16.gguf -cnv
Tool calling
The embedded chat template supports tools (OpenAI-style tools / tool_calls), so apps with a tool picker enable it automatically. The model emits:
<tool_call>
{"name": "web_search", "arguments": {"query": "Tokyo weather today"}}
</tool_call>
Tool results go back as a tool message, which the template renders as a <tool_response> turn. The model then answers from the result. It was trained on calculator, web_search, fetch_url, search_knowledge_base and execute_javascript calls. Remember it's 22M, so double-check arguments.
Examples
Who are you? I'm MiniGPT2, a tiny 22M-parameter model from UltraLabs. I'm small, so I get facts wrong and I'm weak at multi-step arithmetic.
What is the capital of France? The capital of France is Paris.
Write a haiku about the ocean. Waves whisper softly / rolling onto quiet sand / the ocean breathes and rests.
Benchmarks
0-shot, lm-evaluation-harness 0.4.13, fp32. Every model was run through the same harness. Our Supra-50M run reproduces that model's published numbers exactly.
| Task | MiniGPT2 (22M) | Pythia-70M | Supra-50M |
|---|---|---|---|
| ARC-Easy (acc) | 42.2 | 37.5 | 52.2 |
| ARC-Easy (acc_norm) | 40.0 | 35.2 | 46.0 |
| ARC-Challenge (acc_norm) | 23.1 | 21.9 | 25.0 |
| HellaSwag (acc_norm) | 28.2 | 27.4 | 31.8 |
| PIQA (acc_norm) | 58.1 | 59.2 | 62.1 |
| SciQ (acc) | 66.4 | 64.0 | 77.2 |
| Winogrande (acc) | 50.0 | 53.0 | 51.0 |
| BLiMP (acc) | 74.2 | 74.2 | 76.3 |
| LAMBADA OpenAI (acc) | 23.1 | 22.7 | 25.9 |
| LAMBADA standard (acc) | 13.5 | 15.5 | 17.2 |
| WikiText (bits/byte, lower is better) | 1.176 | 1.080 | 1.027 |
MiniGPT2 beats Pythia-70M (3.2x its size) on 6 of 11 tasks and ties BLiMP. It trails Supra-50M, a model 2.3x larger.
Conversion
Converted from the original JAX checkpoint to HF Llama, then to GGUF. Verified:
- The HF model matches the training forward pass (max logit diff 4e-5).
- llama.cpp tokenization matches the HF tokenizer exactly.
- The F16 GGUF's logits match HF (top-1 agreement 100%, mean KL 7.6e-6).
Limitations
This is a 22M model. It makes factual mistakes, is weak at arithmetic and multi-step reasoning, and can be verbose. Do not rely on it for anything important.
- Downloads last month
- 133
16-bit