Instructions to use sirunchained/VibeThinker-1.5B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use sirunchained/VibeThinker-1.5B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Use Docker
docker model run hf.co/sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use sirunchained/VibeThinker-1.5B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sirunchained/VibeThinker-1.5B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sirunchained/VibeThinker-1.5B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
- Ollama
How to use sirunchained/VibeThinker-1.5B-gguf with Ollama:
ollama run hf.co/sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
- Unsloth Desktop
- Pi
How to use sirunchained/VibeThinker-1.5B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sirunchained/VibeThinker-1.5B-gguf:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use sirunchained/VibeThinker-1.5B-gguf with Docker Model Runner:
docker model run hf.co/sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
- Lemonade
How to use sirunchained/VibeThinker-1.5B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Run and chat with the model
lemonade run user.VibeThinker-1.5B-gguf-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use sirunchained/VibeThinker-1.5B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sirunchained/VibeThinker-1.5B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sirunchained/VibeThinker-1.5B-gguf:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sirunchained/VibeThinker-1.5B-gguf:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
VibeThinker-1.5B GGUF
This is a quantized version of the WeiboAI/VibeThinker-1.5B model, converted to GGUF format for use with llama.cpp and compatible tools.
Quantization Details
The model was quantized using llama.cpp to the following formats:
- FP16: Converted from the original checkpoint as an intermediate step.
- Q4_K_M: Quantized using the Q4_K_M method.
- Q5_K_M: Quantized using the Q5_K_M method.
- Q6_K: Quantized using the Q6_K method.
- Q8_0: Quantized using the Q8_0 method.
Original Model Card Summary
- Model ID:
WeiboAI/VibeThinker-1.5B - Original Repository: https://huggingface.co/WeiboAI/VibeThinker-1.5B
Files Provided
VibeThinker-1.5B-f16.gguf(FP16)VibeThinker-1.5B-q4_k_m.gguf(Q4_K_M)VibeThinker-1.5B-q5_k_m.gguf(Q5_K_M)VibeThinker-1.5B-q6_k.gguf(Q6_K)VibeThinker-1.5B-q8_0.gguf(Q8_0)
Q4_K_M Benchmark Results
The following results were obtained with the Q4_K_M quantization. Evaluation settings: temperature=0.6, top_p=0.95, max_gen_toks=40960, seed=1234. These settings match the official recommendation from the WeiboAI model card.
| Benchmark | Metric | Q4_K_M Result | Sample Size | Original Result |
|---|---|---|---|---|
| custom_aime24 | Accuracy | 80.0% (±7.43%) | 30 | 80.3 |
| custom_hmmt25 | Accuracy | 46.67% (±9.26%) | 30 | 50.4 |
| custom_lcb_v6 | Pass@1 | 41.98% (±4.33%) | 131 | 51.1 |
| custom_math500 | Accuracy | 96.80% (±0.79%) | 500 | 95.0 |
Footnotes:
- The "Original Result" column refers to scores reported in the VibeThinker-1.5B technical report (Tiny Model, Big Logic, arXiv:2511.06221). MATH-500 (95.0) is from Table 1; the report evaluates it using average pass rate across 64 sampling trials.
- The
custom_*task names refer to the official evaluation datasets loaded through a rewritten YAML configuration with a modified scorer. The datasets themselves are the standard benchmarks. - The original model's benchmark scores were obtained with the same generation parameters recommended for reproduction:
temperature=0.6,max_token_length=40960,top_p=0.95.
Key observations:
- AIME24: Q4_K_M (80.0%) vs. Original (80.3%) - essentially identical. 4-bit quantization preserves mathematical reasoning on this benchmark with negligible degradation.
- HMMT25: Q4_K_M (46.67%) vs. Original (50.4%) - a ~3.7-point drop, consistent with expected precision loss on a harder competition math benchmark.
- LiveCodeBench v6: Q4_K_M (41.98%) vs. Original (51.1%) - code generation is more sensitive to quantization, and the ~9-point gap reflects that.
- MATH-500: Q4_K_M (96.8%) actually exceeds the Original score (95.0). This benchmark has sufficient headroom that the quantized model retains full capability.
Original Model Performance
From the technical report:
| Benchmark | Score |
|---|---|
| AIME24 | 80.3 |
| AIME25 | 74.4 |
| MATH-500 | 95.0 |
| HMMT25 | 50.4 |
| LiveCodeBench v5 | 55.9 |
| LiveCodeBench v6 | 51.1 |
These scores place VibeThinker-1.5B on par with or above models 100×–600× larger, including DeepSeek R1 (671B) and Kimi K2 (1000B+).
Model tree for sirunchained/VibeThinker-1.5B-gguf
- Base model: Qwen/Qwen2.5-1.5B
- Quantized from: WeiboAI/VibeThinker-1.5B
Notes on the Results
The evaluation JSON files are labeled custom_aime24, custom_hmmt25, custom_lcb_v6, and custom_math500. These use the official benchmark datasets, run through a rewritten YAML configuration with a modified scorer (the default scorer would otherwise fail on these tasks). The generation parameters match the recommended inference settings from the original model card: temperature=0.6, max_token_length=40960, top_p=0.95.
Standard errors are included in the results table. The relatively large standard errors for the 30-sample benchmarks (AIME24, HMMT25) reflect the small sample size; MATH-500 is much more statistically stable with 500 problems.
- Downloads last month
- 205
4-bit
5-bit
6-bit
8-bit
16-bit