How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
# Run inference directly in the terminal:
llama cli -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
# Run inference directly in the terminal:
llama cli -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
# Run inference directly in the terminal:
./llama-cli -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
Use Docker
docker model run hf.co/perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF:
Quick Links

RCRC Hanifa GGUF (Gemma 3 1B)

Quantized GGUF of perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa for llama.cpp and Ollama. Original FP16 weights remain at the source repo.

This is the chatbot for RCRC's Hanifa Standard knowledge base, fine-tuned with RAG-style instruction data across MSA, Najdi, Hijazi, and English.

Files

File Quant Size (approx) Notes
rcrc-hanifa-gemma3-1b-F16.gguf F16 ~2.0 GB Reference precision
rcrc-hanifa-gemma3-1b-Q8_0.gguf Q8_0 ~1.0 GB Near-lossless
rcrc-hanifa-gemma3-1b-Q5_K_M.gguf Q5_K_M ~720 MB Balanced
rcrc-hanifa-gemma3-1b-Q4_K_M.gguf Q4_K_M ~620 MB Recommended for laptops

Quick start β€” llama.cpp

hf download perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF rcrc-hanifa-gemma3-1b-Q4_K_M.gguf --local-dir .
./llama-cli -m rcrc-hanifa-gemma3-1b-Q4_K_M.gguf -cnv

Quick start β€” Ollama

hf download perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF rcrc-hanifa-gemma3-1b-Q4_K_M.gguf Modelfile --local-dir ./rcrc-hanifa
cd rcrc-hanifa
ollama create rcrc-hanifa -f Modelfile
ollama run rcrc-hanifa

Chat template (Gemma 3)

<start_of_turn>user
{user_message}<end_of_turn>
<start_of_turn>model
{response}<end_of_turn>
Downloads last month
63
GGUF
Model size
1.0B params
Architecture
gemma3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for perfectPresentation/rcrc-chat-v3-gemma-1b-rag-hanifa-GGUF

Quantized
(1)
this model