Instructions to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S # Run inference directly in the terminal: llama cli -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S # Run inference directly in the terminal: llama cli -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S # Run inference directly in the terminal: ./llama-cli -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Use Docker
docker model run hf.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
- LM Studio
- Jan
- vLLM
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Eliasfpv28/Kolibri-1-Q3_K_S-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Eliasfpv28/Kolibri-1-Q3_K_S-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
- Ollama
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with Ollama:
ollama run hf.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
- Unsloth Desktop
- Pi
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with Docker Model Runner:
docker model run hf.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
- Lemonade
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Run and chat with the model
lemonade run user.Kolibri-1-Q3_K_S-GGUF-Q3_K_S
List all available models
lemonade list
- Hermes Agent
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Eliasfpv28/Kolibri-1-Q3_K_S-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download runtime-source/README.md from Eliasfpv28/Kolibri-1-Q3_K_S-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 3.43 kB
-
https://huggingface.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF/resolve/main/runtime-source/README.md
- Command line
-
hf download hf://Eliasfpv28/Kolibri-1-Q3_K_S-GGUF/runtime-source/README.md
-
curl -L -o README.md https://huggingface.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF/resolve/main/runtime-source/README.md
Experimental Kolibri1 port for llama.cpp
This source patch is required for the GGUF in the parent directory. It is based on llama.cpp commit edd6e2bbdad5930899a93db8fa73c3b61c7b9bcc. Its architecture reference is the separately Apache-2.0 licensed Aleph Alpha inference code at 049a6a7bd2405b27d6d280d256bd3d585191c7ae.
See THIRD_PARTY_NOTICES.txt, LICENSE-MIT.txt, and LICENSE-Apache-2.0.txt. The model weights have their own Apache-2.0 license in the parent directory.
Build from source
Requirements: Git, CMake, a C++ compiler, and the Vulkan SDK including glslc for the tested backend. A Windows build was compiled with GCC 16.2.0 in w64devkit and Vulkan SDK 1.4.357.0. Other compilers/platforms have not been tested for this port.
Run these commands from this runtime-source directory:
git clone https://github.com/ggml-org/llama.cpp.git llama.cpp
git -C llama.cpp checkout edd6e2bbdad5930899a93db8fa73c3b61c7b9bcc
git -C llama.cpp apply --check ../kolibri1-runtime.patch
git -C llama.cpp apply ../kolibri1-runtime.patch
cmake -S llama.cpp -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON -DGGML_CUDA=OFF -DGGML_BLAS=OFF
cmake --build build --config Release --parallel 8 --target llama-server llama-cli llama-quantize
For Windows with w64devkit, use its compiler environment and add -G "MinGW Makefiles" to the CMake configure command. Keep the generated runtime DLLs beside the executable. Binary layout may depend on the selected CMake generator.
Serve the model
The following is the tested two-GPU configuration. Device IDs can differ on other machines: first run build/bin/llama-server --list-devices, then adapt --device and --tensor-split. On Windows, use the .exe suffix as needed.
build/bin/llama-server -m ../Kolibri-1-Q3_K_S.gguf --alias Kolibri-1-Q3_K_S --device Vulkan0,Vulkan2 --split-mode layer --tensor-split 1,2 -ngl 99 -c 4096 -b 256 -ub 64 -fa on --cache-type-k q8_0 --cache-type-v q8_0 --host 127.0.0.1 --port 8081 --parallel 1 --jinja --reasoning off --no-warmup
The browser UI is at http://127.0.0.1:8081 once the model is loaded. The command binds to localhost. The tested devices were RTX 3060 and Intel Arc Pro B60, with the integrated AMD GPU excluded. The full model needs about 31.54 GiB for its file, plus KV cache, working buffers, and device/host overhead. Smaller devices may require CPU offloading and enough host memory.
Only the 4,096-token configuration and non-reasoning answers were tested on the full model. No long-context, tool-calling, or reasoning-mode compatibility claim is made.
Reproduce the weight conversion
stream-convert.py expects the complete pinned BF16 model snapshot and the patched llama.cpp directory beside the script. It needs Python and NumPy. Conversion produces an intermediate file of approximately 156 GB, so allow sufficient storage in addition to the original snapshot and final GGUF.
python stream-convert.py PATH_TO_KOLIBRI_BF16 Kolibri-1-BF16.gguf
build/bin/llama-quantize --tensor-type ffn_gate_inp=f32 --max-buffer-size 1024 Kolibri-1-BF16.gguf Kolibri-1-Q3_K_S.gguf Q3_K_S 6
The converter rejects existing outputs and validates the complete tensor mapping. The public GGUF also has publication notices in its header, so a fresh conversion's whole-file hash will differ even when the tensor data matches. See ../MODIFICATIONS.md and ../provenance.json.