How to use from
Docker Model Runner
docker model run hf.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF:Q3_K_S
Quick Links

Kolibri 1 — Q3_K_S GGUF

Unofficial, experimental GGUF conversion and quantization of Aleph Alpha Kolibri-1-BF16.

Requires the included Kolibri1 llama.cpp source patch. The unmodified llama.cpp revision used as the base for this port does not support this architecture. Compatibility with other releases, Ollama, or LM Studio has not been verified. This is an independent conversion; Aleph Alpha and the llama.cpp project have not endorsed it.

Deutsch: Dies ist eine unabhängige 3-Bit-GGUF-Version von Kolibri 1. Sie benötigt die hier dokumentierte experimentelle llama.cpp-Erweiterung. Das vollständige Modell wurde lokal auf einer RTX 3060 und einer Intel Arc Pro B60 gestartet; eine umfassende Bewertung der Antwortqualität steht aus.

File and quantization

Property Value
File Kolibri-1-Q3_K_S.gguf
Size Approximately 33.87 GB / 31.54 GiB; exact size in provenance.json
Format GGUF v3, architecture kolibri1
Scheme Q3_K_S; mixed precision, approximately 3.47 bits per parameter overall
Tensor types 501 Q3_K, 401 F32, 1 Q6_K
Router projections, norms, biases F32
Output matrix Q6_K
Importance matrix None
Further training None

“3-bit” describes the quantization scheme. Some tensors deliberately retain higher precision. All experts are included; active parameters per token do not represent the memory needed to store the model.

The original BF16 weights were converted with the included streaming converter and quantized with patched llama.cpp. On 2026-10-03, the publication copy received embedded license, source, and modification notices. Its entire quantized tensor payload is unchanged from the locally tested GGUF. See MODIFICATIONS.md, provenance.json, and SHA256SUMS.

Context

The original model has a native context of 262,144 tokens. Aleph Alpha reports extended-context validation up to 1,048,576 tokens and recommends at most 262,144 for efficiency and complex tasks. Upstream model card

This GGUF port has been tested locally only at 4,096 tokens. Longer contexts, reasoning mode, and tool calling have not been validated in this port. The local 4,096-token setting is a serving configuration, not an inherent limit of weight quantization. Larger contexts need additional KV-cache memory and runtime validation.

Running

Build the pinned llama.cpp revision with the supplied patch using runtime-source/README.md. Use that resulting llama-server or llama-cli binary to load this file. Hardware selection and memory requirements depend on the machine; the model file alone needs approximately 31.54 GiB before runtime buffers and KV cache.

Tested server settings: 4,096-token context, one request slot, Q8_0 key/value cache, flash attention, Vulkan layer split 1:2 across NVIDIA RTX 3060 12 GiB and Intel Arc Pro B60 24 GiB. The AMD integrated GPU was excluded. See the runtime source instructions for an example command.

Validation and limits

Prior checks covered all 903 tensor names, shapes, offsets, types, and finite F32 router/norm values. The port was compared numerically against an independent full-precision reference on a small fixture, and tokenizer comparisons passed 169 cases. The full quantized model answered “Paris” to a capital question and “17 mal 23 ergibt 391.” to a multiplication question. Details are in validation.json.

These checks establish a limited functional result. They are not a language-quality benchmark, a safety evaluation, or evidence that upstream benchmark results transfer to this quantization. Quantization can reduce accuracy. No claim of unchanged capabilities or validated long-context behavior is made.

License and attribution

Original model provider: Aleph Alpha GmbH. Original model developer: Aleph Alpha Research GmbH. The original weights and published configurations are Apache-2.0 licensed. This quantized distribution includes the unchanged LICENSE, attribution in NOTICE, and modification details.

The model repository's license grant covers its published weights and configuration files. Other artifacts have their own rights and licenses. The included runtime source patch uses separately licensed Apache-2.0 inference reference code and MIT-licensed llama.cpp code; see runtime-source/THIRD_PARTY_NOTICES.txt.

The original model's intended uses, limitations, and responsible-use information remain relevant; consult its model card. Users remain responsible for complying with applicable law. Distribution is subject to the included licenses and their warranty disclaimers.

Sources

Downloads last month
1,481
GGUF
Model size
78B params
Architecture
kolibri1
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Eliasfpv28/Kolibri-1-Q3_K_S-GGUF

Quantized
(19)
this model

Space using Eliasfpv28/Kolibri-1-Q3_K_S-GGUF 1