Eliasfpv28's picture
Publish licenses and runtime instructions while the GGUF upload proceeds
eb053f2 verified
|
Raw History Blame Contribute Delete
3.43 kB

Experimental Kolibri1 port for llama.cpp

This source patch is required for the GGUF in the parent directory. It is based on llama.cpp commit edd6e2bbdad5930899a93db8fa73c3b61c7b9bcc. Its architecture reference is the separately Apache-2.0 licensed Aleph Alpha inference code at 049a6a7bd2405b27d6d280d256bd3d585191c7ae.

See THIRD_PARTY_NOTICES.txt, LICENSE-MIT.txt, and LICENSE-Apache-2.0.txt. The model weights have their own Apache-2.0 license in the parent directory.

Build from source

Requirements: Git, CMake, a C++ compiler, and the Vulkan SDK including glslc for the tested backend. A Windows build was compiled with GCC 16.2.0 in w64devkit and Vulkan SDK 1.4.357.0. Other compilers/platforms have not been tested for this port.

Run these commands from this runtime-source directory:

git clone https://github.com/ggml-org/llama.cpp.git llama.cpp
git -C llama.cpp checkout edd6e2bbdad5930899a93db8fa73c3b61c7b9bcc
git -C llama.cpp apply --check ../kolibri1-runtime.patch
git -C llama.cpp apply ../kolibri1-runtime.patch
cmake -S llama.cpp -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON -DGGML_CUDA=OFF -DGGML_BLAS=OFF
cmake --build build --config Release --parallel 8 --target llama-server llama-cli llama-quantize

For Windows with w64devkit, use its compiler environment and add -G "MinGW Makefiles" to the CMake configure command. Keep the generated runtime DLLs beside the executable. Binary layout may depend on the selected CMake generator.

Serve the model

The following is the tested two-GPU configuration. Device IDs can differ on other machines: first run build/bin/llama-server --list-devices, then adapt --device and --tensor-split. On Windows, use the .exe suffix as needed.

build/bin/llama-server -m ../Kolibri-1-Q3_K_S.gguf --alias Kolibri-1-Q3_K_S --device Vulkan0,Vulkan2 --split-mode layer --tensor-split 1,2 -ngl 99 -c 4096 -b 256 -ub 64 -fa on --cache-type-k q8_0 --cache-type-v q8_0 --host 127.0.0.1 --port 8081 --parallel 1 --jinja --reasoning off --no-warmup

The browser UI is at http://127.0.0.1:8081 once the model is loaded. The command binds to localhost. The tested devices were RTX 3060 and Intel Arc Pro B60, with the integrated AMD GPU excluded. The full model needs about 31.54 GiB for its file, plus KV cache, working buffers, and device/host overhead. Smaller devices may require CPU offloading and enough host memory.

Only the 4,096-token configuration and non-reasoning answers were tested on the full model. No long-context, tool-calling, or reasoning-mode compatibility claim is made.

Reproduce the weight conversion

stream-convert.py expects the complete pinned BF16 model snapshot and the patched llama.cpp directory beside the script. It needs Python and NumPy. Conversion produces an intermediate file of approximately 156 GB, so allow sufficient storage in addition to the original snapshot and final GGUF.

python stream-convert.py PATH_TO_KOLIBRI_BF16 Kolibri-1-BF16.gguf
build/bin/llama-quantize --tensor-type ffn_gate_inp=f32 --max-buffer-size 1024 Kolibri-1-BF16.gguf Kolibri-1-Q3_K_S.gguf Q3_K_S 6

The converter rejects existing outputs and validates the complete tensor mapping. The public GGUF also has publication notices in its header, so a fresh conversion's whole-file hash will differ even when the tensor data matches. See ../MODIFICATIONS.md and ../provenance.json.