Kolibri-1 NVFP4 W4A16
An experimental, weight-only NVFP4 conversion of Aleph-Alpha/Kolibri-1, published by audreyt. Aleph Alpha trained the original model.
What changed
- Routed and shared experts use ModelOpt
W4A16_NVFP4, with group size 16. - Attention projections are dequantized from the source block-FP8 weights to BF16.
- Original BF16 embeddings, output head, routers, routing biases, and norms are preserved.
- Activations remain BF16. No activation-calibration dataset was used.
- Each expert's gate/up projections share a global scale, as required by vLLM's fused loader.
This is a conversion of already-quantized FP8 weights, not a quantization directly from the BF16 training checkpoint.
Conversion measurements
| Measurement | Result |
|---|---|
| Source revision | e52eb4627d11516b0c01de49210ab5a4e4061444 |
| Producer | NVIDIA ModelOpt 0.47.0 |
| Tensor payload | 47,396,140,120 bytes / 44.14 GiB |
| Output tensors | 173,853 |
| Safetensors shards | 45 |
| Quantized expert projections | 57,750 |
| BF16 attention projections | 200 |
| Aggregate relative weight RMSE | 0.0946875 |
| Maximum projection relative RMSE | 0.0955360 |
| Conversion time | 519.78 seconds |
| Peak conversion RSS | 6,147,420 KiB |
Conversion ran on an Intel Core Ultra 9 285K with eight CPU threads, no GPU, a 6 GiB container memory limit, and output buffers capped at 1 GiB per shard. The five numerical and machine-format tests passed. The converter also passed strict type checking and lint checks.
These are weight-reconstruction measurements, not language-model benchmark scores.
Inference status
Full generation and long-context validation are not yet complete.
The first RTX 5090 / WSL2 trial used vLLM 0.29.0, the official
aleph-alpha-inference 1.0.0 plugin, Marlin, and selective UVA offloading.
vLLM reported 22.06 GiB of expert parameters offloaded to RAM, then failed with
CUDA out-of-memory during post-load Marlin expert repacking.
This repository does not claim measured throughput, successful 1M-token inference, or quality equivalence to the original model. The base model's native trained context is 262,144 tokens; its advertised 1,048,576-token context is validated extrapolation. Those limits have not been verified for this conversion.
Reproduce the conversion
The recipe/ directory contains the CPU converter, numerical tests, type
interfaces, checker configuration, and a Dockerfile pinned to the vLLM 0.29.0
base image. The root config.json and hf_quant_config.json contain matching
ModelOpt metadata.
Download the recipe without downloading the converted weights:
hf download audreyt/Kolibri-1-NVFP4-W4A16 \
--include 'recipe/**' config.json hf_quant_config.json .dockerignore \
--local-dir kolibri-nvfp4-recipe
hf download Aleph-Alpha/Kolibri-1 \
--revision e52eb4627d11516b0c01de49210ab5a4e4061444 \
--local-dir kolibri-fp8-source
cd kolibri-nvfp4-recipe
docker build -f recipe/Dockerfile -t kolibri-nvfp4-converter .
mkdir -p ../kolibri-nvfp4-output
docker run --rm --memory 6g --memory-swap 8g --cpus 8 \
-e CUDA_VISIBLE_DEVICES= -e OMP_NUM_THREADS=8 \
-v "$(realpath ../kolibri-fp8-source):/source:ro" \
-v "$(realpath ../kolibri-nvfp4-output):/output" \
--entrypoint python3 kolibri-nvfp4-converter \
/work/convert_kolibri.py /source /output
The converter streams source tensors rather than loading a full model state dictionary. It checks projection reconstruction error, payload size, and tensor count before reporting completion. An output directory containing existing safetensors is rejected.
Dependencies and provenance
- Original model
- Official vLLM plugin
- NVIDIA ModelOpt
- vLLM 0.29.0; PyTorch 2.13.0+cu130; safetensors 0.8.0
- Base image:
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
The original model's Apache-2.0 license is preserved in LICENSE.
- Downloads last month
- 376