Kolibri-1 NVFP4 W4A16

An experimental, weight-only NVFP4 conversion of Aleph-Alpha/Kolibri-1, published by audreyt. Aleph Alpha trained the original model.

What changed

  • Routed and shared experts use ModelOpt W4A16_NVFP4, with group size 16.
  • Attention projections are dequantized from the source block-FP8 weights to BF16.
  • Original BF16 embeddings, output head, routers, routing biases, and norms are preserved.
  • Activations remain BF16. No activation-calibration dataset was used.
  • Each expert's gate/up projections share a global scale, as required by vLLM's fused loader.

This is a conversion of already-quantized FP8 weights, not a quantization directly from the BF16 training checkpoint.

Conversion measurements

Measurement Result
Source revision e52eb4627d11516b0c01de49210ab5a4e4061444
Producer NVIDIA ModelOpt 0.47.0
Tensor payload 47,396,140,120 bytes / 44.14 GiB
Output tensors 173,853
Safetensors shards 45
Quantized expert projections 57,750
BF16 attention projections 200
Aggregate relative weight RMSE 0.0946875
Maximum projection relative RMSE 0.0955360
Conversion time 519.78 seconds
Peak conversion RSS 6,147,420 KiB

Conversion ran on an Intel Core Ultra 9 285K with eight CPU threads, no GPU, a 6 GiB container memory limit, and output buffers capped at 1 GiB per shard. The five numerical and machine-format tests passed. The converter also passed strict type checking and lint checks.

These are weight-reconstruction measurements, not language-model benchmark scores.

Inference status

Full generation and long-context validation are not yet complete.

The first RTX 5090 / WSL2 trial used vLLM 0.29.0, the official aleph-alpha-inference 1.0.0 plugin, Marlin, and selective UVA offloading. vLLM reported 22.06 GiB of expert parameters offloaded to RAM, then failed with CUDA out-of-memory during post-load Marlin expert repacking.

This repository does not claim measured throughput, successful 1M-token inference, or quality equivalence to the original model. The base model's native trained context is 262,144 tokens; its advertised 1,048,576-token context is validated extrapolation. Those limits have not been verified for this conversion.

Reproduce the conversion

The recipe/ directory contains the CPU converter, numerical tests, type interfaces, checker configuration, and a Dockerfile pinned to the vLLM 0.29.0 base image. The root config.json and hf_quant_config.json contain matching ModelOpt metadata.

Download the recipe without downloading the converted weights:

hf download audreyt/Kolibri-1-NVFP4-W4A16 \
  --include 'recipe/**' config.json hf_quant_config.json .dockerignore \
  --local-dir kolibri-nvfp4-recipe

hf download Aleph-Alpha/Kolibri-1 \
  --revision e52eb4627d11516b0c01de49210ab5a4e4061444 \
  --local-dir kolibri-fp8-source

cd kolibri-nvfp4-recipe
docker build -f recipe/Dockerfile -t kolibri-nvfp4-converter .
mkdir -p ../kolibri-nvfp4-output

docker run --rm --memory 6g --memory-swap 8g --cpus 8 \
  -e CUDA_VISIBLE_DEVICES= -e OMP_NUM_THREADS=8 \
  -v "$(realpath ../kolibri-fp8-source):/source:ro" \
  -v "$(realpath ../kolibri-nvfp4-output):/output" \
  --entrypoint python3 kolibri-nvfp4-converter \
  /work/convert_kolibri.py /source /output

The converter streams source tensors rather than loading a full model state dictionary. It checks projection reconstruction error, payload size, and tensor count before reporting completion. An output directory containing existing safetensors is rejected.

Dependencies and provenance

The original model's Apache-2.0 license is preserved in LICENSE.

Downloads last month
376
Safetensors
Model size
40B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for audreyt/Kolibri-1-NVFP4-W4A16

Quantized
(21)
this model