Instructions to use 0xSero/DeepSeek-V4.1-Flash-Spark with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Trellis
How to use 0xSero/DeepSeek-V4.1-Flash-Spark with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
DeepSeek-V4.1-Flash-Spark
DeepSeek-V4.1-Flash with every backbone routed expert re-encoded as EXL3 trellis (MUL1 codebook) at a per-expert K2 / K3 / K5 mix, sized to serve on two NVIDIA DGX Spark (GB10, 2 x 128 GB unified memory, tensor parallel 2). Average routed-expert rate 2.77 bpw. Everything else — attention, shared experts, routers, Engram tables, the DSpark/MTP draft layers and their experts, the vision tower, embeddings and head — is the upstream checkpoint, byte for byte, in its native FP8 / BF16 / MXFP4 formats.
EXL3 experts need the EXL3 MoE path in the companion vLLM image (below). Stock vLLM and plain
transformerscannot run this checkpoint. Two Engram shards are stored as split parts (HF limits files to 50 GB); run./reassemble.shonce after download.
At a glance
| Base | deepseek-ai/DeepSeek-V4.1-Flash @ fb2764a5cf321eaa5070ca8f9e892818f477c16d |
| Architecture | 40 layers, 384 routed experts (top-6) + 1 shared, hidden 5120; 3 DSpark/MTP draft layers; Engram on layers 1 and 14; 32-layer vision tower |
| Routed experts (layers 0-39) | EXL3 trellis, MUL1 codebook: K2 7,279 · K3 6,299 · K5 1,782 of 15,360 experts; 2.77 bpw average |
| Everything else | upstream tensors, byte-identical (FP8 e4m3 / ue8m0 scales, BF16, MXFP4 for the draft experts) |
| Engram tables | native FP8, 2 x 98.3 GB, read from disk at serve time |
| Size | native part 221.5 GB + EXL3 banks 240.4 GB = ~462 GB |
| Context | 1,048,576 native; served at 262,144 |
| Vision / tools / reasoning | yes / yes / yes (smoke-tested on the served stack) |
| License | MIT (DeepSeek), carried from the base |
Quality
Offline full-vocabulary token-wise KL divergence against the native checkpoint (teacher), 64 windows per panel, bootstrap 95 % CI. Lower KLD and higher top-1 agreement are better.
| Panel | Mean KLD (nats) [95 % CI] | Top-1 agreement |
|---|---|---|
| v3.1 panel | 0.0742 [0.0644, 0.0865] | 0.917 |
| legacy panel | 0.0449 | 0.926 |
Paired against the previous allocation at 2.68 bpw (same K2/K3 banks, 1,782 experts left as native MXFP4), this build moves −1.5 % on v3.1 and +0.8 % on legacy KLD, both inside the noise band; it was accepted because every routed expert now uses one kernel family.
Serving on 2 x DGX Spark
Historical measurements on host-built image sha256:5668e35e5ee021da4b9ce88a1964caf0e082e8ca01d81d1d64b6213a55c6add5
and the full upstream checkpoint layout, TP2 over the Sparks' RoCE link, FP8 KV cache, DSpark speculative
decoding (7 draft tokens), Engram tables on disk, CUDA graphs FULL_DECODE_ONLY, 2 concurrent sequences:
| KV cache | 2,023,717 tokens |
| Max context | 262,144 |
| Prefill | 2,050 tok/s at 8k · 2,061 tok/s at 32k |
| Code decode, 1 stream | ~39-41 tok/s |
| Code decode, 2 streams | 60 tok/s aggregate |
| Prose decode, 1 stream | ~29 tok/s |
| Load time | ~11-13 min (EXL3 banks ~15 s per layer) |
All decode runs used the model's default sampling and stopped naturally (3k-15k tokens); no output caps. Smoke checks passed for text, tool calls and vision.
Published image and validation status
The public CI image is
ghcr.io/0xsero/deepseek-v4.1-flash-spark@sha256:3cbc8ec016f5fbfc82eba3480de12399cbce31e7b76aa3108b8fe8246c42a79d.
GitHub-hosted ARM64 build 37138684406
published it from source 93e001cf4ecafeb71c6e737eeefad9a9b4ce9235 on main, with verified signed provenance.
Fresh model acceptance on this exact image and this repository's compact layout is still pending in
registry PR #154. The historical measurements above
have not been transferred to this image or layout.
Run it
- Download and rebuild the two Engram shards:
hf download 0xSero/DeepSeek-V4.1-Flash-Spark --revision 08ac8b3defc9a239ba0baf51059687c689406519 --local-dir ~/models/DeepSeek-V4.1-Flash-Spark
cd ~/models/DeepSeek-V4.1-Flash-Spark && ./reassemble.sh # cat parts, sha256 check
Allow about 665 GB per node during download and reassembly, plus image and cache space.
The two-Spark launcher mounts the repository root at
/modeland writes a separate plan copy whose bank paths are absolute/model/exl3/...paths. It mounts that copy at/plansand keeps the checkpoint read-only. The pending registry recipe instead uses the original relative plan with workdir/model/exl3.To exercise the CI image, set
IMGto the exact public digest above when runningscripts/launch.shfrom the launcher repository; also setWORKERto your second Spark's SSH target as its instructions describe. The launcher's default still points to the historical host image. This is a candidate configuration pending the acceptance run linked above. Key flags:
vllm serve /model --tensor-parallel-size 2 --nnodes 2 --dtype bfloat16 \
--kv-cache-dtype fp8 --block-size 256 --swa-block-size 128 --kv-cache-memory-bytes 2700000000 \
--max-model-len 262144 --max-num-seqs 2 --max-num-batched-tokens 2048 \
--attention-backend B12X --linear-backend b12x --moe-backend b12x \
--engram-config '{"table_memory":"disk","cpu_offload":false}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_tensor_parallel_size":2}' \
--tokenizer-mode deepseek_v41 --reasoning-parser deepseek_v41 \
--tool-call-parser deepseek_v41 --enable-auto-tool-choice
# launcher env: ST_EXL3_PLAN=/plans/J268-ho-v31-K5n.json (absolute qdir), ST_EXL3_PREFILL=st
The validated runs mounted the full upstream checkpoint as /model; this repo carries the same tensors
minus the backbone routed experts, which the EXL3 path never reads.
Files
| Path | What it is |
|---|---|
model-00001..00004-of-00006.safetensors |
upstream non-expert tensors, byte copies (18.4 GB) |
model-00005/00006-of-00006.safetensors.part00..05 |
Engram shards (layers 1 and 14) split into ≤17 GB parts |
model.safetensors.index.json |
weight map for the six reassembled shards |
reassemble.sh, sha256-manifest.txt |
Engram rebuild + whole-file hashes of the native shards |
exl3/qn/K5/L??.part.* |
K5 expert banks, layers 0-39 |
exl3/q31/K{2,3}/L??.part.* |
K2 / K3 expert banks, layers 0-11 |
exl3/q31s/K{2,3}/L??.part.* |
K2 / K3 expert banks, layers 12-39 |
exl3/plan/*.json |
per-layer, per-expert K allocation (qdir relative to exl3/) |
config.json, tokenizer*.json, encoding/, inference/, LICENSE |
upstream, unchanged |
repack-report.json |
repack verification receipt |
Verification
- Dropped exactly the 92,160 backbone routed-expert tensors (
layers.{0-39}.ffn.experts.{0-383}.w{1,2,3}.{weight,scale}= 40 x 384 x 6). Nothing else was dropped; themtp.*draft experts are kept. - Every kept tensor's bytes were hashed at read time and re-hashed from the written shard: all 3,913
non-Engram tensors match, dtypes and shapes match. Engram shards are verified by whole-file sha256
against the source. Receipt:
repack-report.json.
Provenance
| Source | deepseek-ai/DeepSeek-V4.1-Flash rev fb2764a5cf321eaa5070ca8f9e892818f477c16d |
| Allocation plan | J268-ho-v31-K5n (night-5 deploy): K2/K3 banks from the v3.1 calibration run, the remaining 1,782 experts re-encoded at K5 |
| Encoder / kernels | exllamav3 1.5.3 d3739fd393337b1ff4d6c2a342b12f0c87a9592f |
| Serving runtime | local-inference-lab vLLM r38 66c293578412417476f842c1da5805d3a3d959a8 + EXL3 MoE method |
| Pipeline code | a9ea365bbb8a3b5d4eaaf7097c3f3995e217f85e |
License and credits
MIT, carried from DeepSeek's release; this repo adds no terms. Thanks to DeepSeek for the model, turboderp for ExLlamaV3 / EXL3 (MIT) — the trellis encoder and kernels used here — and Local Inference Lab and the vLLM community for the Spark serving stack.
- Downloads last month
- 201
Model tree for 0xSero/DeepSeek-V4.1-Flash-Spark
Base model
deepseek-ai/DeepSeek-V4.1-Flash