Instructions to use igorls/Qwen3.8-Flash-Next-mixed-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use igorls/Qwen3.8-Flash-Next-mixed-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Qwen3.8-Flash-Next mixed quantization for NInfer
An early-access engineering preview of Qwen3.8-Flash-Next in the native .ninfer
format, prepared for the NInfer workstation fork.
This release distributes the artifact used by our Windows RTX PRO 6000 service,
including its text, Vision, MTP, tokenizer, and prompt-template data.
The artifact combines Primitive AI's NVFP4/FP8 compute weights with its INT4 PLE table. It preserves all 512 routed experts per layer. This is a format conversion and source splice, with no fine-tuning or expert pruning performed by this release. Credit for the base model belongs to Qwen and for the source quantizations to Primitive AI; see NOTICE.md.
Use the pinned NInfer fork below. This file cannot be loaded by Transformers, vLLM, llama.cpp, or Ollama. There is no Hugging Face hosted inference integration.
Update 2026-10-10: a NInfer v3 build of this model is now published alongside the v2 file. See NInfer v3 artifact below. The v2 file, its SHA256 and the
preview-2026-09-07revision described on this page are unchanged. The current NInferworkstationruntime loads only v3.
Download and identity
| Property | Value |
|---|---|
| File | qwen3_8_flash_next_mixed.ninfer |
| Download size | 113,298,397,952 bytes — 113.30 GB / 105.52 GiB |
| Container | NInfer v2 |
| Model identity | qwen3.8-flash-next |
| Weight identity | mixed-nvfp4-fp8-ple-int4 |
| Release revision | preview-2026-09-07 |
| Compatible engine revision | c81c88f6d66652be4607fb775333a6c63132f3f5 |
Using the Hugging Face CLI:
hf download igorls/Qwen3.8-Flash-Next-mixed-NInfer `
qwen3_8_flash_next_mixed.ninfer artifact-manifest.json SHA256SUMS `
LICENSE NOTICE.md LICENSE-APACHE-2.0.txt README.md `
--revision preview-2026-09-07 --local-dir .\models\flash-next
(Get-FileHash .\models\flash-next\qwen3_8_flash_next_mixed.ninfer -Algorithm SHA256).Hash
Expected SHA256:
3d383e51963aafd4318dfd04c8dc63ee7df11768de19d9ab58dbba44460d1d02
The manifest and checksum identify the exact release bytes. Keep them with the artifact when reproducing results.
Storage and memory
The on-disk inventory has 1,566 objects: 1,560 tensors and six frontend resources.
| Component | Stored format | Runtime placement |
|---|---|---|
| Main routed expert banks | 96 NVFP4 tensors; K16 scales and expert divisors | GPU |
| Main QSA/GDN projection parents | 96 FP8 E4M3 tensors with FP32 row scales | GPU |
| PLE embedding table | 128 INT4 group-16 tensors with FP16 scales | Host file mapping / RAM page cache |
| PLE indexing metadata | Three I64 tensors | Host file mapping |
| MTP expert banks | Two BF16 tensors, included in the BF16 count below | Converted to NVFP4 by the loader when MTP is enabled |
| Other weights, including Vision and MTP tails | BF16; 1,237 BF16 tensors in total | GPU, subject to enabled features |
| Tokenizer, template and media configuration | Six embedded raw resources | Host |
MTP is BF16 on disk and NVFP4 during execution. The two stored MTP expert banks occupy 5,033,164,800 bytes. The pinned loader retains those source banks in the file mapping and creates 1,415,581,696 bytes of NVFP4 device payload using NInfer's weights-only quantizer. No quantization calibration dataset was added in this conversion. The optional separately spliced artifact with NVFP4 MTP already on disk is not the file published here.
With Vision and MTP enabled and the default BF16 embedding/output head, the other device-bound tensor payload totals 76,251,938,528 bytes (71.02 GiB), before the MTP NVFP4 buffers, alignment, execution workspaces, CUDA graphs, recurrent state and KV cache. This is a payload accounting figure, not total VRAM usage. The PLE table plus its indexing metadata occupies 32,000,162,072 bytes (29.80 GiB) of host-mapped data. The runtime warms the table's pages at startup; reclaiming this cache under memory pressure can hurt latency.
The development workstation has one RTX PRO 6000 Blackwell Workstation Edition (96 GB) and 128 GB-class system RAM (125.64 GiB visible to Windows). These are the observed host specifications, not a measured minimum-RAM requirement. Use an SSD with more than 114 GB free for this file and metadata; leave additional room for the build and download cache. We have not qualified this artifact on smaller GPUs.
Windows quick start
Build from an x64 Visual Studio developer shell with CUDA installed. This example
uses the workstation toolchain: Visual Studio 2026, CUDA 13.3, and CMake 4.3. NInfer
targets sm_120a; this release is intended for the RTX PRO 6000 Blackwell platform.
git clone https://github.com/igorls/ninfer.git
cd ninfer
git checkout c81c88f6d66652be4607fb775333a6c63132f3f5
cmake -S . -B build-win -G "Visual Studio 18 2026" -A x64 -DNINFER_BUILD_MEDIA=OFF
cmake --build build-win --config Release -j
Download the model into models/flash-next inside that checkout using the command
above, then start a bounded text-serving profile:
.\build-win\apps\Release\ninfer-serve.exe .\models\flash-next\qwen3_8_flash_next_mixed.ninfer `
--host 127.0.0.1 --port 8010 --model-id qwen3.8-flash-next `
--max-context 32768 --kv-capacity 65536 --max-concurrency 2 `
--prefill-chunk 8192 --desktop-reserve-gib 3 `
--kv-dtype fp8 --gdn-state-dtype bf16 --spec mtp --draft-tokens 4 `
--preserve-thinking
This is a conservative starting profile, not the settings used for every published
measurement. Wait for the server-ready message: model materialization, host-table
warm-up and graph creation happen before serving. --max-context limits each
request; --kv-capacity is shared across active requests. The engine supports up
to eight active requests, but a KV pool sized for fewer full-context requests can
queue long requests. MTP currently speculates at decode batch size one; larger
batches use ordinary batched decode.
For image/video input, build with NINFER_BUILD_MEDIA=ON and the FFmpeg/libcurl
dependencies described in the fork README,
then add --vision. The text-only build above does not enable image/video input.
The artifact always contains Vision and MTP weights, even when a feature is disabled.
An OpenAI-compatible structured-output request:
$body = @{
model = "qwen3.8-flash-next"
messages = @(@{role = "user"; content = 'Return a JSON object with a "greeting" string in Portuguese.'})
max_tokens = 128
reasoning_effort = "none"
response_format = @{type = "json_object"}
} | ConvertTo-Json -Depth 6
Invoke-RestMethod http://127.0.0.1:8010/v1/chat/completions `
-Method Post -ContentType "application/json" -Body $body
The fork also serves OpenAI Responses and Anthropic Messages, tool calls, streaming, and a supported JSON Schema subset using constrained decoding. Details and explicit unsupported-schema behavior are in the serving reference. JSON conformance and application-level answer correctness are separate properties.
Provenance and validation
| Input | Pinned revision | Contribution |
|---|---|---|
| Primitive mixed NVFP4/FP8 | a4e813ed3cfbbcc61e2929699eccb864a4dfa843 |
Compute backbone, BF16 MTP/Vision/tails and frontend; BF16 PLE shards omitted |
| Primitive INT4 PLE | da8b39586016d8325ac619be28ad77d6296625ec |
All 128 ples_int4 table shards |
| Official Qwen reference | de4b8e4d43b917e7706784d8bb445c9af86a3540 |
Retained license and byte-identical frontend reference |
The converter aggregates expert banks, arranges projection rows and scale planes
for NInfer, substitutes the INT4 PLE table, and embeds the six frontend resources.
It consumes the 296,502 non-PLE source tensors recorded in the mixed checkpoint
index and replaces its 128 BF16 PLE tensors. The sidecar recorded recipe
qwen3_8_flash_next_mixed-v1 and the two source revisions above, but did not record
the historical converter Git revision. That historical revision is unknown; the
compatible engine pin is not presented as the original converter revision. The
original BF16 checkpoint revision used by Primitive is also not independently
established by this release.
Release checks on September 7, 2026 read the actual file with NInfer's container reader, validating the v2 directory, format/shape byte sizes, offsets, alignments and file bounds. The full-file SHA256 was computed, and every embedded frontend resource matched both the recorded conversion hashes and the pinned official Qwen reference. This verifies artifact structure and identity, not full numerical equivalence to Qwen's original BF16 inference implementation.
Earlier Windows RTX PRO 6000 engineering checks cover serving, multi-position MTP acceptance, prefix reuse, Vision smoke requests, and constrained output. Workload definitions, timings and limitations are in the performance reference. This publication adds no benchmark campaign or quality score. Real application evaluation, including manual review of complex legal/evidence workflows, remains in progress. This preview does not establish production readiness for those tasks.
NInfer v3 artifact (added 2026-10-10)
The NInfer workstation runtime (igorls/ninfer, branch workstation, commit
2f08a0b2d16a21ce2e8e456961c535ddc3c78f40)
runs Flash-Next as a native qwen4_exp v3 package (text, MTP, Vision) and rejects v2 files. The v3
artifact below is the upgrade of the v2 file above; every v2 file in this repository is unchanged.
| Property | Value |
|---|---|
| Files | qwen3_8_flash_next_mixed.v3.ninfer plus .part-0001, .part-0002, .part-0003 (keep all four in one directory; load the entry file) |
| Total size | 109,681,105,664 bytes (102.15 GiB), payload 109,680,552,704 bytes |
| Container | NInfer v3, 1,566 objects, 1,668 bindings, 959 uses |
artifact_id |
1e7e026e9f634926ae26b80f0fc8591e |
| Source | the v2 file above, SHA256 3d383e51963aafd4318dfd04c8dc63ee7df11768de19d9ab58dbba44460d1d02 (checked by the upgrade) |
| Checksums | SHA256SUMS.v3, artifact-manifest.v3.json, qwen3_8_flash_next_mixed.v3.ninfer.upgrade.json |
4a3faef781244fa431af35523d5df00d4c255f0271061c8cfe6718b17f9c66dd qwen3_8_flash_next_mixed.v3.ninfer
e5d74e6a4874ef270f000799db5e171786dd685fb4c0d14c0baa075cae13cda8 qwen3_8_flash_next_mixed.v3.ninfer.part-0001
67a30b1e718f2205d45ed22235ea93d1c0f9f0bf7cba963155002b5bccbf7baa qwen3_8_flash_next_mixed.v3.ninfer.part-0002
8bc02c7328ff27ce22858ce447afeb5a60d4d01f1f3a12955f9e2a5220780565 qwen3_8_flash_next_mixed.v3.ninfer.part-0003
What changed from v2. 1,564 of 1,566 objects are the v2 bytes (checked against per-object v2
digests). The two BF16 MTP expert banks are stored as NVFP4 expert banks, the form the v2 loader
built on the device at every start, so the file is 3.62 GB smaller and v3 does no load-time weight
repacking. Layout is otherwise per the workstation reference docs/maintainer/qwen3.8-flash-next-artifact.md.
Derivable. The workstation docs define v3 as derived, never distributed: every consumer rebuilds it from the published v2 file. This repository publishes the derivation output as a convenience; the derivation remains the reference:
python -m tools.convert.qwen4_exp.upgrade qwen3_8_flash_next_mixed.ninfer qwen3_8_flash_next_mixed.v3.ninfer
python -m tools.convert.qwen4_exp.verify qwen3_8_flash_next_mixed.v3.ninfer --file-digests
Known digest difference (disclosed). Parts 1 to 3 match the digests pinned in the workstation
tree (tools/convert/qwen4_exp/source.py). The entry file does not: the tree pins
0de7b5f6c3ac9719e1c6e5811a765fda15fb4618444343502fc798dff2282714, while this file is
4a3faef781244fa431af35523d5df00d4c255f0271061c8cfe6718b17f9c66dd, so verify --file-digests
reports file digests differ from the pinned qualified v3 artifact. The derivation is
deterministic: two independent runs (Linux, Python 3.13, torch CPU; Windows, Python 3.11.15,
torch 2.14.1, numpy 2.4.6) produced identical bytes, the same artifact_id, 1,564 objects equal to
v2 and passing every other verify check. The cause of the difference from the pin has not been
isolated; the pin is expected to be stale. The published file is the one validated below.
Validation (Windows, RTX PRO 6000 Blackwell, workstation @ 2f08a0b2, greedy, 64K context).
Needle retrieval at 3k, 12k and 31k tokens is correct. All 10 test prompts produce text identical to
the same v3 file on a Colab RTX PRO 6000, and first-token logprobs are identical. Against v2 outputs
(from the previous runtime) 7 of 10 outputs are identical; the others differ only in wording.
Plain v3 decode is about 110 tok/s on short chat prompts; load takes about 38 s. MTP speculation
produced the same text but decoded slower on this short-chat test (about 50 tok/s). With production-like flags (--max-context 131072 --kv-capacity 262144 --max-concurrency 4 --prefill-chunk 8192 --kv-dtype fp8 --vision, no MTP) the same file loads in about 38 s, keeps needle retrieval at 3k/12k/31k tokens correct, decodes at about 100 to 106 tok/s on short chat, and reaches about 286 tok/s aggregate with four concurrent 200-token requests. 8 of the 10 prompts match the bf16-KV run exactly; the other two differ only in wording, as expected with fp8 KV. The bytes served were the uploaded files (the Hub LFS SHA256 of every v3 file equals SHA256SUMS.v3). This is an engineering smoke test, not a benchmark or quality score.
Run it with the workstation build at the commit above, for example:
hf download igorls/Qwen3.8-Flash-Next-mixed-NInfer `
qwen3_8_flash_next_mixed.v3.ninfer qwen3_8_flash_next_mixed.v3.ninfer.part-0001 `
qwen3_8_flash_next_mixed.v3.ninfer.part-0002 qwen3_8_flash_next_mixed.v3.ninfer.part-0003 `
SHA256SUMS.v3 artifact-manifest.v3.json --local-dir .\models\flash-next-v3
.\build-win\apps\Release\ninfer-serve.exe .\models\flash-next-v3\qwen3_8_flash_next_mixed.v3.ninfer `
--model-id qwen3.8-flash-next --max-context 65536 --kv-capacity 65536 --kv-dtype fp8 --vision
The v2 revision preview-2026-09-07 and the instructions above it still describe the v2 file.
License
The Qwen-derived model weights are distributed with the original Qwen Community License 1.0, reproduced verbatim in LICENSE. Its commercial-use conditions apply to derivatives; this is not an Apache-2.0-only model release. Primitive's PLE repository separately advertises Apache-2.0 metadata; the accompanying Apache license text and attribution notice preserve that information without replacing Qwen's conditions. The NInfer engine source has its own Apache-2.0 license.
- Downloads last month
- 1,651
Model tree for igorls/Qwen3.8-Flash-Next-mixed-NInfer
Base model
Qwen/Qwen3.8-Flash-Next