Instructions to use WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3 with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Qwen3.8-Flash-Next GSQ-RCO Coder IQ1_M, NInfer v3 artifact
A NInfer v3 artifact of ISTA-DASLab's
Qwen3.8-Flash-Next GSQ-RCO Coder
for NInfer-all, the master branch of iamwavecut/ninfer-all.
The Coder release removes half of Qwen3.8-Flash-Next's routed experts, keeping the 256 of 512 per
layer that code, agentic tool use, vision and spatial reasoning need (chosen with RCO against the
full model), and stores the rest at about 3.5 bits per weight; IQ1_M names the effective 1.89 bits
per parameter of the original model, not a block type. This artifact keeps each tensor's ggml
blocks byte for byte (the hyper-connection projections excepted: they are re-quantized from BF16 to Q8_0), and the engine multiplies the blocks in place.
Stock NInfer builds refuse these files: the model family and the gguf_* formats exist only in that
line; build master from commit 5395ec30b on.
NInfer serves it with up to eight concurrent requests, prompt-prefix reuse, structured output and
its Vision tower; This file includes the MTP block. Enable drafting with --spec mtp --draft-tokens 4.
What is inside
| component | representation |
|---|---|
| routed experts, 48 layers × 256 | the GGUF's own blocks, one type per layer and projection: gate and up IQ2_S (20 layers), IQ3_XXS (17), IQ3_S (10), IQ4_XS (1); down IQ4_NL (39), Q2_0 (9); expert-major, so one expert is one contiguous range of bytes |
| shared experts | the GGUF's own blocks per tensor: IQ4_NL and Q8_0 down, Q4_K, Q5_K, Q6_K and IQ4_XS gate and up |
| 36 Gated DeltaNet and 12 sparse-attention (QSA) layers | the GGUF's own blocks, Q4_K to Q6_K and IQ4_XS per tensor |
| output head, token table | Q6_K, IQ4_XS |
| router (256 rows) | BF16, as in the GGUF |
| hyper-connection projections | ggml Q8_0: the GGUF's BF16 matrices re-quantized (October 9, 2026), the MTP block's as stored in its GGUF |
GDN controls, norms, convolution, A_log, dt_bias |
BF16/FP32, restored from llama.cpp's exporter conventions |
| Vision tower | the release's mmproj-Qwen3.8-Flash-Next-BF16.gguf in BF16 (0.9 GB), the same file as the unpruned releases': the Qwen3.5/3.6 tower, 27 blocks of width 1152, merging 2×2 patches onto the text model's 2,560, with the checkpoint's image and video preprocessor configurations |
| n-gram table | not in this file: the model names its table by SHA-256 (dd55c289…6cc3), the same table as every Flash-Next release, published once as WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 |
| chat template | Qwen/Qwen3.8-Flash-Next's chat_template.jinja |
One file, Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer, of 32,685,469,696 bytes
(30.44 GiB), next to its conversion report, SHA256SUMS, NOTICE and LICENSE. Original text/Vision conversion
command, before the MTP attachment:
python3 -m tools.convert --model Qwen3.8-Flash-Next \
--recipe qwen3_8_flash_next_gguf --components text,vision \
--source gguf=Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf \
--source ngram=Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf \
--source vision=mmproj-Qwen3.8-Flash-Next-BF16.gguf \
--device cpu --rows-per-chunk 65536 \
--name qwen3.8-flash-next-coder --out Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer
The converter reads the table shard to record its digest; the rows themselves are in the table
repository (this conversion read the Q2_0 release's shard, the same file). --model needs the
configuration, tokenizer and preprocessor files of
Qwen/Qwen3.8-Flash-Next; the converter takes the
256-expert count from the GGUF.
Running
Download the artifact and the n-gram table, then serve them from a build of master
(its README; CMAKE_CUDA_ARCHITECTURES is
86, 89 or 120a, and 120a needs CUDA 13.1 or newer):
hf download WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3 \
Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer --local-dir models
hf download WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 \
Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer --local-dir models
M=models/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer
T=models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer
# One 24 GB GPU: the experts in page-locked host memory (25.1 GB), the most used of them cached on the GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next-coder --expert-residency host --max-context 32768
# One 24 GB GPU and little RAM: the experts stay in the file and stream into a GPU cache.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next-coder --expert-residency disk --max-context 32768
The Docker image (an NVIDIA driver of the CUDA
13 branch, 580 or newer, and the NVIDIA Container Toolkit) runs the same commands as serve with
the files under /models, and answers on http://localhost:8080/v1:
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
-v "$PWD/models:/models" -v ninfer-cache:/cache \
ghcr.io/iamwavecut/ninfer-all serve /models/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer \
--ngram-table /models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer \
--model-id qwen3.8-flash-next-coder --expert-residency host --max-context 32768
A model without its n-gram table is refused; --no-ngram-table runs it without, a non-standard
experimental mode with no practical use (see the table's card). --vision adds the Vision tower
(0.9 GB on the GPU) for images and video.
Qwen3.8-Flash-Next
describes the placements and the execution.
October 9, 2026: Q8_0 hyper-connection projections
The hyper-connection projections were the largest read of a decoded token, 1.27 GB in BF16. This
update stores every one of them as ggml Q8_0 (32 inputs of a row share a binary16 scale): the text
model's, BF16 in the GGUF, quantized with ggml's reference rounding, and the MTP block's as the
Unsloth MTP GGUF stores them (this file had decoded them to BF16). The file is 618,159,104 bytes
smaller; every other object is byte for byte the previous one, which the rewrite checked on
readback. hc-q8-requantization.json records the transformation and its error: over the text
matrices the values differ from the BF16 ones by at most 0.0334 (4.79e-4 RMS). Builds before
commit 944811b5c refuse this file.
Perplexity over two 40 KB slices of the repository's corpora (int8 KV, host experts), the start of the WikiText stream of perplexity-1m and of the held-out English Wikipedia stream, moved from 4.2254 to 4.1985 (-0.64%) and from 8.9087 to 8.9281 (+0.22%). Over so few tokens one quantization moves perplexity by up to 0.7% either way; the Q2_0 build, which shares these matrices, moved by +0.015% and +0.021% over the two whole quick corpora. Earlier quality results below were measured with the BF16 projections.
Decode speed was not measured for this model; the Q2_0 build decoded 3.6-6.0% faster on one RTX 3090 with host experts, and its dense weights shrank by 0.55 GiB, which the host expert cache uses. The measurements below predate this update.
The minimum build above also fixes host experts' prompts: from commit 5118e069d (October 9) on, a
prompt call of 256 tokens or more with --expert-residency host could multiply some layers by
other layers' experts. Over 40 KB of WikiText the Q2_0 model read a perplexity of about 4.0
instead of 2.54 with host experts; disk and device experts were not affected.
MTP attachment and qualification
The October 8, 2026 update adds 28 MTP objects (2,791,415,296 bytes) and 1,567 bindings. Every previous text weight, Vision weight, tokenizer resource and IQ4 table descriptor passed byte-preservation checks. The separate 28.80 GB IQ4 table is unchanged.
MTP comes from the Unsloth shared-Q8_0 release, through the pinned donor recorded in mtp-attachment.json. The original text/Vision conversion report remains in the repository. The base quantization labels describe the text model, not the added MTP block.
The public Engine passed host-to-GPU and disk-to-GPU MTP requests on one RTX 3090, with Vision enabled, int8 KV, 4,096 context capacity, 128-token prefill chunks and an 8 GiB device expert cache. All expert arithmetic ran on the GPU. Each mode used two fixed prompts, three greedy repetitions and 64 generated tokens per request, with prefix reuse disabled. Each fixed mode repeated exactly and accepted MTP drafts. Image requests identified the red test image with thinking disabled. Flash-Next disables MTP for media requests and uses plain decoding. A fresh text request resumed MTP after each image check. An initial check incorrectly required MTP for media and failed. Source inspection confirmed this existing limitation; the engine behavior was not changed.
Samples include the first request. Cache contents are not reset between repetitions. Other artifacts were being prepared on the same host during this campaign, so these timings do not isolate the performance effect of MTP.
The measurements below cover only those short requests. The OS page cache was not controlled. They do not establish cold-disk speed, low-RAM operation or RTX 3080 support. Different draft widths or modes can produce different tokens because their arithmetic rounds differently. Prior quality results below were measured without this MTP attachment.
Prompt 1 requests a Python merge function after a repeated prose prefix. Prompt 2 requests a simple explanation of the blue sky. Each row summarizes the three repetitions of that prompt.
| prompt | expert placement | draft tokens | minimum draft probability | median decode tok/s | range tok/s |
|---|---|---|---|---|---|
| 1 | host | 4 | 0 | 27.01 | 19.27–28.95 |
| 2 | host | 4 | 0 | 18.66 | 14.69–20.90 |
| 1 | disk | 4 | 0 | 12.06 | 11.79–12.31 |
| 2 | disk | 4 | 0 | 9.54 | 9.46–9.70 |
The probability-floor experiment uses the tested development snapshot. The normal MTP command above uses the default floor. mtp-qualification.json records every sample and the tested source snapshot; the minimum commit above identifies artifact support.
Quality
The Coder card reports SWE-bench Verified 75.60 and LiveCodeBench v6 86.28 at xhigh reasoning effort in llama.cpp (BF16 base: 82.80, 87.43). These benchmarks have not been run in NInfer.
Speed
October 2026, greedy decoding, CUDA 12.8 (an sm_86 build), the repository's generate test
(ninfer_qwen4_exp_generate_real) on one NVIDIA L40S (48 GB, PCIe) with the table read from its own
artifact. Decode is measured over the 70 tokens of a short answer to a 20-token prompt (the expert
cache still filling) and over the first five tokens after a 4,463-token prompt; prefill is that
prompt in 512-token chunks. On 48 GB the device expert cache (39.9 GB) holds every expert of this
build once it has seen them, so the long-prompt figures are an upper bound for a 24 GB card, whose
cache holds about half of them.
| placement | decode, short answer | decode after 4,463 tokens | prefill |
|---|---|---|---|
| experts in pinned host memory (25.1 GB) | 34.7 tok/s | 80.4 tok/s | 1,648 tok/s |
| experts on disk, the file in the page cache | 41.6 tok/s | 74.5 tok/s | 1,938 tok/s |
Credits and license
Quantized and pruned weights: ISTA-DASLab,
Qwen3.8-Flash-Next GSQ-RCO Coder GGUF,
produced with GSQ and RCO
from Qwen/Qwen3.8-Flash-Next. The model is
released under the Qwen Community License 1.0, reproduced in LICENSE, and its conditions
apply to these files; ISTA-DASLab publish their GGUFs under Apache-2.0. Attribution notices are
collected in NOTICE. This repository only re-packs those weights into NInfer's container.
- Downloads last month
- 145
Model tree for WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3
Base model
Qwen/Qwen3.8-Flash-Next