Qwen3.8-Flash-Next GSQ-RCO Coder IQ1_M, NInfer v3 artifact

A NInfer v3 artifact of ISTA-DASLab's Qwen3.8-Flash-Next GSQ-RCO Coder for NInfer-all, the master branch of iamwavecut/ninfer-all. The Coder release removes half of Qwen3.8-Flash-Next's routed experts, keeping the 256 of 512 per layer that code, agentic tool use, vision and spatial reasoning need (chosen with RCO against the full model), and stores the rest at about 3.5 bits per weight; IQ1_M names the effective 1.89 bits per parameter of the original model, not a block type. This artifact keeps each tensor's ggml blocks byte for byte (the hyper-connection projections excepted: they are re-quantized from BF16 to Q8_0), and the engine multiplies the blocks in place.

Stock NInfer builds refuse these files: the model family and the gguf_* formats exist only in that line; build master from commit 5395ec30b on. NInfer serves it with up to eight concurrent requests, prompt-prefix reuse, structured output and its Vision tower; This file includes the MTP block. Enable drafting with --spec mtp --draft-tokens 4.

What is inside

component representation
routed experts, 48 layers × 256 the GGUF's own blocks, one type per layer and projection: gate and up IQ2_S (20 layers), IQ3_XXS (17), IQ3_S (10), IQ4_XS (1); down IQ4_NL (39), Q2_0 (9); expert-major, so one expert is one contiguous range of bytes
shared experts the GGUF's own blocks per tensor: IQ4_NL and Q8_0 down, Q4_K, Q5_K, Q6_K and IQ4_XS gate and up
36 Gated DeltaNet and 12 sparse-attention (QSA) layers the GGUF's own blocks, Q4_K to Q6_K and IQ4_XS per tensor
output head, token table Q6_K, IQ4_XS
router (256 rows) BF16, as in the GGUF
hyper-connection projections ggml Q8_0: the GGUF's BF16 matrices re-quantized (October 9, 2026), the MTP block's as stored in its GGUF
GDN controls, norms, convolution, A_log, dt_bias BF16/FP32, restored from llama.cpp's exporter conventions
Vision tower the release's mmproj-Qwen3.8-Flash-Next-BF16.gguf in BF16 (0.9 GB), the same file as the unpruned releases': the Qwen3.5/3.6 tower, 27 blocks of width 1152, merging 2×2 patches onto the text model's 2,560, with the checkpoint's image and video preprocessor configurations
n-gram table not in this file: the model names its table by SHA-256 (dd55c289…6cc3), the same table as every Flash-Next release, published once as WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3
chat template Qwen/Qwen3.8-Flash-Next's chat_template.jinja

One file, Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer, of 32,685,469,696 bytes (30.44 GiB), next to its conversion report, SHA256SUMS, NOTICE and LICENSE. Original text/Vision conversion command, before the MTP attachment:

python3 -m tools.convert --model Qwen3.8-Flash-Next \
  --recipe qwen3_8_flash_next_gguf --components text,vision \
  --source gguf=Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf \
  --source ngram=Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf \
  --source vision=mmproj-Qwen3.8-Flash-Next-BF16.gguf \
  --device cpu --rows-per-chunk 65536 \
  --name qwen3.8-flash-next-coder --out Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer

The converter reads the table shard to record its digest; the rows themselves are in the table repository (this conversion read the Q2_0 release's shard, the same file). --model needs the configuration, tokenizer and preprocessor files of Qwen/Qwen3.8-Flash-Next; the converter takes the 256-expert count from the GGUF.

Running

Download the artifact and the n-gram table, then serve them from a build of master (its README; CMAKE_CUDA_ARCHITECTURES is 86, 89 or 120a, and 120a needs CUDA 13.1 or newer):

hf download WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3 \
  Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer --local-dir models
hf download WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 \
  Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer --local-dir models
M=models/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer
T=models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer

# One 24 GB GPU: the experts in page-locked host memory (25.1 GB), the most used of them cached on the GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next-coder --expert-residency host --max-context 32768

# One 24 GB GPU and little RAM: the experts stay in the file and stream into a GPU cache.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next-coder --expert-residency disk --max-context 32768

The Docker image (an NVIDIA driver of the CUDA 13 branch, 580 or newer, and the NVIDIA Container Toolkit) runs the same commands as serve with the files under /models, and answers on http://localhost:8080/v1:

docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
  -v "$PWD/models:/models" -v ninfer-cache:/cache \
  ghcr.io/iamwavecut/ninfer-all serve /models/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-ninfer-v3.ninfer \
  --ngram-table /models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer \
  --model-id qwen3.8-flash-next-coder --expert-residency host --max-context 32768

A model without its n-gram table is refused; --no-ngram-table runs it without, a non-standard experimental mode with no practical use (see the table's card). --vision adds the Vision tower (0.9 GB on the GPU) for images and video. Qwen3.8-Flash-Next describes the placements and the execution.

October 9, 2026: Q8_0 hyper-connection projections

The hyper-connection projections were the largest read of a decoded token, 1.27 GB in BF16. This update stores every one of them as ggml Q8_0 (32 inputs of a row share a binary16 scale): the text model's, BF16 in the GGUF, quantized with ggml's reference rounding, and the MTP block's as the Unsloth MTP GGUF stores them (this file had decoded them to BF16). The file is 618,159,104 bytes smaller; every other object is byte for byte the previous one, which the rewrite checked on readback. hc-q8-requantization.json records the transformation and its error: over the text matrices the values differ from the BF16 ones by at most 0.0334 (4.79e-4 RMS). Builds before commit 944811b5c refuse this file.

Perplexity over two 40 KB slices of the repository's corpora (int8 KV, host experts), the start of the WikiText stream of perplexity-1m and of the held-out English Wikipedia stream, moved from 4.2254 to 4.1985 (-0.64%) and from 8.9087 to 8.9281 (+0.22%). Over so few tokens one quantization moves perplexity by up to 0.7% either way; the Q2_0 build, which shares these matrices, moved by +0.015% and +0.021% over the two whole quick corpora. Earlier quality results below were measured with the BF16 projections.

Decode speed was not measured for this model; the Q2_0 build decoded 3.6-6.0% faster on one RTX 3090 with host experts, and its dense weights shrank by 0.55 GiB, which the host expert cache uses. The measurements below predate this update.

The minimum build above also fixes host experts' prompts: from commit 5118e069d (October 9) on, a prompt call of 256 tokens or more with --expert-residency host could multiply some layers by other layers' experts. Over 40 KB of WikiText the Q2_0 model read a perplexity of about 4.0 instead of 2.54 with host experts; disk and device experts were not affected.

MTP attachment and qualification

The October 8, 2026 update adds 28 MTP objects (2,791,415,296 bytes) and 1,567 bindings. Every previous text weight, Vision weight, tokenizer resource and IQ4 table descriptor passed byte-preservation checks. The separate 28.80 GB IQ4 table is unchanged.

MTP comes from the Unsloth shared-Q8_0 release, through the pinned donor recorded in mtp-attachment.json. The original text/Vision conversion report remains in the repository. The base quantization labels describe the text model, not the added MTP block.

The public Engine passed host-to-GPU and disk-to-GPU MTP requests on one RTX 3090, with Vision enabled, int8 KV, 4,096 context capacity, 128-token prefill chunks and an 8 GiB device expert cache. All expert arithmetic ran on the GPU. Each mode used two fixed prompts, three greedy repetitions and 64 generated tokens per request, with prefix reuse disabled. Each fixed mode repeated exactly and accepted MTP drafts. Image requests identified the red test image with thinking disabled. Flash-Next disables MTP for media requests and uses plain decoding. A fresh text request resumed MTP after each image check. An initial check incorrectly required MTP for media and failed. Source inspection confirmed this existing limitation; the engine behavior was not changed.

Samples include the first request. Cache contents are not reset between repetitions. Other artifacts were being prepared on the same host during this campaign, so these timings do not isolate the performance effect of MTP.

The measurements below cover only those short requests. The OS page cache was not controlled. They do not establish cold-disk speed, low-RAM operation or RTX 3080 support. Different draft widths or modes can produce different tokens because their arithmetic rounds differently. Prior quality results below were measured without this MTP attachment.

Prompt 1 requests a Python merge function after a repeated prose prefix. Prompt 2 requests a simple explanation of the blue sky. Each row summarizes the three repetitions of that prompt.

prompt expert placement draft tokens minimum draft probability median decode tok/s range tok/s
1 host 4 0 27.01 19.27–28.95
2 host 4 0 18.66 14.69–20.90
1 disk 4 0 12.06 11.79–12.31
2 disk 4 0 9.54 9.46–9.70

The probability-floor experiment uses the tested development snapshot. The normal MTP command above uses the default floor. mtp-qualification.json records every sample and the tested source snapshot; the minimum commit above identifies artifact support.

Quality

The Coder card reports SWE-bench Verified 75.60 and LiveCodeBench v6 86.28 at xhigh reasoning effort in llama.cpp (BF16 base: 82.80, 87.43). These benchmarks have not been run in NInfer.

Speed

October 2026, greedy decoding, CUDA 12.8 (an sm_86 build), the repository's generate test (ninfer_qwen4_exp_generate_real) on one NVIDIA L40S (48 GB, PCIe) with the table read from its own artifact. Decode is measured over the 70 tokens of a short answer to a 20-token prompt (the expert cache still filling) and over the first five tokens after a 4,463-token prompt; prefill is that prompt in 512-token chunks. On 48 GB the device expert cache (39.9 GB) holds every expert of this build once it has seen them, so the long-prompt figures are an upper bound for a 24 GB card, whose cache holds about half of them.

placement decode, short answer decode after 4,463 tokens prefill
experts in pinned host memory (25.1 GB) 34.7 tok/s 80.4 tok/s 1,648 tok/s
experts on disk, the file in the page cache 41.6 tok/s 74.5 tok/s 1,938 tok/s

Credits and license

Quantized and pruned weights: ISTA-DASLab, Qwen3.8-Flash-Next GSQ-RCO Coder GGUF, produced with GSQ and RCO from Qwen/Qwen3.8-Flash-Next. The model is released under the Qwen Community License 1.0, reproduced in LICENSE, and its conditions apply to these files; ISTA-DASLab publish their GGUFs under Apache-2.0. Attribution notices are collected in NOTICE. This repository only re-packs those weights into NInfer's container.

Downloads last month
145
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3

Papers for WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3