How to use from
Ollama
# Gated model: Login with a HF token with gated access permission
hf auth login
ollama run hf.co/Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Quick Links

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This is an abliterated research model with substantially reduced safety alignment. It may produce harmful, illegal, offensive, biased, or otherwise unsafe content. Access is provided for legitimate research and controlled evaluation. You are responsible for lawful use, downstream safeguards, and compliance with the Qwen Community License 1.0.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.8-Flash-Next-Uncensored AD-4.27 GGUF

Compatibility notice

For standard/mainline llama.cpp builds, download all 33 files containing mainline in their names, keep the filenames unchanged, and pass shard 1 to --model.

The original 34-shard files contain an attached experimental MTP/NextN layer and require the unmerged llama.cpp Qwen3.8 MTP implementation from PR #28243, commit d1a92352cbd417fd840b4e765c0b82f5fe3d1d89. They are not compatible with standard llama.cpp b10941.

Do not mix the 33-shard mainline files with the original 34-shard attached-MTP files.

Research artifact with substantially reduced safety alignment. This model is derived from an abliterated checkpoint and may comply with harmful, unethical, illegal, offensive, biased, or otherwise unsafe requests that the aligned model would refuse. It has no dependable built-in guardrails. Do not expose it to end users or production traffic without independently designed safety, moderation, access-control, logging, and abuse-prevention measures. You are responsible for how you use it and for compliance with applicable law.

A reproducible, tensor-specific mixed GGUF quantization of orcarouter/Qwen3.8-Flash-Next-Uncensored, pinned at revision 8336e613ea508b13c2159bd0f68965d97a606b95.

Why this quantization exists

The practical target of this build is to run Qwen3.8 Flash-Next—including its vision projector and, when using the experimental variant, its matching MTP draft path—on a machine with 64 GB of aggregate VRAM. Each model variant is larger than aggregate VRAM, so the complete model is not meant to reside in VRAM. The 38.4 GB PLE n-gram table is isolated in shard 2 and left SSD-pageable through mmap; the rest of the model can be placed on GPU while retaining PLE, native long-context configuration, and vision support.

This build applies AtomicChat's published AD-4.27bpw-Q4_K_M-M64 tensor recipe and BF16 importance matrix to the target model only. The matching uncensored MTP/NextN weights were exported separately through llama.cpp's official --mtp path and then attached without applying the target-only importance matrix to MTP tensors. The F16 vision projector is included.

This is an independent community build. It is not produced, endorsed, or warranted by Qwen, Alibaba, OrcaRouter, AtomicChat, or llama.cpp.

Contents

Component Format Notes
Recommended mainline target-only model 33 GGUF shards Standard/mainline llama.cpp compatible; filenames contain mainline; 94,525,395,584 bytes (about 88.03 GiB)
Experimental target + attached MTP model 34 GGUF shards Requires the unmerged Qwen3.8 MTP implementation from llama.cpp PR #28243; 98,208,562,784 bytes (about 91.47 GiB)
PLE n-gram table Isolated in shard 2 of each variant Q5_1, intended to remain SSD-pageable with mmap enabled
Vision projector F16 GGUF mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf
Mainline checksums SHA256SUMS-MAINLINE SHA-256 for all 33 recommended mainline shards
Experimental checksums SHA256SUMS SHA-256 for the original attached-MTP GGUF set

Tensor recipe

Tensor group Quantization
per_layer_token_embd PLE table Q5_1
ffn_gate_exps and ffn_up_exps, blocks 0–3 and 40–47 IQ3_S
Remaining ffn_gate_exps and ffn_up_exps IQ2_S
ffn_down_exps IQ4_NL
Other quantized target tensors Predominantly Q8_0
MTP/NextN tensors Exported separately from the pinned uncensored BF16 checkpoint; not quantized with the target imatrix

The 4.27 bpw name describes the measured mixed target recipe, not a uniform tensor type. Some tools may display a representative GGUF ftype such as IQ2_S; that does not describe the full tensor mix.

Verified compatibility and validation

The two distributed variants have deliberately different runtime requirements:

Variant Expected loader Verified result
33-shard mainline target-only set Standard llama.cpp Full GPU load and token generation succeeded with llama.cpp b10941, commit 4a89937354190cef5a97baf8eeb17336105eb72d
Original 34-shard attached-MTP set Experimental PR #28243 build Target + attached MTP loads and runs with commit d1a92352cbd417fd840b4e765c0b82f5fe3d1d89
Original 34-shard attached-MTP set Standard llama.cpp b10941 Expected incompatibility reproduced: wrong number of tensors; expected 1256, got 1224

The corrected mainline set was verified as 48 target layers, 1,224 tensors, and 33 shards, with no blk.48.* MTP tensors and no nextn_predict_layers metadata. The validation run loaded the model across two NVIDIA GPUs (approximately 27.1 GiB and 26.4 GiB allocated) and generated tokens successfully. That mainline validation run was deliberately limited to a very short smoke test to prove that the corrected GGUF could load and generate. Its timing is not reported as a mainline performance benchmark because startup and short-generation overhead dominate such a small sample.

Measured production throughput: attached-MTP variant

The original 34-shard attached-MTP variant is also used continuously in our own llama.cpp deployment. The following are measured server results from the deployed model, not estimates. The server used the pinned PR #28243 revision, MTP speculative decoding, two NVIDIA GPUs with 64 GiB aggregate VRAM, mmap for the SSD-pageable PLE table, and a roughly 31K-token working conversation.

Workload Newly processed input Reused KV cache Generated output Server processing time Generation speed
Short incremental turn over a warm long-context cache, non-reasoning 73 tokens 30,600 tokens 491 tokens 10.42 s 54.94 tok/s
Long-context cold request, non-reasoning 30,592 tokens 0 7 tokens 209.35 s 29.39 tok/s*
Long-context cold request, reasoning enabled 31,607 tokens 0 66 tokens 204.76 s 46.31 tok/s
Short incremental turn over a warm long-context cache, reasoning enabled 70 tokens 31,673 tokens 50 tokens 2.48 s 48.25 tok/s

* The 29.39 tok/s row generated only seven tokens, so that generation-rate sample is not representative of sustained decoding. Its useful measurement is the 209.35-second server time to ingest and answer from approximately 30.6K uncached input tokens. With the long-context KV cache retained, normal follow-up turns in the same conversation completed in approximately 2.48–10.42 seconds while decoding at about 48–55 tok/s.

These production figures describe the experimental attached-MTP deployment, not a same-condition benchmark of the target-only mainline set. Throughput varies with output length, GPU generation, PCIe topology, context length, batch size, cache state, and tensor placement.

Recommended: standard/mainline llama.cpp (33 shards)

Use the 33 target-only files whose names contain mainline. Keep all shard filenames unchanged, place them together, and pass the first shard to --model. This corrected variant contains 48 target layers and 1,224 tensors, with no attached blk.48 MTP tensors and no nextn_predict_layers metadata. It was load-and-generation tested with standard llama.cpp b10941.

Keep mmap enabled so the isolated PLE shard can remain SSD-pageable. Use --fit off to preserve the intended placement.

llama-server \
  --model Qwen3.8-Flash-Next-Uncensored-AD-4.27-mainline-00001-of-00033.gguf \
  --mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf \
  --no-mmproj-offload --image-min-tokens 1024 \
  --ctx-size 262144 --cache-type-k q8_0 --cache-type-v q8_0 \
  --gpu-layers 999 --flash-attn on --fit off --jinja

Verify this set with SHA256SUMS-MAINLINE. Do not add --spec-type draft-mtp when using the target-only mainline set.

Experimental: attached MTP/NextN (34 shards)

The original 34-shard files remain available as an experimental attached-MTP variant. They require the unmerged Qwen3.8 MTP implementation from llama.cpp PR #28243, specifically commit:

d1a92352cbd417fd840b4e765c0b82f5fe3d1d89

They do not load correctly with standard llama.cpp b10941. Use this path only if you intentionally built that experimental implementation.

Build the required experimental llama.cpp revision

The 34-shard attached-MTP files require the exact unmerged PR revision below. A normal checkout of standard llama.cpp is not sufficient.

git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/28243/head:qwen38-flash-next-mtp
git checkout d1a92352cbd417fd840b4e765c0b82f5fe3d1d89

cmake -S . -B build \
  -DGGML_CUDA=ON \
  -DLLAMA_CURL=OFF \
  -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j --target llama-server

./build/bin/llama-server --version

Confirm that the reported commit begins with d1a92352 before launching the 34-shard model. On systems where CUDA architecture detection is unsuitable, add the architecture explicitly; for example, Tesla P40 is compute capability 6.1, so add -DCMAKE_CUDA_ARCHITECTURES=61 to the CMake configure command. Multi-GPU/NCCL options are hardware- and topology-specific and are not required merely to parse the attached-MTP GGUF layout.

llama-server \
  --model Qwen3.8-Flash-Next-Uncensored-AD-4.27-main-00001-of-00034.gguf \
  --mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf \
  --no-mmproj-offload --image-min-tokens 1024 \
  --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.75 \
  --spec-draft-type-k f16 --spec-draft-type-v f16 \
  --ctx-size 262144 --cache-type-k q8_0 --cache-type-v q8_0 \
  --gpu-layers 999 --flash-attn on --fit off --jinja

Verify the original set with SHA256SUMS. Never mix shards from the two variants.

Hardware-specific flags such as --tensor-split, device placement, batch sizes, and tensor overrides must be adapted to the host. For a dual-32-GB setup, begin with one slot, --tensor-split 0.50,0.50, target Q8 KV, and—only for the experimental attached-MTP variant—draft F16 KV.

Provenance and attribution

  1. Qwen / Alibaba: Qwen/Qwen3.8-Flash-Next, the upstream model and architecture.
  2. OrcaRouter: orcarouter/Qwen3.8-Flash-Next-Uncensored, revision 8336e613ea508b13c2159bd0f68965d97a606b95, the BF16 abliterated source checkpoint, including the matching vision and MTP weights.
  3. AtomicChat: AtomicChat/Qwen3.8-Flash-Next-GGUF, the published AD-4.27 tensor recipe and BF16 importance matrix. The matrix used here had SHA-256 5591ce3dc3bf0b73d3c074bc588c90b6c4f7b3c273de6b10e50d111b75f05487.
  4. llama.cpp: conversion, quantization, sharding, GGUF loading, multimodal inference, and MTP runtime.
  5. Navin Model Repository: independent conversion of the pinned OrcaRouter checkpoint, target-only recipe application, separate MTP export and attachment, sharding, and checksums.

See REPRODUCIBILITY.md for the exact construction path and ATTRIBUTION.md for notices.

License and access conditions

The repository metadata of the OrcaRouter source says apache-2.0, but the actual LICENSE file distributed in the pinned source checkpoint—and the upstream Qwen model's current license—is Qwen Community License 1.0. To avoid granting rights that the publisher may not possess, this repository applies and includes the actual Qwen Community License 1.0. The more permissive Apache label is not relied upon here.

The Qwen Community License 1.0 permits use, copying, modification, publication, distribution, sublicensing, sale, deployment, hosting, fine-tuning, and derivative works, subject to its conditions. Among other requirements:

  • retain the Qwen copyright and permission notice in copies or substantial portions;
  • comply with applicable laws and third-party intellectual-property rights;
  • prominently display the applicable model name when the license's large-service threshold applies;
  • obtain a separate Qwen license before certain commercial uses if the licensee or an affiliate conducts a Model-as-a-Service or AI Work Assistant business, as defined in the license.

Read the complete LICENSE; this summary is not a substitute for it and is not legal advice. No patent, trademark, endorsement, warranty, or other right is granted beyond the included license and applicable source terms.

The OrcaRouter source access notice states that the abliterated model is released strictly for legitimate research and that downloading or using it acknowledges the stated safety warning and responsibility. This gated repository preserves that notice and is intended for legitimate research, interpretability, AI-safety/refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments.

By requesting access to, downloading, or using this release, you acknowledge the safety notice above, accept the included Qwen Community License 1.0, and assume responsibility for lawful use and appropriate downstream safeguards.

Warranty disclaimer

The model, projector, metadata, documentation, and outputs are provided "AS IS", without warranty of any kind. To the maximum extent permitted by applicable law, the contributors and upstream authors disclaim liability for claims, damages, misuse, or other consequences arising from use. This notice does not limit obligations or rights that cannot legally be limited.

Downloads last month
1,724
GGUF
Model size
180B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF

Quantized
(39)
this model