Instructions to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16 # Run inference directly in the terminal: llama cli -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16 # Run inference directly in the terminal: llama cli -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Use Docker
docker model run hf.co/Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
- LM Studio
- Jan
- vLLM
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
- Ollama
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with Ollama:
ollama run hf.co/Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
- Unsloth Desktop
- Pi
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with Docker Model Runner:
docker model run hf.co/Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
- Lemonade
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This is an abliterated research model with substantially reduced safety alignment. It may produce harmful, illegal, offensive, biased, or otherwise unsafe content. Access is provided for legitimate research and controlled evaluation. You are responsible for lawful use, downstream safeguards, and compliance with the Qwen Community License 1.0.
Log in or Sign Up to review the conditions and access this model content.
Qwen3.8-Flash-Next-Uncensored AD-4.27 GGUF
Compatibility notice
For standard/mainline llama.cpp builds, download all 33 files containing
mainlinein their names, keep the filenames unchanged, and pass shard 1 to--model.The original 34-shard files contain an attached experimental MTP/NextN layer and require the unmerged llama.cpp Qwen3.8 MTP implementation from PR #28243, commit
d1a92352cbd417fd840b4e765c0b82f5fe3d1d89. They are not compatible with standard llama.cpp b10941.Do not mix the 33-shard
mainlinefiles with the original 34-shard attached-MTP files.
Research artifact with substantially reduced safety alignment. This model is derived from an abliterated checkpoint and may comply with harmful, unethical, illegal, offensive, biased, or otherwise unsafe requests that the aligned model would refuse. It has no dependable built-in guardrails. Do not expose it to end users or production traffic without independently designed safety, moderation, access-control, logging, and abuse-prevention measures. You are responsible for how you use it and for compliance with applicable law.
A reproducible, tensor-specific mixed GGUF quantization of orcarouter/Qwen3.8-Flash-Next-Uncensored, pinned at revision 8336e613ea508b13c2159bd0f68965d97a606b95.
Why this quantization exists
The practical target of this build is to run Qwen3.8 Flash-Next—including its vision projector and, when using the experimental variant, its matching MTP draft path—on a machine with 64 GB of aggregate VRAM. Each model variant is larger than aggregate VRAM, so the complete model is not meant to reside in VRAM. The 38.4 GB PLE n-gram table is isolated in shard 2 and left SSD-pageable through mmap; the rest of the model can be placed on GPU while retaining PLE, native long-context configuration, and vision support.
This build applies AtomicChat's published AD-4.27bpw-Q4_K_M-M64 tensor recipe and BF16 importance matrix to the target model only. The matching uncensored MTP/NextN weights were exported separately through llama.cpp's official --mtp path and then attached without applying the target-only importance matrix to MTP tensors. The F16 vision projector is included.
This is an independent community build. It is not produced, endorsed, or warranted by Qwen, Alibaba, OrcaRouter, AtomicChat, or llama.cpp.
Contents
| Component | Format | Notes |
|---|---|---|
| Recommended mainline target-only model | 33 GGUF shards | Standard/mainline llama.cpp compatible; filenames contain mainline; 94,525,395,584 bytes (about 88.03 GiB) |
| Experimental target + attached MTP model | 34 GGUF shards | Requires the unmerged Qwen3.8 MTP implementation from llama.cpp PR #28243; 98,208,562,784 bytes (about 91.47 GiB) |
| PLE n-gram table | Isolated in shard 2 of each variant | Q5_1, intended to remain SSD-pageable with mmap enabled |
| Vision projector | F16 GGUF | mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf |
| Mainline checksums | SHA256SUMS-MAINLINE |
SHA-256 for all 33 recommended mainline shards |
| Experimental checksums | SHA256SUMS |
SHA-256 for the original attached-MTP GGUF set |
Tensor recipe
| Tensor group | Quantization |
|---|---|
per_layer_token_embd PLE table |
Q5_1 |
ffn_gate_exps and ffn_up_exps, blocks 0–3 and 40–47 |
IQ3_S |
Remaining ffn_gate_exps and ffn_up_exps |
IQ2_S |
ffn_down_exps |
IQ4_NL |
| Other quantized target tensors | Predominantly Q8_0 |
| MTP/NextN tensors | Exported separately from the pinned uncensored BF16 checkpoint; not quantized with the target imatrix |
The 4.27 bpw name describes the measured mixed target recipe, not a uniform tensor type. Some tools may display a representative GGUF ftype such as IQ2_S; that does not describe the full tensor mix.
Verified compatibility and validation
The two distributed variants have deliberately different runtime requirements:
| Variant | Expected loader | Verified result |
|---|---|---|
33-shard mainline target-only set |
Standard llama.cpp | Full GPU load and token generation succeeded with llama.cpp b10941, commit 4a89937354190cef5a97baf8eeb17336105eb72d |
| Original 34-shard attached-MTP set | Experimental PR #28243 build | Target + attached MTP loads and runs with commit d1a92352cbd417fd840b4e765c0b82f5fe3d1d89 |
| Original 34-shard attached-MTP set | Standard llama.cpp b10941 | Expected incompatibility reproduced: wrong number of tensors; expected 1256, got 1224 |
The corrected mainline set was verified as 48 target layers, 1,224 tensors, and 33 shards, with no blk.48.* MTP tensors and no nextn_predict_layers metadata. The validation run loaded the model across two NVIDIA GPUs (approximately 27.1 GiB and 26.4 GiB allocated) and generated tokens successfully. That mainline validation run was deliberately limited to a very short smoke test to prove that the corrected GGUF could load and generate. Its timing is not reported as a mainline performance benchmark because startup and short-generation overhead dominate such a small sample.
Measured production throughput: attached-MTP variant
The original 34-shard attached-MTP variant is also used continuously in our own llama.cpp deployment. The following are measured server results from the deployed model, not estimates. The server used the pinned PR #28243 revision, MTP speculative decoding, two NVIDIA GPUs with 64 GiB aggregate VRAM, mmap for the SSD-pageable PLE table, and a roughly 31K-token working conversation.
| Workload | Newly processed input | Reused KV cache | Generated output | Server processing time | Generation speed |
|---|---|---|---|---|---|
| Short incremental turn over a warm long-context cache, non-reasoning | 73 tokens | 30,600 tokens | 491 tokens | 10.42 s | 54.94 tok/s |
| Long-context cold request, non-reasoning | 30,592 tokens | 0 | 7 tokens | 209.35 s | 29.39 tok/s* |
| Long-context cold request, reasoning enabled | 31,607 tokens | 0 | 66 tokens | 204.76 s | 46.31 tok/s |
| Short incremental turn over a warm long-context cache, reasoning enabled | 70 tokens | 31,673 tokens | 50 tokens | 2.48 s | 48.25 tok/s |
* The 29.39 tok/s row generated only seven tokens, so that generation-rate sample is not representative of sustained decoding. Its useful measurement is the 209.35-second server time to ingest and answer from approximately 30.6K uncached input tokens. With the long-context KV cache retained, normal follow-up turns in the same conversation completed in approximately 2.48–10.42 seconds while decoding at about 48–55 tok/s.
These production figures describe the experimental attached-MTP deployment, not a same-condition benchmark of the target-only mainline set. Throughput varies with output length, GPU generation, PCIe topology, context length, batch size, cache state, and tensor placement.
Recommended: standard/mainline llama.cpp (33 shards)
Use the 33 target-only files whose names contain mainline. Keep all shard filenames unchanged, place them together, and pass the first shard to --model. This corrected variant contains 48 target layers and 1,224 tensors, with no attached blk.48 MTP tensors and no nextn_predict_layers metadata. It was load-and-generation tested with standard llama.cpp b10941.
Keep mmap enabled so the isolated PLE shard can remain SSD-pageable. Use --fit off to preserve the intended placement.
llama-server \
--model Qwen3.8-Flash-Next-Uncensored-AD-4.27-mainline-00001-of-00033.gguf \
--mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf \
--no-mmproj-offload --image-min-tokens 1024 \
--ctx-size 262144 --cache-type-k q8_0 --cache-type-v q8_0 \
--gpu-layers 999 --flash-attn on --fit off --jinja
Verify this set with SHA256SUMS-MAINLINE. Do not add --spec-type draft-mtp when using the target-only mainline set.
Experimental: attached MTP/NextN (34 shards)
The original 34-shard files remain available as an experimental attached-MTP variant. They require the unmerged Qwen3.8 MTP implementation from llama.cpp PR #28243, specifically commit:
d1a92352cbd417fd840b4e765c0b82f5fe3d1d89
They do not load correctly with standard llama.cpp b10941. Use this path only if you intentionally built that experimental implementation.
Build the required experimental llama.cpp revision
The 34-shard attached-MTP files require the exact unmerged PR revision below. A normal checkout of standard llama.cpp is not sufficient.
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/28243/head:qwen38-flash-next-mtp
git checkout d1a92352cbd417fd840b4e765c0b82f5fe3d1d89
cmake -S . -B build \
-DGGML_CUDA=ON \
-DLLAMA_CURL=OFF \
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j --target llama-server
./build/bin/llama-server --version
Confirm that the reported commit begins with d1a92352 before launching the 34-shard model. On systems where CUDA architecture detection is unsuitable, add the architecture explicitly; for example, Tesla P40 is compute capability 6.1, so add -DCMAKE_CUDA_ARCHITECTURES=61 to the CMake configure command. Multi-GPU/NCCL options are hardware- and topology-specific and are not required merely to parse the attached-MTP GGUF layout.
llama-server \
--model Qwen3.8-Flash-Next-Uncensored-AD-4.27-main-00001-of-00034.gguf \
--mmproj mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf \
--no-mmproj-offload --image-min-tokens 1024 \
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.75 \
--spec-draft-type-k f16 --spec-draft-type-v f16 \
--ctx-size 262144 --cache-type-k q8_0 --cache-type-v q8_0 \
--gpu-layers 999 --flash-attn on --fit off --jinja
Verify the original set with SHA256SUMS. Never mix shards from the two variants.
Hardware-specific flags such as --tensor-split, device placement, batch sizes, and tensor overrides must be adapted to the host. For a dual-32-GB setup, begin with one slot, --tensor-split 0.50,0.50, target Q8 KV, and—only for the experimental attached-MTP variant—draft F16 KV.
Provenance and attribution
- Qwen / Alibaba:
Qwen/Qwen3.8-Flash-Next, the upstream model and architecture. - OrcaRouter:
orcarouter/Qwen3.8-Flash-Next-Uncensored, revision8336e613ea508b13c2159bd0f68965d97a606b95, the BF16 abliterated source checkpoint, including the matching vision and MTP weights. - AtomicChat:
AtomicChat/Qwen3.8-Flash-Next-GGUF, the published AD-4.27 tensor recipe and BF16 importance matrix. The matrix used here had SHA-2565591ce3dc3bf0b73d3c074bc588c90b6c4f7b3c273de6b10e50d111b75f05487. - llama.cpp: conversion, quantization, sharding, GGUF loading, multimodal inference, and MTP runtime.
- Navin Model Repository: independent conversion of the pinned OrcaRouter checkpoint, target-only recipe application, separate MTP export and attachment, sharding, and checksums.
See REPRODUCIBILITY.md for the exact construction path and ATTRIBUTION.md for notices.
License and access conditions
The repository metadata of the OrcaRouter source says apache-2.0, but the actual LICENSE file distributed in the pinned source checkpoint—and the upstream Qwen model's current license—is Qwen Community License 1.0. To avoid granting rights that the publisher may not possess, this repository applies and includes the actual Qwen Community License 1.0. The more permissive Apache label is not relied upon here.
The Qwen Community License 1.0 permits use, copying, modification, publication, distribution, sublicensing, sale, deployment, hosting, fine-tuning, and derivative works, subject to its conditions. Among other requirements:
- retain the Qwen copyright and permission notice in copies or substantial portions;
- comply with applicable laws and third-party intellectual-property rights;
- prominently display the applicable model name when the license's large-service threshold applies;
- obtain a separate Qwen license before certain commercial uses if the licensee or an affiliate conducts a Model-as-a-Service or AI Work Assistant business, as defined in the license.
Read the complete LICENSE; this summary is not a substitute for it and is not legal advice. No patent, trademark, endorsement, warranty, or other right is granted beyond the included license and applicable source terms.
The OrcaRouter source access notice states that the abliterated model is released strictly for legitimate research and that downloading or using it acknowledges the stated safety warning and responsibility. This gated repository preserves that notice and is intended for legitimate research, interpretability, AI-safety/refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments.
By requesting access to, downloading, or using this release, you acknowledge the safety notice above, accept the included Qwen Community License 1.0, and assume responsibility for lawful use and appropriate downstream safeguards.
Warranty disclaimer
The model, projector, metadata, documentation, and outputs are provided "AS IS", without warranty of any kind. To the maximum extent permitted by applicable law, the contributors and upstream authors disclaim liability for claims, damages, misuse, or other consequences arising from use. This notice does not limit obligations or rights that cannot legally be limited.
- Downloads last month
- 1,724
We're not able to determine the quantization variants.
Model tree for Navin-Models/Qwen3.8-Flash-Next-Uncensored-AD-4.27-GGUF
Base model
Qwen/Qwen3.8-Flash-Next