Qwen3.8 27B-Uncensored-MC
Qwen3.8-27B-Uncensored, OrcaRouter's abliterated BF16 build of Qwen's Qwen3.8-27B, converted to MegaCapybara's MX weight formats. It comes in five sizes for one NVIDIA RTX 5090, made the same way as perkel/Qwen3.8-27B-MC.
Safety alignment largely removed. The source model had its refusal behaviour removed by abliteration. It follows requests that the original Qwen3.8-27B would refuse, including harmful ones, and has no built-in guardrails.
It is meant for research, red-teaming, refusal and interpretability studies, and controlled experiments. Do not serve it to others without your own safety and moderation layers. You are responsible for how you use it and for what it generates, under the Apache 2.0 license and the laws that apply to you.
Run it with MegaCapybara
These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).
- Download MegaCapybara from its releases and unpack it.
- Run
MegaCapybaraLauncherand press the arrow next to the model to download a size from this repository's list (Qwen3.8-27B-Uncensored-MC); what it needs beside it comes with it. - Pick a preset, press Load server, and point your client at
http://127.0.0.1:8080/v1.
Without the launcher, put the files in the package's model folder and start a preset: the
usage guide has the details.
The source
- What OrcaRouter changed (from their model card, after Arditi et al. 2024, Refusal in Language Models Is Mediated
by a Single Direction):
- One refusal direction was estimated at layer 38.
- It was orthogonalized out of every matrix that writes to the residual stream: the attention and DeltaNet output projections, the MLP down projections, the embedding and the MTP head's. That is 131 matrices.
- The vision tower is untouched.
- What stays the same: the architecture, tensor names, tokenizer and chat template are Qwen's.
- On our test text, its full-precision outputs match the original's within 0.5% in perplexity on reasoning, math, code and wiki. They differ on chat (where refusals live) and on tool-call text.
Files
| File | Launcher name | Weight formats | On disk | In VRAM | KL | Top-1 |
|---|---|---|---|---|---|---|
qwen3.8-27b-uncensored-tiny.mcapy |
Qwen3.8 27B-Uncensored-MCTiny | all MXFP4 | 16.6 GB | 12.70 GiB | 0.0662 | 92.7% |
qwen3.8-27b-uncensored-small.mcapy |
Qwen3.8 27B-Uncensored-MCSmall | MXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs | 17.6 GB | 13.67 GiB | 0.0574 | 93.7% |
qwen3.8-27b-uncensored-medium.mcapy |
Qwen3.8 27B-Uncensored-MCMedium | MXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs | 19.5 GB | 15.39 GiB | 0.0261 | 95.6% |
qwen3.8-27b-uncensored-large.mcapy |
Qwen3.8 27B-Uncensored-MCLarge | all MXFP6 | 22.9 GB | 18.66 GiB | 0.0074 | 97.9% |
qwen3.8-27b-uncensored-xxl.mcapy |
Qwen3.8 27B-Uncensored-MCXXL | MXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms) | 28.2 GB | 23.58 GiB | 0.0057 | 98.2% |
You need one model file. These support files are shared by all five:
| File | What it is | On disk |
|---|---|---|
qwen3.8-27b-dflash2.mcapy |
DFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster) | 1.3 GB |
qwen3.8-27b-uncensored-mtp.mcapy |
the uncensored model's own MTP layer (abliterated with it) as a smaller, slower drafter | 0.28 GB |
qwen3.8-27b-vision-fp8.mcapy |
the vision tower in FP8 (images; nearly exact; abliteration left it unchanged) | 0.57 GB |
qwen3.8-27b-vision-bf16.mcapy |
the original BF16 vision tower | 1.0 GB |
- Using both repositories: they can share one
model\folder. The launcher lists each model by its name, and gives each the MTP layer of its own source. - The drafter: DFlash2 was trained on the original model and is used here as is. Drafting never changes the answers, since the model checks every drafted token. It may change speed; see below.
Accuracy against its own original
- What KL and top-1 measure here: KL divergence and top-1 agreement against the uncensored model's own BF16 weights, computed with FP32 math. This is how much the conversion loses; the uncensoring itself is not counted.
- The test text: 81,880 held-out tokens of reasoning traces, math, code, chat and wiki.
- Settings: the engine at 16-bit activations and an unquantized KV cache.
- The chart:
- Unsloth's points are their published top-1 for GGUFs of Qwen's original model, on their own text. They are there to show what each size costs, not to compare the two models.
- Sizes are the bytes on the GPU (no embedding table or MTP layer).
- Against our conversion of the original: top-1 is the same at every size, within 0.4 points.
KL per kind of text
| Size | Reasoning | Math | Code | Chat | Wiki |
|---|---|---|---|---|---|
| Tiny | 0.0138 | 0.0156 | 0.0508 | 0.2094 | 0.0414 |
| Small | 0.0107 | 0.0139 | 0.0450 | 0.1899 | 0.0277 |
| Medium | 0.0059 | 0.0080 | 0.0282 | 0.0689 | 0.0196 |
| Large | 0.0009 | 0.0015 | 0.0092 | 0.0227 | 0.0028 |
| XXL | 0.0007 | 0.0014 | 0.0081 | 0.0162 | 0.0022 |
Speed on an RTX 5090
Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.
| Size | Coding, 1 conversation | Prose, 1 conversation | Coding, 8 at once (total) | Reading an 8K-token prompt |
|---|---|---|---|---|
| Tiny | 493 | 250 | 1,856 | 7,854 |
| Small | 488 | 236 | 1,828 | 7,697 |
| Medium | 445 | 219 | 1,669 | 7,365 |
| Large | 375 | 178 | 1,444 | 6,982 |
| XXL | 312 | 149 | 1,314 | 6,212 |
- Against our conversion of the original: the same within run-to-run differences, except Tiny's coding: 493 against 540. These are single runs.
- Conditions: measured with nothing else on the card. Its memory was overclocked to 16.8 GHz (+20%); a stock card is slower.
How the files were made
Source: orcarouter/Qwen3.8-27B-Uncensored, revision
404ea47, BF16. It is the reference for every score above.Recipes: the same as for perkel/Qwen3.8-27B-MC.
- The weights are rounded layer by layer with GPTQ and MX block scales (32 weights share a power-of-two scale).
- The calibration text is the same 514K tokens.
- Each size uses the same per-layer format choices.
Storage changes for the engine:
- norm gains are folded into the following matrices;
- the attention value rows are Hadamard-rotated;
- the output head's rows are ordered by token frequency.
None of these change the model's function. XXL's 16-bit parts are stored as two FP8 E4M3 terms.
License
Apache 2.0, as the original Qwen3.8-27B, OrcaRouter's
Qwen3.8-27B-Uncensored and
Qwen3.8-27B-DFlash2. The changes are listed in NOTICE.
Converted by Perkel's Software Corner.
- Downloads last month
- 56
