Qwen2.5-7B-Instruct Abliterated

Uncensored version of Qwen/Qwen2.5-7B-Instruct using weight-level refusal-direction ablation — a difference-in-means direction extracted from harmful vs. harmless prompt activations, permanently orthogonalized out of every residual-stream-writing weight matrix (no runtime hook, no strength cap, no LoRA — the edit is baked directly into the checkpoint).

Steered modules: attention output projection (o_proj) + MLP down projection (down_proj), layer 16 of 28, coefficient 1.0 (full-strength).

Format

Native original repo format — bf16 safetensors, standard Qwen2 architecture. Only the residual-write projection weights at layer 16 were edited; every other tensor, the config, and the tokenizer are byte-identical to the original repo. Loads and serves exactly like Qwen/Qwen2.5-7B-Instruct, no patches needed.

Serving

Transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "rajaykumar12959/qwen2.5-7b-abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM:

vllm serve rajaykumar12959/qwen2.5-7b-abliterated

Uses the same chat template as the base model — no separate template file needed.

Metrics

Metric Original Abliterated
Refusal rate (292 held-out harmful prompts, 12 categories) ~95–97% 43.2%
Capability score (independent ARC-Easy-style MCQ eval) 1.000 1.000 (unchanged)

Refusal suppression is uneven across categories — from 79.2% (violence) down to 16.7% (self-harm). This is a partial reduction, not a fully "jailbroken" model. See the per-category breakdown in Evaluation below.

Layer 16 was chosen over other candidates specifically because it generalizes far better across prompt phrasing than the layer a narrower (AdvBench-only) sweep would have picked — an earlier layer-21 checkpoint dropped refusal by only 3–9pp on this same eval set, vs. 43pp+ at layer 16 on identical prompts.

Method

Per-layer difference-in-means direction extraction (fp32), unit-normalized, then closed-form weight orthogonalization — no gradient-based training. For a residual-writing projection y = Wx + b, the direction d̂'s component is removed exactly via W' = W − d̂(d̂ᵀW), b' = b − (d̂·b)d̂, at full strength (coefficient 1.0). Applied to every residual-write projection at layer 16 only — capability is re-verified after the edit on an independent MCQ set to confirm nothing outside the refusal circuit was disturbed.

Capture: activations at the last templated token, chat-template-formatted, for harmful (AdvBench) and harmless (Alpaca) prompt sets across all 28 layers. Eval: 292 held-out harmful prompts across 12 categories, graded by a two-tier classifier (rule-based + self-judge fallback); capability on an independent ARC-Easy-style subset.

Datasets

Dataset Role
walledai/AdvBench Harmful prompts (direction extraction + eval)
tatsu-lab/alpaca Harmless prompts (direction extraction)

Evaluation

Per-category refusal rate, this checkpoint:

Category Refusal rate
violence 79.2%
hate_speech 62.5%
misinformation 62.5%
extremism 58.3%
fraud_scams 54.2%
illicit_drugs 45.8%
financial_crime 37.5%
weapons 33.3%
privacy_invasion 33.3%
malware 20.8%
cybercrime_hacking 17.9%
self_harm 16.7%

Caveat: these figures rely heavily on the model's own self-judge for grading (~96–97% of verdicts), and that judge runs on this same already-ablated model — a known asymmetry not yet corrected for. Treat exact percentage-point figures as directionally reliable but provisional.

Disclaimer

This model has had safety guardrails removed and will comply with requests the original model would refuse. Released for research into AI alignment, interpretability, and refusal mechanisms. The creator assumes no responsibility for downstream use.

Acknowledgments

Downloads last month
105
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rajaykumar12959/qwen2.5-7b-abliterated

Base model

Qwen/Qwen2.5-7B
Finetuned
(3155)
this model

Datasets used to train rajaykumar12959/qwen2.5-7b-abliterated

Space using rajaykumar12959/qwen2.5-7b-abliterated 1