Qwen3.8 27B AEON Ultimate Uncensored โ€” 4-bit MLX

This is a 4-bit MLX quantization for Apple Silicon of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.

Original Uncensored model made by @SpaceTimeViking , complete credit to him and all his work https://x.com/SpaceTimeViking

All model and fine-tuning credit belongs to the original author, AEON-7, and the upstream Qwen team. This repository only provides an MLX conversion and quantization for easier local use on Apple Silicon. No additional fine-tuning or intentional behavioral changes were made.

Quantization

  • Format: MLX
  • Quantization: 4-bit affine
  • Group size: 64
  • Approximate download size: 14 GB
  • Intended platform: Apple Silicon macOS
  • Recommended unified memory: 32 GB or more

Install

pip install -U mlx-lm

Run an interactive chat

mlx_lm.chat \
  --model choppedgarlic/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-4bit-MLX \
  --max-tokens 8192 \
  --temp 0.6 \
  --top-p 0.95

The first launch downloads the model from Hugging Face. Later launches use the local Hugging Face cache.

Run an OpenAI-compatible local server

Thinking enabled:

mlx_lm.server \
  --model choppedgarlic/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-4bit-MLX \
  --host 127.0.0.1 \
  --port 8081 \
  --max-tokens 8192 \
  --temp 0.6 \
  --top-p 0.95 \
  --chat-template-args '{"enable_thinking":true}'

Thinking disabled:

mlx_lm.server \
  --model choppedgarlic/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-4bit-MLX \
  --host 127.0.0.1 \
  --port 8081 \
  --max-tokens 8192 \
  --temp 0.6 \
  --top-p 0.95 \
  --chat-template-args '{"enable_thinking":false}'

The API base URL is then http://127.0.0.1:8081/v1.

Verified locally

  • Text generation
  • Interactive mlx_lm.chat
  • Thinking and non-thinking modes
  • mlx_lm.server through its OpenAI-compatible API
  • Local use through the Pi coding agent

Limitations

  • This conversion is intended for Apple Silicon and requires MLX.
  • The current conversion contains no vision_config or vision preprocessor assets. Image input and image generation are not supported or claimed.
  • Quantization may reduce quality compared with the BF16 source model.
  • The model can produce inaccurate, unsafe, or objectionable output. Validate outputs before using them in production or high-stakes settings.
  • Please also read the original model card for its intended use, behavior, and limitations.

License and attribution

Released under the Apache License 2.0, matching the source repository's declared license. See LICENSE.

Downloads last month
2,928
Safetensors
Model size
27B params
Tensor type
U32
ยท
BF16
ยท
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for choppedgarlic/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-4bit-MLX

Base model

Qwen/Qwen3.8-27B
Quantized
(33)
this model

Space using choppedgarlic/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-4bit-MLX 1