Why only 6bit? Protection floors (MTP + vision BF16, etc.) force both the 4.8 and ~6.0 BPW budgets to **6.97 BPW** with identical weights. There is no smaller AXQ-4bit sibling for Qwen3.5-9B — use this pack only.

AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

Development evidence — not a certified AXQuant release. This package has conversion and artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed, or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim.

Stable-name v2. main serves the audited v2 artifact for backward compatibility. The same revision is tagged v2; the replaced artifact remains recoverable at legacy-pre-v2.

Model details

Property Value
Base model Qwen/Qwen3.5-9B
Source revision c202236235762e1c871ad0ccb60c8ee5ba337b9a
Product family qwen3.5
Source architecture Qwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters 9.41B logical parameters
Quantizer AXQuant 1.2.0
Hub budget class 6bit
Artifact edition v2
AXQuant base precision class 7p0bpw
Planned storage-adjusted BPW 6.9700
Measured main-model BPW 6.7367
Measured total BPW, including MTP 6.9701
Safetensors weight size 8.41 GB
Approximate complete download 8.43 GB
Configured maximum context 262,144 tokens; practical limits depend on unified memory
MLX-LM compatibility Standard text inference, compatibility level B
AX Engine native execution Not established; no validated native manifest is included
MTP present True
Vision sidecar present True

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Why there is no AXQ-4bit pack

On this base, AXQuant protection floors (BF16 MTP sidecar, BF16 vision tower, and other protected tensors) raise both the low-memory (~4.8 BPW) and 6 BPW planning budgets to the same effective target of about 6.97 BPW.

The former …-AXQ-4bit-MTP sibling therefore contained the same mixed plan and the same weight files as this 6bit pack (~8.4 GB download). Publishing both would look like a size trade-off when none exists. AutomatosX keeps only this repository.

Measured main-model BPW ~6.74
Measured total BPW (incl. MTP) ~6.97
Why not 4bit Floor-collapsed; identical to this pack

Choosing an AXQ pack

AXQ 4bit / 6bit names are storage-budget product classes, not a promise that every tensor uses that width. On this base there is no distinct 4bit Hub pack — see Why there is no AXQ-4bit pack above.

Sibling Intended trade-off
(none published) This 6bit pack is the only public AXQ checkpoint for this base.

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP --local-dir ./AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP

Allow at least 8.43 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine status

This package does not include a validated native model-manifest.json, so AX Engine execution is not established by this release. The AX Engine fields in axquant_runtime.json describe the intended compatibility contract, not observed runtime evidence. Use the MLX-LM path above for standard text/backbone inference. The artifact records AX Engine version not recorded, but version discovery alone is not a runtime check.

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precision Parameters Share
4bit 6.92B 71.67%
8bit 1.02B 10.54%
bf16 1.72B 17.79%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 243.29M parameters, 0.49 GB, BF16.
  • Vision sidecar: 333 tensors, 456.01M parameters, 0.91 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

Check Status
Planning evidence architecture_prior
Calibration none; the allocation is based on architecture priors
Quantizer execution 249/249 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifest not included
Quality versus BF16 or uniform baselines Not published; no quality-retention claim
MTP acceptance and speed not measured; no MTP speedup claim
AX Engine kernel evidence unmeasured
Vision-language quality Not evaluated or claimed; vision tensors are preserved at BF16
Long-context quality 262,144-token capacity is config metadata, not a validated claim
Release certification Not certified; formal AXQuant M0-M8 gates are not closed

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish exactness or speed certification for those runtimes.

  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • AX Engine execution is not established because this package has no validated native manifest.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.5-9B model card for license terms, model limitations, and responsible-use guidance.

Downloads last month
307
Safetensors
Model size
9B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP

Finetuned
Qwen/Qwen3.5-9B
Quantized
(500)
this model

Collection including AutomatosX/AX-Qwen3.5-9B-MLX-AXQ-6bit-MTP