GEV-26B-Decide, MLX 4-bit

This is autotrust/GEV-26B-Decide (System 1), converted to MLX for Apple silicon. It answers typed questions with a calibrated probability for every option, in one forward pass per question. A question can be yes/no, a pick from 2–256 options, or a rating from 0 to 5.

This is not a chat model. The System 1 LoRA is merged into the Gemma-4-26B-A4B backbone. mlx_lm.generate will load it, but the text it produces is not meaningful. Decisions come from the original 24-slot decision head, which the bundled gev_mlx.py runs.

This port is text only and System 1 only. The reference model's adaptive thinking (System 2, which is the unmodified Gemma 4 reasoning) is not included.

Variants

Variant Download Peak memory Median per decision (M1 Max) Same top answer as the bf16 reference
GEV-26B-Decide-mlx-8bit 25 GB 26.9 GB 0.18 s yes/no, 0.24 s 4 options 118/121
GEV-26B-Decide-mlx-4bit 13 GB 14.3 GB 0.19 s yes/no, 0.25 s 4 options 118/121

Timings are for short prompts on an M1 Max (64 GB). On that chip the two builds run at the same speed, so the 4-bit build is mainly for Macs with less memory. By default macOS lets the GPU wire only about 70–75% of RAM, so plan on a 48 GB Mac for the 8-bit build and a 24 GB Mac for the 4-bit build.

Usage

pip install "mlx==0.32.3" "mlx-lm==0.32.0" huggingface_hub   # no torch needed
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("Avicennasis/GEV-26B-Decide-mlx-4bit")
sys.path.insert(0, path)
import gev_mlx

model = gev_mlx.load(path)
model.decide("choice", "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
             "How much does the ball cost?", ["$0.10", "$0.05", "$1.00", "$0.55"])
# {'choice': '$0.05', 'probabilities': [0.001, 0.999, 0.0, 0.0], ...}
model.decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
             "Is the customer asking for a refund?")

decide() returns the reference /v1/decide response shape: options, probabilities, choice and choice_index. state can be a string or any JSON value. kind is one of:

  • noul: probabilities for ["false", "true"].
  • score: a rating from 0 to 5.
  • choice: 2–256 options. Above 16 options, the reference tournament reads groups of up to 16 and then a final of 16.

systemone(request) answers a Jev/SystemOne POST /v1/systemone body. GEV reads one question per pass, so each question is a separate decision:

  • Choice options are shown to the model as id: description, or as the id alone when there is no description.
  • A score question with exactly six levels uses the 0–5 slots. Any other score scale is read as a choice.

Calibration: gev_mlx.load(path, calibration="gold") selects calibration_gold.json. The reference card recommends that file when automatic actions are gated on confidence. The default, calibration.json, is what the reference Decision Index run used.

Command line and local server

cd "$(hf download Avicennasis/GEV-26B-Decide-mlx-4bit --quiet)"
python gev_mlx.py predict --kind choice --state "Checkout fails for every customer since the deploy." \
  --question "Which team?" --option billing --option engineering --option legal
python gev_mlx.py serve --port 8000     # binds 127.0.0.1 only by default

The server exposes these endpoints, and runs requests one at a time:

  • POST /v1/decide, with the reference request shape. thinking must be off.
  • POST /v1/systemone.
  • GET /health and GET /v1/models.

There is no authentication. Keep the default 127.0.0.1 binding unless you put your own proxy in front.

Parity with the reference implementation

The reference is the model card's own transformers + peft path: bf16, adapter merged in memory, head.safetensors, run under PyTorch MPS. Both the reference and these builds were scored on 121 typed decisions from a private evaluation set. The decisions come from shell-command guardrails and inbox triage, and are yes/no, 3-way and 19-way choices.

Same top answer as the reference Largest Δprob per decision, median / max Accuracy on that set
Reference (PyTorch bf16) n/a n/a 109/121
8-bit 118/121 0.006 / 0.113 110/121
4-bit 118/121 0.012 / 0.386 112/121

The three decisions that changed sat at 0.41–0.59 confidence on both sides, so they were coin flips. The accuracy differences are those flips, not a change in quality. On the same M1 Max the PyTorch MPS reference took a median of 4.7 s per evaluation record (2–4 decisions); the 8-bit build takes 0.43 s.

How it was converted

  1. The System 1 LoRA (adapter/) was merged into the bf16 backbone with peft's merge_and_unload (transformers 5.17, peft 0.21). The script checks that the adapter really loaded, that a targeted weight moved, and that an untargeted weight did not.
  2. The original config.json was restored. Current transformers re-saves Gemma 4's global-attention settings (global_head_dim, num_global_key_value_heads) as a per_layer_config block, which mlx-lm 0.32 does not read. The merge changes no shapes, so the original config is exact.
  3. The model was converted with mlx_lm.convert -q --q-bits 4 --q-group-size 64, affine (mlx-lm 0.32.0, mlx 0.32.3). The vision tower is dropped, since this build is text only, and the MoE routers stay at 8-bit.
  4. head.safetensors, judge_config.json, calibration.json and calibration_gold.json were copied unchanged.

The read-out matches the reference, and nothing is generated:

  1. The prompt is <bos>[kind] … [state] … [question] … [options] … [decision]:.
  2. The final-norm hidden state of the last token goes through the linear fp32 head (2,816 → 24 slots).
  3. The result is soft-capped at 30, sliced to the kind's slots, divided by the per-kind temperature, and softmaxed.

Limitations

  • System 1 only. There is no adaptive thinking.
  • Text only. The vision tower is not in this build.
  • One decision per forward pass. There is no batching, and multi-question requests run one question after another.
  • Custom code. gev_mlx.py is Python that you import and run. Read it before use if that matters in your environment.
  • The reference card's limitations still apply. In particular: it can be confidently wrong where its teacher was wrong, it is English-centric, and it is not for high-stakes decisions without confidence gating.

License

Apache-2.0. This follows autotrust/GEV-26B-Decide (adapter, head and calibration files) and google/gemma-4-26B-A4B-it (Gemma 4 is released under Apache 2.0). This port is not affiliated with autotrust, AutoTrust AI, TypeSafe AI or Google.

Downloads last month
138
Safetensors
Model size
25B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Avicennasis/GEV-26B-Decide-mlx-4bit

Quantized
(7)
this model