Instructions to use Avicennasis/GEV-26B-Decide-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Avicennasis/GEV-26B-Decide-mlx-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Avicennasis/GEV-26B-Decide-mlx-4bit --local-dir GEV-26B-Decide-mlx-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
GEV-26B-Decide, MLX 4-bit
This is autotrust/GEV-26B-Decide (System 1), converted to MLX for Apple silicon. It answers typed questions with a calibrated probability for every option, in one forward pass per question. A question can be yes/no, a pick from 2–256 options, or a rating from 0 to 5.
This is not a chat model. The System 1 LoRA is merged into the Gemma-4-26B-A4B backbone. mlx_lm.generate will
load it, but the text it produces is not meaningful. Decisions come from the original 24-slot decision head, which the
bundled gev_mlx.py runs.
This port is text only and System 1 only. The reference model's adaptive thinking (System 2, which is the unmodified Gemma 4 reasoning) is not included.
Variants
| Variant | Download | Peak memory | Median per decision (M1 Max) | Same top answer as the bf16 reference |
|---|---|---|---|---|
| GEV-26B-Decide-mlx-8bit | 25 GB | 26.9 GB | 0.18 s yes/no, 0.24 s 4 options | 118/121 |
| GEV-26B-Decide-mlx-4bit | 13 GB | 14.3 GB | 0.19 s yes/no, 0.25 s 4 options | 118/121 |
Timings are for short prompts on an M1 Max (64 GB). On that chip the two builds run at the same speed, so the 4-bit build is mainly for Macs with less memory. By default macOS lets the GPU wire only about 70–75% of RAM, so plan on a 48 GB Mac for the 8-bit build and a 24 GB Mac for the 4-bit build.
Usage
pip install "mlx==0.32.3" "mlx-lm==0.32.0" huggingface_hub # no torch needed
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("Avicennasis/GEV-26B-Decide-mlx-4bit")
sys.path.insert(0, path)
import gev_mlx
model = gev_mlx.load(path)
model.decide("choice", "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
"How much does the ball cost?", ["$0.10", "$0.05", "$1.00", "$0.55"])
# {'choice': '$0.05', 'probabilities': [0.001, 0.999, 0.0, 0.0], ...}
model.decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
"Is the customer asking for a refund?")
decide() returns the reference /v1/decide response shape: options, probabilities, choice and choice_index.
state can be a string or any JSON value. kind is one of:
noul: probabilities for["false", "true"].score: a rating from 0 to 5.choice: 2–256 options. Above 16 options, the reference tournament reads groups of up to 16 and then a final of 16.
systemone(request) answers a Jev/SystemOne POST /v1/systemone body. GEV reads one question per pass, so each question
is a separate decision:
- Choice options are shown to the model as
id: description, or as the id alone when there is no description. - A score question with exactly six levels uses the 0–5 slots. Any other score scale is read as a choice.
Calibration: gev_mlx.load(path, calibration="gold") selects calibration_gold.json. The reference card recommends
that file when automatic actions are gated on confidence. The default, calibration.json, is what the reference
Decision Index run used.
Command line and local server
cd "$(hf download Avicennasis/GEV-26B-Decide-mlx-4bit --quiet)"
python gev_mlx.py predict --kind choice --state "Checkout fails for every customer since the deploy." \
--question "Which team?" --option billing --option engineering --option legal
python gev_mlx.py serve --port 8000 # binds 127.0.0.1 only by default
The server exposes these endpoints, and runs requests one at a time:
POST /v1/decide, with the reference request shape.thinkingmust be off.POST /v1/systemone.GET /healthandGET /v1/models.
There is no authentication. Keep the default 127.0.0.1 binding unless you put your own proxy in front.
Parity with the reference implementation
The reference is the model card's own transformers + peft path: bf16, adapter merged in memory, head.safetensors,
run under PyTorch MPS. Both the reference and these builds were scored on 121 typed decisions from a private evaluation
set. The decisions come from shell-command guardrails and inbox triage, and are yes/no, 3-way and 19-way choices.
| Same top answer as the reference | Largest Δprob per decision, median / max | Accuracy on that set | |
|---|---|---|---|
| Reference (PyTorch bf16) | n/a | n/a | 109/121 |
| 8-bit | 118/121 | 0.006 / 0.113 | 110/121 |
| 4-bit | 118/121 | 0.012 / 0.386 | 112/121 |
The three decisions that changed sat at 0.41–0.59 confidence on both sides, so they were coin flips. The accuracy differences are those flips, not a change in quality. On the same M1 Max the PyTorch MPS reference took a median of 4.7 s per evaluation record (2–4 decisions); the 8-bit build takes 0.43 s.
How it was converted
- The System 1 LoRA (
adapter/) was merged into the bf16 backbone with peft'smerge_and_unload(transformers 5.17, peft 0.21). The script checks that the adapter really loaded, that a targeted weight moved, and that an untargeted weight did not. - The original
config.jsonwas restored. Current transformers re-saves Gemma 4's global-attention settings (global_head_dim,num_global_key_value_heads) as aper_layer_configblock, which mlx-lm 0.32 does not read. The merge changes no shapes, so the original config is exact. - The model was converted with
mlx_lm.convert -q --q-bits 4 --q-group-size 64, affine (mlx-lm 0.32.0, mlx 0.32.3). The vision tower is dropped, since this build is text only, and the MoE routers stay at 8-bit. head.safetensors,judge_config.json,calibration.jsonandcalibration_gold.jsonwere copied unchanged.
The read-out matches the reference, and nothing is generated:
- The prompt is
<bos>[kind] … [state] … [question] … [options] … [decision]:. - The final-norm hidden state of the last token goes through the linear fp32 head (2,816 → 24 slots).
- The result is soft-capped at 30, sliced to the kind's slots, divided by the per-kind temperature, and softmaxed.
Limitations
- System 1 only. There is no adaptive thinking.
- Text only. The vision tower is not in this build.
- One decision per forward pass. There is no batching, and multi-question requests run one question after another.
- Custom code.
gev_mlx.pyis Python that you import and run. Read it before use if that matters in your environment. - The reference card's limitations still apply. In particular: it can be confidently wrong where its teacher was wrong, it is English-centric, and it is not for high-stakes decisions without confidence gating.
License
Apache-2.0. This follows autotrust/GEV-26B-Decide (adapter, head and calibration files) and google/gemma-4-26B-A4B-it (Gemma 4 is released under Apache 2.0). This port is not affiliated with autotrust, AutoTrust AI, TypeSafe AI or Google.
- Downloads last month
- 138
4-bit
Model tree for Avicennasis/GEV-26B-Decide-mlx-4bit
Base model
google/gemma-4-26B-A4B