Instructions to use Quazim0t0/Byrne-86M-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Quazim0t0/Byrne-86M-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Quazim0t0/Byrne-86M-Base", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Quazim0t0/Byrne-86M-Base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Quazim0t0/Byrne-86M-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Quazim0t0/Byrne-86M-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-86M-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Quazim0t0/Byrne-86M-Base
- SGLang
How to use Quazim0t0/Byrne-86M-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Quazim0t0/Byrne-86M-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-86M-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Quazim0t0/Byrne-86M-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-86M-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Quazim0t0/Byrne-86M-Base with Docker Model Runner:
docker model run hf.co/Quazim0t0/Byrne-86M-Base
Byrne-86M-Base
Base of the Byrne family. Distilled step-4000 checkpoint. ~86M SpikeWhaleLM
from scratch - MLA, n-gram engram, hash-lookup, hyper-connections, HRM refine,
MTP. Custom ChatML-aware tokenizer. I use this as a general base to keep
pretraining / SFT.
Trained with Modal credits during the Small Models, Big Adventures Hackathon.
Related: chat β Byrne-86M
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Quazim0t0/Byrne-86M-Base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("Quazim0t0/Byrne-86M-Base", trust_remote_code=True)
Architecture
SpikeWhaleLM, ~86M, 16 layers, hidden 640, 4096 context, 16,512 vocab, tied embeddings.
- Multi-head Latent Attention (MLA + XSA) - Q and O LoRA-compressed (rank 128); each head splits RoPE dim 16 / NoPE dim 48; 10 query heads, one KV head (MQA); QK-norm.
- Engram n-gram memory - gated table, hashes local n-grams (up to trigrams) into 4,096 rows, mixes back into the residual.
- Hash-lookup layers (Γ2) - content-addressable features next to the token embeddings.
- Hyper-Connections - learned width-expanded residuals, Sinkhorn routing, instead of a plain add.
- HRM refinement - extra latent pass over hidden states before the output head.
- Multi-Token Prediction (MTP) - DeepSeek-V3-style extra head, more than one next token. Training only.
- FFN is dense. The block can do MoE; MoE is off in this release.
JEPA vs HRM. Byrne is Non-JEPA: HRM refine only (
use_hrm_refine=True,use_jepa=False). Escarda adds JEPA on top of HRM.
Tokenizer
SpikeTokenizer. Byte-level length-max (greedy longest-match), 16,512 vocab.
Not BPE. Text β UTF-8 β latin-1 bytes β longest vocab key that fits. ChatML-aware.
Atomic specials: <|im_start|>, <|im_end|>, <think>/</think>,
<begin_solution>/<end_solution>, tool-call markers, plus <bos>/<eos>/<pad>/<unk>.
PreTrainedTokenizer in spike_tokenizer.py. Load with
AutoTokenizer.from_pretrained(..., trust_remote_code=True).
Evaluation
Zero-shot multiple-choice, continuation log-likelihood (acc_norm =
byte-length-normalized).
| Task | acc | acc_norm |
|---|---|---|
| arc_easy | 0.4205 | 0.3931 |
| arc_challenge | 0.1877 | 0.2389 |
| hellaswag | 0.2792 | 0.2927 |
| winogrande | 0.5193 | - |
| piqa | 0.5941 | 0.5860 |
| openbookqa | 0.1420 | 0.2820 |
| boolq | 0.6171 | - |
ArithMark-2.0 (AxiomicLabs)
- official metric is raw
acc: 0.2732.
Language modeling: WikiText-2 byte_ppl (β) 2.3753 Β· BLiMP (β) 0.7356.
Citation
If you use this model, please cite:
@misc{byrne86mbase,
title = {Byrne-86M-Base: A ~86M-parameter SpikeWhaleLM},
author = {Dean Byrne (Quazim0t0)},
year = {2026},
howpublished = {HuggingFace, \url{https://huggingface.co/Quazim0t0/Byrne-86M-Base}},
note = {Quazim0t0/Byrne-86M-Base}
}
Update: engram repair (behavior-preserving)
The n-gram Engram in the original weights was degenerate: frozen LSH compressor at init scale hashed every token to bucket 0, so only one table row ever got gradient. This revision rescales the (frozen) compressor and broadcasts the learned bucket-0 vector across all table rows.
Outputs are bit-identical to the previous revision (verified: max logit difference 0.0 across a prompt battery). The only change: the Engram hash now spreads across the full table and every bucket is independently trainable - so if you distill or SFT on top of this base, the n-gram memory will actually learn instead of staying a constant bias.
- Downloads last month
- 719