Text Generation
Transformers
TensorBoard
Safetensors
English
alicet5_moe
text2text-generation
pretrained
from-scratch
tiny-llm-ablation
custom_code
ul2
Mixture of Experts
encoder-decoder
Eval Results (legacy)
Instructions to use d0rj/t5-moe-55M-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use d0rj/t5-moe-55M-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="d0rj/t5-moe-55M-base", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("d0rj/t5-moe-55M-base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use d0rj/t5-moe-55M-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "d0rj/t5-moe-55M-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/d0rj/t5-moe-55M-base
- SGLang
How to use d0rj/t5-moe-55M-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "d0rj/t5-moe-55M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "d0rj/t5-moe-55M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use d0rj/t5-moe-55M-base with Docker Model Runner:
docker model run hf.co/d0rj/t5-moe-55M-base
File size: 18,461 Bytes
322d767 1e82aa5 322d767 d2d2f87 322d767 de3885b 322d767 de3885b 322d767 de3885b 322d767 d2d2f87 de3885b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 | ---
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- pretrained
- from-scratch
- tiny-llm-ablation
- custom_code
- tensorboard
- ul2
- moe
- encoder-decoder
datasets:
- HuggingFaceFW/fineweb-edu
model-index:
- name: t5-moe-55M-base
results:
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: Rowan/hellaswag
name: HellaSwag
config: default
split: validation
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.2779326827325234
args:
ci95_low: 0.2692569969516344
ci95_high: 0.28677820246076313
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/ai2_arc
name: ARC-Easy
config: ARC-Easy
split: test
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.38930976430976433
args:
ci95_low: 0.3698977221380954
ci95_high: 0.40907915127907174
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/ai2_arc
name: ARC-Challenge
config: ARC-Challenge
split: test
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.2380546075085324
args:
ci95_low: 0.21455235182352161
ci95_high: 0.26326840760453163
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: baber/piqa
name: PIQA
config: default
split: validation
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.5560391730141458
args:
ci95_low: 0.5332313404997734
ci95_high: 0.5786132479763156
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/winogrande
name: WinoGrande
config: winogrande_xl
split: validation
args:
num_few_shot: 0
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12)
value: 0.4846093133385951
args:
ci95_low: 0.4571789703267686
ci95_high: 0.512132701298903
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/openbookqa
name: OpenBookQA
config: main
split: test
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.262
args:
ci95_low: 0.22537626941723765
ci95_high: 0.30225291664246134
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: aps/super_glue
name: BoolQ
config: boolq
split: validation
args:
num_few_shot: 0
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12)
value: 0.499388379204893
args:
ci95_low: 0.48226179513749795
ci95_high: 0.5165163985990305
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: EleutherAI/lambada_openai
name: LAMBADA OpenAI
config: default
split: test
args:
num_few_shot: 0
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12)
value: 0.14826314768096255
args:
ci95_low: 0.13882265856941103
ci95_high: 0.1582276717641896
ci_method: Wilson
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: AxiomicLabs/Arithmark-3.0
name: ArithMark-3
config: default
split: train
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.342
args:
dtype: bfloat16
num_few_shot: 0
max_length: 1024
standard_error: 0.015008706182121804
evaluation_date: '2026-10-01'
ci95_low: 0.3132528988050859
ci95_high: 0.37195635687634954
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: pkavumba/balanced-copa
name: Balanced COPA
config: default
split: train
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.538
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.015773547629015002
evaluation_date: '2026-10-01'
ci95_low: 0.5070132970793998
ci95_high: 0.5686958692756982
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: tau/commonsense_qa
name: CommonsenseQA
config: default
split: validation
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.20065520065520065
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.011466011466011467
evaluation_date: '2026-10-01'
ci95_low: 0.17914588148536167
ci95_high: 0.2240421844184298
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: allenai/sciq
name: SciQ (with support)
config: default
split: test
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.654
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.01505026612756434
evaluation_date: '2026-10-01'
ci95_low: 0.6239780184885133
ci95_high: 0.6828433398979359
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: truthfulqa/truthful_qa
name: TruthfulQA MC2
config: multiple_choice
split: validation
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.4656644987453161
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.016000577772697092
evaluation_date: '2026-10-01'
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: BananaMind/BananaMind-Base-Bench-1.1
name: BananaMind Base 1.1
config: default
split: test
metrics:
- type: raw_accuracy
name: raw_accuracy (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.36
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.025693810923465083
evaluation_date: '2026-10-01'
ci95_low: 0.3114835744745811
ci95_high: 0.41155622892601307
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: cais/mmlu
name: MMLU continuation
config: 57 subjects
split: test
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.24932345819683804
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.003640204268569072
evaluation_date: '2026-10-01'
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: nyu-mll/blimp
name: BLiMP
config: 67 minimal-pair subsets
split: train
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.6988955223880597
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.0015386258071003309
evaluation_date: '2026-10-01'
---
# T5 MoE 55M Base (UL2)
A **54,858,240-parameter** English base model in the **Tiny llm ablation** experiment. Trained from scratch on exactly **3,932,160,000 source tokens** over **15,000 optimizer steps**. The token count measures processed input blocks, not unique text or supervised target tokens.
## Architecture and references
6 encoder + 6 decoder layers, width 512; encoder 8-head attention, decoder 8 query / 2 KV heads; 8 experts per MoE layer, top-2 routing, expert width 160; RoPE, RMSNorm, FP32 residuals and tied shared embeddings. Maximum encoder length 2050 including controls; raw training blocks 2048.
Architecture inspiration: [yandex/AliceAI-T5-35B-A0.6B](https://huggingface.co/yandex/AliceAI-T5-35B-A0.6B). Tokenizer foundation: [q-project/Q-50M-Base](https://huggingface.co/q-project/Q-50M-Base), preserving all 32,768 original IDs and adding 3 mode tokens + 512 sentinels (33,283 entries). Weights were initialized randomly. This small adaptation does not reproduce Alice’s corpus, optimizer or routing recipe.
## Training
- Data: [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), `sample-10BT`, streamed from local Parquet shards; shuffle buffer 100,000.
- Objective: UL2 with seven equally likely denoisers: R(15%, mean span 3/8), S(suffix), X(50%,3), X(50%,8), X(15%,64), X(50%,64). S masks a uniformly sampled suffix of length 1..L/2. Targets contain corrupted spans and control tokens. Training adds router auxiliary loss with coefficient 0.01. These sampler choices are explicit local choices; see [UL2](https://arxiv.org/abs/2205.05131).
- Batch: 16 sequences × 8 accumulation × 2048 tokens = 262,144 source tokens per step.
- Fused AdamW; peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 (no decay for bias/norm/1D parameters), gradient clipping 1.0. Linear warmup for 150 steps, then cosine decay to 10% of peak LR.
- BF16 compute on one RTX 5070 Ti (16 GB), seed 2026; checkpoints retain FP32 weights. [Exact training configuration](training_config.json).
## Evaluation
Full selected task splits, **no added few-shot examples**, **lm-eval 0.4.12**, no chat template, BF16 on RTX 5070 Ti, maximum context 2048 (ArithMark: 1024). Accuracy is a percentage. **± is one standard error; the separate bracketed column is the 95% Wilson confidence interval.** Intervals describe finite evaluation-sample uncertainty, not variation across training seeds; no multiple-comparison correction is applied.
| Dataset | Split | Examples | Metric | Score ± SE (%) | 95% CI (%) |
|---|---|---:|---|---:|---:|
| [HellaSwag](https://huggingface.co/datasets/Rowan/hellaswag) | validation | 10,042 | `acc_norm` | 27.79 ± 0.45 | [26.93, 28.68] |
| [ARC-Easy](https://huggingface.co/datasets/allenai/ai2_arc) | test | 2,376 | `acc_norm` | 38.93 ± 1.00 | [36.99, 40.91] |
| [ARC-Challenge](https://huggingface.co/datasets/allenai/ai2_arc) | test | 1,172 | `acc_norm` | 23.81 ± 1.24 | [21.46, 26.33] |
| [PIQA](https://huggingface.co/datasets/baber/piqa) | validation | 1,838 | `acc_norm` | 55.60 ± 1.16 | [53.32, 57.86] |
| [WinoGrande](https://huggingface.co/datasets/allenai/winogrande) | validation | 1,267 | `acc` | 48.46 ± 1.40 | [45.72, 51.21] |
| [OpenBookQA](https://huggingface.co/datasets/allenai/openbookqa) | test | 500 | `acc_norm` | 26.20 ± 1.97 | [22.54, 30.23] |
| [BoolQ](https://huggingface.co/datasets/aps/super_glue) | validation | 3,270 | `acc` | 49.94 ± 0.87 | [48.23, 51.65] |
| [LAMBADA OpenAI](https://huggingface.co/datasets/EleutherAI/lambada_openai) | test | 5,153 | `acc` | 14.83 ± 0.50 | [13.88, 15.82] |
| [ArithMark-3](https://huggingface.co/datasets/AxiomicLabs/Arithmark-3.0) | train | 1,000 | `acc_norm` | 34.20 ± 1.50 | [31.33, 37.20] |
| [Balanced COPA](https://huggingface.co/datasets/pkavumba/balanced-copa) | train | 1,000 | `acc` | 53.80 ± 1.58 | [50.70, 56.87] |
| [CommonsenseQA](https://huggingface.co/datasets/tau/commonsense_qa) | validation | 1,221 | `acc` | 20.07 ± 1.15 | [17.91, 22.40] |
| [SciQ (with support)](https://huggingface.co/datasets/allenai/sciq) | test | 1,000 | `acc_norm` | 65.40 ± 1.51 | [62.40, 68.28] |
| [TruthfulQA MC2](https://huggingface.co/datasets/truthfulqa/truthful_qa) | validation | 817 | `acc` | 46.57 ± 1.60 | — |
| [BananaMind Base 1.1](https://huggingface.co/datasets/BananaMind/BananaMind-Base-Bench-1.1) | test | 350 | `raw_accuracy` | 36.00 ± 2.57 | [31.15, 41.16] |
| [MMLU continuation](https://huggingface.co/datasets/cais/mmlu) | test | 14,042 | `acc` | 24.93 ± 0.36 | — |
| [BLiMP](https://huggingface.co/datasets/nyu-mll/blimp) | train | 67,000 | `acc` | 69.89 ± 0.15 | — |
T5 uses UL2 S-mode: encoder S + prefix + sentinel + EOS; decoder BOS + sentinel + shifted answer. Only answer text is scored; control tokens and router loss are excluded, with the full vocabulary retained in the softmax. Its encoder sees at most 2047 text-prefix tokens after reserving controls. LAMBADA accuracy requires the complete final-word token sequence. `acc_norm` is harness length-normalized option scoring; raw accuracy is also stored in [results.json](evaluation/results.json).
**[WikiText-2 raw test](https://huggingface.co/datasets/Salesforce/wikitext), conditional continuation:** CPU FP32 re-evaluation on 291 nonoverlapping blocks (512 prefix + 512 scored suffix tokens), 148,992 scored tokens; 335 tail tokens excluded. NLL **3.920357**, 95% CI **[3.879299, 3.960742]**; token PPL **50.418**, 95% CI **[48.390, 52.496]**. Percentile block bootstrap, 10,000 resamples, seed 2026; exponentiate NLL endpoints for PPL. Blocks are the resampling unit; this does not model all within-document dependence. This is not standard rolling AR or word PPL. The earlier BF16 point is retained separately in TensorBoard, with no borrowed FP32 interval.
The metadata contains author-reported `model-index` scores. The evaluated dataset repositories had no registered `eval.yaml` on 2026-09-20, so no `.eval_results` leaderboard entry or verified badge is claimed. [Machine-readable results and provenance](evaluation/results.json).
Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti;
context cap 2048 (ArithMark 1024), TF32 disabled, no chat template and no added
few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble.
ArithMark and BananaMind normalize by continuation token count; ordinary harness
acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo.
SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item
train-named evaluation split; cRia's split was inferred, not confirmed.
MMLU scores full answer continuations across 57 subjects, weighted by item count;
BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained
from each evaluator. Wilson intervals are reported only where the runner logged
binary item accuracy; MC2 is probability mass, not binary accuracy. These
intervals do not model dependence between paired/templated examples or training
seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented
conditional scoring protocols; PLL exposes the other answer tokens and is not
autoregressive likelihood. cRia's published scores used a different precision
and benchmark-adapted checkpoint; this completes our comparison coverage, not
an independent reproduction of cRia or an official leaderboard submission.
[Full results, provenance and group scores](evaluation/comparison-20261001/results.json). [Updated machine-readable results](evaluation/results.json). [TensorBoard events](tensorboard/) contain these new scores at step 15,000.
## Usage
Install `requirements.txt` (tested with Transformers 5.17.0 / PyTorch 2.11.0). Custom model code is included; `trust_remote_code=True` is required. This example runs on CPU.
```python
import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
repo = "d0rj/t5-moe-55M-base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
c = model.config.ul2
prefix = tokenizer.encode("The capital of France is", add_special_tokens=False)
inputs = torch.tensor([[c["mode_ids"]["S"], *prefix, c["sentinel_ids"][0], c["eos_id"]]])
decoder = torch.tensor([[model.config.decoder_start_token_id, c["sentinel_ids"][0]]])
output = model.generate(input_ids=inputs, decoder_input_ids=decoder,
max_new_tokens=64, do_sample=False)
print(tokenizer.decode(output[0, decoder.shape[1]:], skip_special_tokens=True))
```
To reproduce the core evaluation from a downloaded repository, install `evaluation/requirements.txt` and run:
```bash
python evaluation/run_core.py --device cuda:0 --dtype bfloat16 --batch-size 16 --output evaluation-rerun
```
To reproduce after downloading this model repository, accept the BananaMind dataset terms, authenticate with `hf auth login`, then run in a suitable CUDA environment:
```bash
pip install -r evaluation/comparison-20261001/repro/requirements.txt
python evaluation/comparison-20261001/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output comparison-rerun
```
The bundled runner uses the published model classes with the exact evaluation adapters and tokenizer. `--limit` produces smoke results only. Raw dataset examples are not included in this release.
## TensorBoard and limitations
[TensorBoard event files](tensorboard/) include training telemetry and `eval/<task>/<metric>` at step 15,000, plus separate CI bounds. Training telemetry covers steps 20–15,000 (750 loss points), including token CE, router loss, gradient norm, throughput, memory, padding and denoiser fractions.
These are small English continuation models, not instruction-tuned assistants. Equal source-token budgets do not imply equal target-token supervision or FLOPs. Benchmark contamination was not audited; results are from one training seed. Reference-model scores from different prompts, tokenizers or corpora are not directly interchangeable.
|