Instructions to use d0rj/t5-moe-55M-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use d0rj/t5-moe-55M-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="d0rj/t5-moe-55M-base", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("d0rj/t5-moe-55M-base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use d0rj/t5-moe-55M-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "d0rj/t5-moe-55M-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/d0rj/t5-moe-55M-base
- SGLang
How to use d0rj/t5-moe-55M-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "d0rj/t5-moe-55M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "d0rj/t5-moe-55M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use d0rj/t5-moe-55M-base with Docker Model Runner:
docker model run hf.co/d0rj/t5-moe-55M-base
Download README.md from d0rj/t5-moe-55M-base: direct link, hf CLI and curl.
- Browser
- Download file 18.5 kB
-
https://huggingface.co/d0rj/t5-moe-55M-base/resolve/main/README.md
- Command line
-
hf download hf://d0rj/t5-moe-55M-base/README.md
-
curl -L -o README.md https://huggingface.co/d0rj/t5-moe-55M-base/resolve/main/README.md
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- pretrained
- from-scratch
- tiny-llm-ablation
- custom_code
- tensorboard
- ul2
- moe
- encoder-decoder
datasets:
- HuggingFaceFW/fineweb-edu
model-index:
- name: t5-moe-55M-base
results:
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: Rowan/hellaswag
name: HellaSwag
config: default
split: validation
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.2779326827325234
args:
ci95_low: 0.2692569969516344
ci95_high: 0.28677820246076313
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/ai2_arc
name: ARC-Easy
config: ARC-Easy
split: test
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.38930976430976433
args:
ci95_low: 0.3698977221380954
ci95_high: 0.40907915127907174
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/ai2_arc
name: ARC-Challenge
config: ARC-Challenge
split: test
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.2380546075085324
args:
ci95_low: 0.21455235182352161
ci95_high: 0.26326840760453163
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: baber/piqa
name: PIQA
config: default
split: validation
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.5560391730141458
args:
ci95_low: 0.5332313404997734
ci95_high: 0.5786132479763156
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/winogrande
name: WinoGrande
config: winogrande_xl
split: validation
args:
num_few_shot: 0
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12)
value: 0.4846093133385951
args:
ci95_low: 0.4571789703267686
ci95_high: 0.512132701298903
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: allenai/openbookqa
name: OpenBookQA
config: main
split: test
args:
num_few_shot: 0
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12)
value: 0.262
args:
ci95_low: 0.22537626941723765
ci95_high: 0.30225291664246134
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: aps/super_glue
name: BoolQ
config: boolq
split: validation
args:
num_few_shot: 0
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12)
value: 0.499388379204893
args:
ci95_low: 0.48226179513749795
ci95_high: 0.5165163985990305
ci_method: Wilson
- task:
type: text-generation
name: Zero-shot continuation likelihood
dataset:
type: EleutherAI/lambada_openai
name: LAMBADA OpenAI
config: default
split: test
args:
num_few_shot: 0
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12)
value: 0.14826314768096255
args:
ci95_low: 0.13882265856941103
ci95_high: 0.1582276717641896
ci_method: Wilson
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: AxiomicLabs/Arithmark-3.0
name: ArithMark-3
config: default
split: train
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.342
args:
dtype: bfloat16
num_few_shot: 0
max_length: 1024
standard_error: 0.015008706182121804
evaluation_date: '2026-10-01'
ci95_low: 0.3132528988050859
ci95_high: 0.37195635687634954
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: pkavumba/balanced-copa
name: Balanced COPA
config: default
split: train
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.538
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.015773547629015002
evaluation_date: '2026-10-01'
ci95_low: 0.5070132970793998
ci95_high: 0.5686958692756982
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: tau/commonsense_qa
name: CommonsenseQA
config: default
split: validation
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.20065520065520065
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.011466011466011467
evaluation_date: '2026-10-01'
ci95_low: 0.17914588148536167
ci95_high: 0.2240421844184298
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: allenai/sciq
name: SciQ (with support)
config: default
split: test
metrics:
- type: acc_norm
name: acc_norm (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.654
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.01505026612756434
evaluation_date: '2026-10-01'
ci95_low: 0.6239780184885133
ci95_high: 0.6828433398979359
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: truthfulqa/truthful_qa
name: TruthfulQA MC2
config: multiple_choice
split: validation
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.4656644987453161
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.016000577772697092
evaluation_date: '2026-10-01'
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: BananaMind/BananaMind-Base-Bench-1.1
name: BananaMind Base 1.1
config: default
split: test
metrics:
- type: raw_accuracy
name: raw_accuracy (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.36
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.025693810923465083
evaluation_date: '2026-10-01'
ci95_low: 0.3114835744745811
ci95_high: 0.41155622892601307
ci_method: Wilson 95%; item independence approximation
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: cais/mmlu
name: MMLU continuation
config: 57 subjects
split: test
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.24932345819683804
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.003640204268569072
evaluation_date: '2026-10-01'
- task:
type: text-generation
name: Continuation likelihood
dataset:
type: nyu-mll/blimp
name: BLiMP
config: 67 minimal-pair subsets
split: train
metrics:
- type: acc
name: acc (fraction; lm-eval 0.4.12 comparison protocol)
value: 0.6988955223880597
args:
dtype: bfloat16
num_few_shot: 0
max_length: 2048
standard_error: 0.0015386258071003309
evaluation_date: '2026-10-01'
T5 MoE 55M Base (UL2)
A 54,858,240-parameter English base model in the Tiny llm ablation experiment. Trained from scratch on exactly 3,932,160,000 source tokens over 15,000 optimizer steps. The token count measures processed input blocks, not unique text or supervised target tokens.
Architecture and references
6 encoder + 6 decoder layers, width 512; encoder 8-head attention, decoder 8 query / 2 KV heads; 8 experts per MoE layer, top-2 routing, expert width 160; RoPE, RMSNorm, FP32 residuals and tied shared embeddings. Maximum encoder length 2050 including controls; raw training blocks 2048.
Architecture inspiration: yandex/AliceAI-T5-35B-A0.6B. Tokenizer foundation: q-project/Q-50M-Base, preserving all 32,768 original IDs and adding 3 mode tokens + 512 sentinels (33,283 entries). Weights were initialized randomly. This small adaptation does not reproduce Alice’s corpus, optimizer or routing recipe.
Training
- Data: FineWeb-Edu,
sample-10BT, streamed from local Parquet shards; shuffle buffer 100,000. - Objective: UL2 with seven equally likely denoisers: R(15%, mean span 3/8), S(suffix), X(50%,3), X(50%,8), X(15%,64), X(50%,64). S masks a uniformly sampled suffix of length 1..L/2. Targets contain corrupted spans and control tokens. Training adds router auxiliary loss with coefficient 0.01. These sampler choices are explicit local choices; see UL2.
- Batch: 16 sequences × 8 accumulation × 2048 tokens = 262,144 source tokens per step.
- Fused AdamW; peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 (no decay for bias/norm/1D parameters), gradient clipping 1.0. Linear warmup for 150 steps, then cosine decay to 10% of peak LR.
- BF16 compute on one RTX 5070 Ti (16 GB), seed 2026; checkpoints retain FP32 weights. Exact training configuration.
Evaluation
Full selected task splits, no added few-shot examples, lm-eval 0.4.12, no chat template, BF16 on RTX 5070 Ti, maximum context 2048 (ArithMark: 1024). Accuracy is a percentage. ± is one standard error; the separate bracketed column is the 95% Wilson confidence interval. Intervals describe finite evaluation-sample uncertainty, not variation across training seeds; no multiple-comparison correction is applied.
| Dataset | Split | Examples | Metric | Score ± SE (%) | 95% CI (%) |
|---|---|---|---|---|---|
| HellaSwag | validation | 10,042 | acc_norm |
27.79 ± 0.45 | [26.93, 28.68] |
| ARC-Easy | test | 2,376 | acc_norm |
38.93 ± 1.00 | [36.99, 40.91] |
| ARC-Challenge | test | 1,172 | acc_norm |
23.81 ± 1.24 | [21.46, 26.33] |
| PIQA | validation | 1,838 | acc_norm |
55.60 ± 1.16 | [53.32, 57.86] |
| WinoGrande | validation | 1,267 | acc |
48.46 ± 1.40 | [45.72, 51.21] |
| OpenBookQA | test | 500 | acc_norm |
26.20 ± 1.97 | [22.54, 30.23] |
| BoolQ | validation | 3,270 | acc |
49.94 ± 0.87 | [48.23, 51.65] |
| LAMBADA OpenAI | test | 5,153 | acc |
14.83 ± 0.50 | [13.88, 15.82] |
| ArithMark-3 | train | 1,000 | acc_norm |
34.20 ± 1.50 | [31.33, 37.20] |
| Balanced COPA | train | 1,000 | acc |
53.80 ± 1.58 | [50.70, 56.87] |
| CommonsenseQA | validation | 1,221 | acc |
20.07 ± 1.15 | [17.91, 22.40] |
| SciQ (with support) | test | 1,000 | acc_norm |
65.40 ± 1.51 | [62.40, 68.28] |
| TruthfulQA MC2 | validation | 817 | acc |
46.57 ± 1.60 | — |
| BananaMind Base 1.1 | test | 350 | raw_accuracy |
36.00 ± 2.57 | [31.15, 41.16] |
| MMLU continuation | test | 14,042 | acc |
24.93 ± 0.36 | — |
| BLiMP | train | 67,000 | acc |
69.89 ± 0.15 | — |
T5 uses UL2 S-mode: encoder S + prefix + sentinel + EOS; decoder BOS + sentinel + shifted answer. Only answer text is scored; control tokens and router loss are excluded, with the full vocabulary retained in the softmax. Its encoder sees at most 2047 text-prefix tokens after reserving controls. LAMBADA accuracy requires the complete final-word token sequence. acc_norm is harness length-normalized option scoring; raw accuracy is also stored in results.json.
WikiText-2 raw test, conditional continuation: CPU FP32 re-evaluation on 291 nonoverlapping blocks (512 prefix + 512 scored suffix tokens), 148,992 scored tokens; 335 tail tokens excluded. NLL 3.920357, 95% CI [3.879299, 3.960742]; token PPL 50.418, 95% CI [48.390, 52.496]. Percentile block bootstrap, 10,000 resamples, seed 2026; exponentiate NLL endpoints for PPL. Blocks are the resampling unit; this does not model all within-document dependence. This is not standard rolling AR or word PPL. The earlier BF16 point is retained separately in TensorBoard, with no borrowed FP32 interval.
The metadata contains author-reported model-index scores. The evaluated dataset repositories had no registered eval.yaml on 2026-09-20, so no .eval_results leaderboard entry or verified badge is claimed. Machine-readable results and provenance.
Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled, no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.
Full results, provenance and group scores. Updated machine-readable results. TensorBoard events contain these new scores at step 15,000.
Usage
Install requirements.txt (tested with Transformers 5.17.0 / PyTorch 2.11.0). Custom model code is included; trust_remote_code=True is required. This example runs on CPU.
import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
repo = "d0rj/t5-moe-55M-base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
c = model.config.ul2
prefix = tokenizer.encode("The capital of France is", add_special_tokens=False)
inputs = torch.tensor([[c["mode_ids"]["S"], *prefix, c["sentinel_ids"][0], c["eos_id"]]])
decoder = torch.tensor([[model.config.decoder_start_token_id, c["sentinel_ids"][0]]])
output = model.generate(input_ids=inputs, decoder_input_ids=decoder,
max_new_tokens=64, do_sample=False)
print(tokenizer.decode(output[0, decoder.shape[1]:], skip_special_tokens=True))
To reproduce the core evaluation from a downloaded repository, install evaluation/requirements.txt and run:
python evaluation/run_core.py --device cuda:0 --dtype bfloat16 --batch-size 16 --output evaluation-rerun
To reproduce after downloading this model repository, accept the BananaMind dataset terms, authenticate with hf auth login, then run in a suitable CUDA environment:
pip install -r evaluation/comparison-20261001/repro/requirements.txt
python evaluation/comparison-20261001/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output comparison-rerun
The bundled runner uses the published model classes with the exact evaluation adapters and tokenizer. --limit produces smoke results only. Raw dataset examples are not included in this release.
TensorBoard and limitations
TensorBoard event files include training telemetry and eval/<task>/<metric> at step 15,000, plus separate CI bounds. Training telemetry covers steps 20–15,000 (750 loss points), including token CE, router loss, gradient norm, throughput, memory, padding and denoiser fractions.
These are small English continuation models, not instruction-tuned assistants. Equal source-token budgets do not imply equal target-token supervision or FLOPs. Benchmark contamination was not audited; results are from one training seed. Reference-model scores from different prompts, tokenizers or corpora are not directly interchangeable.