t5-moe-55M-base / README.md
d0rj's picture
Integrate all benchmark scores into the main evaluation table
de3885b verified
|
Raw History Blame Contribute Delete
18.5 kB
metadata
language:
  - en
library_name: transformers
pipeline_tag: text-generation
tags:
  - pretrained
  - from-scratch
  - tiny-llm-ablation
  - custom_code
  - tensorboard
  - ul2
  - moe
  - encoder-decoder
datasets:
  - HuggingFaceFW/fineweb-edu
model-index:
  - name: t5-moe-55M-base
    results:
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: Rowan/hellaswag
          name: HellaSwag
          config: default
          split: validation
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12)
            value: 0.2779326827325234
            args:
              ci95_low: 0.2692569969516344
              ci95_high: 0.28677820246076313
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: allenai/ai2_arc
          name: ARC-Easy
          config: ARC-Easy
          split: test
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12)
            value: 0.38930976430976433
            args:
              ci95_low: 0.3698977221380954
              ci95_high: 0.40907915127907174
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: allenai/ai2_arc
          name: ARC-Challenge
          config: ARC-Challenge
          split: test
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12)
            value: 0.2380546075085324
            args:
              ci95_low: 0.21455235182352161
              ci95_high: 0.26326840760453163
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: baber/piqa
          name: PIQA
          config: default
          split: validation
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12)
            value: 0.5560391730141458
            args:
              ci95_low: 0.5332313404997734
              ci95_high: 0.5786132479763156
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: allenai/winogrande
          name: WinoGrande
          config: winogrande_xl
          split: validation
          args:
            num_few_shot: 0
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12)
            value: 0.4846093133385951
            args:
              ci95_low: 0.4571789703267686
              ci95_high: 0.512132701298903
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: allenai/openbookqa
          name: OpenBookQA
          config: main
          split: test
          args:
            num_few_shot: 0
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12)
            value: 0.262
            args:
              ci95_low: 0.22537626941723765
              ci95_high: 0.30225291664246134
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: aps/super_glue
          name: BoolQ
          config: boolq
          split: validation
          args:
            num_few_shot: 0
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12)
            value: 0.499388379204893
            args:
              ci95_low: 0.48226179513749795
              ci95_high: 0.5165163985990305
              ci_method: Wilson
      - task:
          type: text-generation
          name: Zero-shot continuation likelihood
        dataset:
          type: EleutherAI/lambada_openai
          name: LAMBADA OpenAI
          config: default
          split: test
          args:
            num_few_shot: 0
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12)
            value: 0.14826314768096255
            args:
              ci95_low: 0.13882265856941103
              ci95_high: 0.1582276717641896
              ci_method: Wilson
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: AxiomicLabs/Arithmark-3.0
          name: ArithMark-3
          config: default
          split: train
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.342
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 1024
              standard_error: 0.015008706182121804
              evaluation_date: '2026-10-01'
              ci95_low: 0.3132528988050859
              ci95_high: 0.37195635687634954
              ci_method: Wilson 95%; item independence approximation
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: pkavumba/balanced-copa
          name: Balanced COPA
          config: default
          split: train
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.538
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.015773547629015002
              evaluation_date: '2026-10-01'
              ci95_low: 0.5070132970793998
              ci95_high: 0.5686958692756982
              ci_method: Wilson 95%; item independence approximation
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: tau/commonsense_qa
          name: CommonsenseQA
          config: default
          split: validation
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.20065520065520065
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.011466011466011467
              evaluation_date: '2026-10-01'
              ci95_low: 0.17914588148536167
              ci95_high: 0.2240421844184298
              ci_method: Wilson 95%; item independence approximation
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: allenai/sciq
          name: SciQ (with support)
          config: default
          split: test
        metrics:
          - type: acc_norm
            name: acc_norm (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.654
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.01505026612756434
              evaluation_date: '2026-10-01'
              ci95_low: 0.6239780184885133
              ci95_high: 0.6828433398979359
              ci_method: Wilson 95%; item independence approximation
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: truthfulqa/truthful_qa
          name: TruthfulQA MC2
          config: multiple_choice
          split: validation
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.4656644987453161
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.016000577772697092
              evaluation_date: '2026-10-01'
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: BananaMind/BananaMind-Base-Bench-1.1
          name: BananaMind Base 1.1
          config: default
          split: test
        metrics:
          - type: raw_accuracy
            name: raw_accuracy (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.36
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.025693810923465083
              evaluation_date: '2026-10-01'
              ci95_low: 0.3114835744745811
              ci95_high: 0.41155622892601307
              ci_method: Wilson 95%; item independence approximation
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: cais/mmlu
          name: MMLU continuation
          config: 57 subjects
          split: test
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.24932345819683804
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.003640204268569072
              evaluation_date: '2026-10-01'
      - task:
          type: text-generation
          name: Continuation likelihood
        dataset:
          type: nyu-mll/blimp
          name: BLiMP
          config: 67 minimal-pair subsets
          split: train
        metrics:
          - type: acc
            name: acc (fraction; lm-eval 0.4.12 comparison protocol)
            value: 0.6988955223880597
            args:
              dtype: bfloat16
              num_few_shot: 0
              max_length: 2048
              standard_error: 0.0015386258071003309
              evaluation_date: '2026-10-01'

T5 MoE 55M Base (UL2)

A 54,858,240-parameter English base model in the Tiny llm ablation experiment. Trained from scratch on exactly 3,932,160,000 source tokens over 15,000 optimizer steps. The token count measures processed input blocks, not unique text or supervised target tokens.

Architecture and references

6 encoder + 6 decoder layers, width 512; encoder 8-head attention, decoder 8 query / 2 KV heads; 8 experts per MoE layer, top-2 routing, expert width 160; RoPE, RMSNorm, FP32 residuals and tied shared embeddings. Maximum encoder length 2050 including controls; raw training blocks 2048.

Architecture inspiration: yandex/AliceAI-T5-35B-A0.6B. Tokenizer foundation: q-project/Q-50M-Base, preserving all 32,768 original IDs and adding 3 mode tokens + 512 sentinels (33,283 entries). Weights were initialized randomly. This small adaptation does not reproduce Alice’s corpus, optimizer or routing recipe.

Training

  • Data: FineWeb-Edu, sample-10BT, streamed from local Parquet shards; shuffle buffer 100,000.
  • Objective: UL2 with seven equally likely denoisers: R(15%, mean span 3/8), S(suffix), X(50%,3), X(50%,8), X(15%,64), X(50%,64). S masks a uniformly sampled suffix of length 1..L/2. Targets contain corrupted spans and control tokens. Training adds router auxiliary loss with coefficient 0.01. These sampler choices are explicit local choices; see UL2.
  • Batch: 16 sequences × 8 accumulation × 2048 tokens = 262,144 source tokens per step.
  • Fused AdamW; peak LR 0.001, betas (0.9, 0.95), weight decay 0.1 (no decay for bias/norm/1D parameters), gradient clipping 1.0. Linear warmup for 150 steps, then cosine decay to 10% of peak LR.
  • BF16 compute on one RTX 5070 Ti (16 GB), seed 2026; checkpoints retain FP32 weights. Exact training configuration.

Evaluation

Full selected task splits, no added few-shot examples, lm-eval 0.4.12, no chat template, BF16 on RTX 5070 Ti, maximum context 2048 (ArithMark: 1024). Accuracy is a percentage. ± is one standard error; the separate bracketed column is the 95% Wilson confidence interval. Intervals describe finite evaluation-sample uncertainty, not variation across training seeds; no multiple-comparison correction is applied.

Dataset Split Examples Metric Score ± SE (%) 95% CI (%)
HellaSwag validation 10,042 acc_norm 27.79 ± 0.45 [26.93, 28.68]
ARC-Easy test 2,376 acc_norm 38.93 ± 1.00 [36.99, 40.91]
ARC-Challenge test 1,172 acc_norm 23.81 ± 1.24 [21.46, 26.33]
PIQA validation 1,838 acc_norm 55.60 ± 1.16 [53.32, 57.86]
WinoGrande validation 1,267 acc 48.46 ± 1.40 [45.72, 51.21]
OpenBookQA test 500 acc_norm 26.20 ± 1.97 [22.54, 30.23]
BoolQ validation 3,270 acc 49.94 ± 0.87 [48.23, 51.65]
LAMBADA OpenAI test 5,153 acc 14.83 ± 0.50 [13.88, 15.82]
ArithMark-3 train 1,000 acc_norm 34.20 ± 1.50 [31.33, 37.20]
Balanced COPA train 1,000 acc 53.80 ± 1.58 [50.70, 56.87]
CommonsenseQA validation 1,221 acc 20.07 ± 1.15 [17.91, 22.40]
SciQ (with support) test 1,000 acc_norm 65.40 ± 1.51 [62.40, 68.28]
TruthfulQA MC2 validation 817 acc 46.57 ± 1.60 —
BananaMind Base 1.1 test 350 raw_accuracy 36.00 ± 2.57 [31.15, 41.16]
MMLU continuation test 14,042 acc 24.93 ± 0.36 —
BLiMP train 67,000 acc 69.89 ± 0.15 —

T5 uses UL2 S-mode: encoder S + prefix + sentinel + EOS; decoder BOS + sentinel + shifted answer. Only answer text is scored; control tokens and router loss are excluded, with the full vocabulary retained in the softmax. Its encoder sees at most 2047 text-prefix tokens after reserving controls. LAMBADA accuracy requires the complete final-word token sequence. acc_norm is harness length-normalized option scoring; raw accuracy is also stored in results.json.

WikiText-2 raw test, conditional continuation: CPU FP32 re-evaluation on 291 nonoverlapping blocks (512 prefix + 512 scored suffix tokens), 148,992 scored tokens; 335 tail tokens excluded. NLL 3.920357, 95% CI [3.879299, 3.960742]; token PPL 50.418, 95% CI [48.390, 52.496]. Percentile block bootstrap, 10,000 resamples, seed 2026; exponentiate NLL endpoints for PPL. Blocks are the resampling unit; this does not model all within-document dependence. This is not standard rolling AR or word PPL. The earlier BF16 point is retained separately in TensorBoard, with no borrowed FP32 interval.

The metadata contains author-reported model-index scores. The evaluated dataset repositories had no registered eval.yaml on 2026-09-20, so no .eval_results leaderboard entry or verified badge is claimed. Machine-readable results and provenance.

Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti; context cap 2048 (ArithMark 1024), TF32 disabled, no chat template and no added few-shot examples. TruthfulQA retains the harness's fixed six-QA preamble. ArithMark and BananaMind normalize by continuation token count; ordinary harness acc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo. SciQ includes the support passage. Balanced COPA uses the mirrored 1000-item train-named evaluation split; cRia's split was inferred, not confirmed. MMLU scores full answer continuations across 57 subjects, weighted by item count; BLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained from each evaluator. Wilson intervals are reported only where the runner logged binary item accuracy; MC2 is probability mass, not binary accuracy. These intervals do not model dependence between paired/templated examples or training seed variation. UL2, PrefixLM and experimental diffusion PLL use their documented conditional scoring protocols; PLL exposes the other answer tokens and is not autoregressive likelihood. cRia's published scores used a different precision and benchmark-adapted checkpoint; this completes our comparison coverage, not an independent reproduction of cRia or an official leaderboard submission.

Full results, provenance and group scores. Updated machine-readable results. TensorBoard events contain these new scores at step 15,000.

Usage

Install requirements.txt (tested with Transformers 5.17.0 / PyTorch 2.11.0). Custom model code is included; trust_remote_code=True is required. This example runs on CPU.

import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
repo = "d0rj/t5-moe-55M-base"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
c = model.config.ul2
prefix = tokenizer.encode("The capital of France is", add_special_tokens=False)
inputs = torch.tensor([[c["mode_ids"]["S"], *prefix, c["sentinel_ids"][0], c["eos_id"]]])
decoder = torch.tensor([[model.config.decoder_start_token_id, c["sentinel_ids"][0]]])
output = model.generate(input_ids=inputs, decoder_input_ids=decoder,
                        max_new_tokens=64, do_sample=False)
print(tokenizer.decode(output[0, decoder.shape[1]:], skip_special_tokens=True))

To reproduce the core evaluation from a downloaded repository, install evaluation/requirements.txt and run:

python evaluation/run_core.py --device cuda:0 --dtype bfloat16 --batch-size 16 --output evaluation-rerun

To reproduce after downloading this model repository, accept the BananaMind dataset terms, authenticate with hf auth login, then run in a suitable CUDA environment:

pip install -r evaluation/comparison-20261001/repro/requirements.txt
python evaluation/comparison-20261001/repro/run.py --device cuda:0 --dtype bfloat16 --batch-size 8 --output comparison-rerun

The bundled runner uses the published model classes with the exact evaluation adapters and tokenizer. --limit produces smoke results only. Raw dataset examples are not included in this release.

TensorBoard and limitations

TensorBoard event files include training telemetry and eval/<task>/<metric> at step 15,000, plus separate CI bounds. Training telemetry covers steps 20–15,000 (750 loss points), including token CE, router loss, gradient norm, throughput, memory, padding and denoiser fractions.

These are small English continuation models, not instruction-tuned assistants. Equal source-token budgets do not imply equal target-token supervision or FLOPs. Benchmark contamination was not audited; results are from one training seed. Reference-model scores from different prompts, tokenizers or corpora are not directly interchangeable.