Text Generation
Transformers
TensorBoard
Safetensors
English
alicet5_moe
text2text-generation
pretrained
from-scratch
tiny-llm-ablation
custom_code
ul2
Mixture of Experts
encoder-decoder
Eval Results (legacy)
Instructions to use d0rj/t5-moe-55M-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use d0rj/t5-moe-55M-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="d0rj/t5-moe-55M-base", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("d0rj/t5-moe-55M-base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use d0rj/t5-moe-55M-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "d0rj/t5-moe-55M-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/d0rj/t5-moe-55M-base
- SGLang
How to use d0rj/t5-moe-55M-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "d0rj/t5-moe-55M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "d0rj/t5-moe-55M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0rj/t5-moe-55M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use d0rj/t5-moe-55M-base with Docker Model Runner:
docker model run hf.co/d0rj/t5-moe-55M-base
Download evaluation/comparison-20261001/results.json from d0rj/t5-moe-55M-base: direct link, hf CLI and curl.
- Browser
- Download file 36.1 kB
-
https://huggingface.co/d0rj/t5-moe-55M-base/resolve/main/evaluation/comparison-20261001/results.json
- Command line
-
hf download hf://d0rj/t5-moe-55M-base/evaluation/comparison-20261001/results.json
-
curl -L -o results.json https://huggingface.co/d0rj/t5-moe-55M-base/resolve/main/evaluation/comparison-20261001/results.json
36.1 kB
| { | |
| "date": "2026-10-01", | |
| "model": "d0rj/t5-moe-55M-base", | |
| "protocol": "Full selected splits; lm-eval 0.4.12; seed 1234; BF16 on RTX 5070 Ti;\ncontext cap 2048 (ArithMark 1024), TF32 disabled, no chat template and no added\nfew-shot examples. TruthfulQA retains the harness's fixed six-QA preamble.\nArithMark and BananaMind normalize by continuation token count; ordinary harness\nacc_norm uses its own length normalization. BananaMind is raw accuracy, not Elo.\nSciQ includes the support passage. Balanced COPA uses the mirrored 1000-item\ntrain-named evaluation split; cRia's split was inferred, not confirmed.\nMMLU scores full answer continuations across 57 subjects, weighted by item count;\nBLiMP averages 67 equal-sized minimal-pair subsets. Standard errors are retained\nfrom each evaluator. Wilson intervals are reported only where the runner logged\nbinary item accuracy; MC2 is probability mass, not binary accuracy. These\nintervals do not model dependence between paired/templated examples or training\nseed variation. UL2, PrefixLM and experimental diffusion PLL use their documented\nconditional scoring protocols; PLL exposes the other answer tokens and is not\nautoregressive likelihood. cRia's published scores used a different precision\nand benchmark-adapted checkpoint; this completes our comparison coverage, not\nan independent reproduction of cRia or an official leaderboard submission.", | |
| "results": { | |
| "checkpoint": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/checkpoints/t5", | |
| "checkpoint_sha256": { | |
| "config.json": "8031de3caf99e9d424970b69701b1077d9dff5151bafd4ce214936a50fbd3bd4", | |
| "generation_config.json": "8f1ce3b416c7f594951681c73e2b3c7933baac0a09a4f8d649aad27af21941f6", | |
| "model.safetensors": "1efcd3aeed90f8b87ff06495f29f2cffa36435db1885d34a5e237450d2404b9f", | |
| "tokenizer.json": "1d1f7409de3e53ae51d7085b3009aa5cd491dc08fbe114794c8d314f3fba2af4", | |
| "tokenizer_config.json": "879515efa0fe331af8eeb0ede00fdd13c0bd3c827bff8bb6f4b7f959c02159ec" | |
| }, | |
| "core": "/mnt/d/Projects/tiny_llm/runs/eval/t5-core-full-20260920", | |
| "training_run": "moe_t5-gpu-v1", | |
| "training_step": 15000, | |
| "hf": "d0rj/t5-moe-55M-base", | |
| "tokenizer": "runs/moe_t5-gpu-v1/checkpoint-last", | |
| "tb": "runs/moe_t5-gpu-v1/tb", | |
| "protocol": "UL2 S-mode: encoder S+context+sentinel+EOS; decoder BOS+sentinel+answer; score answer text only, full vocabulary", | |
| "kind": "seq2seq", | |
| "tasks": { | |
| "arithmark3": { | |
| "label": "ArithMark-3", | |
| "dataset": "AxiomicLabs/Arithmark-3.0", | |
| "config": "default", | |
| "split": "train", | |
| "samples": 1000, | |
| "shots": 0, | |
| "primary": "acc_norm", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.346, | |
| "stderr": 0.015050266127564339, | |
| "ci95": [ | |
| 0.3171566601020642, | |
| 0.3760219815114867 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| }, | |
| "acc_norm": { | |
| "value": 0.342, | |
| "stderr": 0.015008706182121804, | |
| "ci95": [ | |
| 0.3132528988050859, | |
| 0.37195635687634954 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/arithmark3.json", | |
| "result_sha256": "0d0d9bc0918204a22db4c2e9acbafcddd494e433c46ae82d0dedfde50aea3eef", | |
| "categories": { | |
| "elementary_school_math_continuation::addition::grades_1_2::easy": { | |
| "count": 128, | |
| "acc_norm,none": 0.25, | |
| "acc_norm_stderr,none": 0.038423663968391995 | |
| }, | |
| "elementary_school_math_continuation::comparison::grades_2_3::medium": { | |
| "count": 44, | |
| "acc_norm,none": 0.2727272727272727, | |
| "acc_norm_stderr,none": 0.06791703342160257 | |
| }, | |
| "elementary_school_math_continuation::comparison_difference::grades_2_3::medium": { | |
| "count": 48, | |
| "acc_norm,none": 0.2916666666666667, | |
| "acc_norm_stderr,none": 0.06629996666059657 | |
| }, | |
| "elementary_school_math_continuation::data::grades_2_3::easy": { | |
| "count": 43, | |
| "acc_norm,none": 0.2558139534883721, | |
| "acc_norm_stderr,none": 0.06732528971681358 | |
| }, | |
| "elementary_school_math_continuation::division::grades_3_4::medium": { | |
| "count": 54, | |
| "acc_norm,none": 0.42592592592592593, | |
| "acc_norm_stderr,none": 0.06792240738867397 | |
| }, | |
| "elementary_school_math_continuation::fractions_counting::grades_3_4::medium": { | |
| "count": 50, | |
| "acc_norm,none": 0.22, | |
| "acc_norm_stderr,none": 0.05917804336345137 | |
| }, | |
| "elementary_school_math_continuation::geometry_area::grades_4_5::medium": { | |
| "count": 52, | |
| "acc_norm,none": 0.5384615384615384, | |
| "acc_norm_stderr,none": 0.06980655484407924 | |
| }, | |
| "elementary_school_math_continuation::geometry_perimeter::grades_4_5::medium": { | |
| "count": 45, | |
| "acc_norm,none": 0.5555555555555556, | |
| "acc_norm_stderr,none": 0.07491109582924912 | |
| }, | |
| "elementary_school_math_continuation::measurement::grades_2_3::easy": { | |
| "count": 76, | |
| "acc_norm,none": 0.35526315789473684, | |
| "acc_norm_stderr,none": 0.05526315789473684 | |
| }, | |
| "elementary_school_math_continuation::money::grades_3_4::medium": { | |
| "count": 64, | |
| "acc_norm,none": 0.296875, | |
| "acc_norm_stderr,none": 0.05756159356351619 | |
| }, | |
| "elementary_school_math_continuation::multiplication::grades_3_4::medium": { | |
| "count": 74, | |
| "acc_norm,none": 0.5, | |
| "acc_norm_stderr,none": 0.05852057359806528 | |
| }, | |
| "elementary_school_math_continuation::patterns::grades_3_4::medium": { | |
| "count": 53, | |
| "acc_norm,none": 0.20754716981132076, | |
| "acc_norm_stderr,none": 0.05623975840347623 | |
| }, | |
| "elementary_school_math_continuation::subtraction::grades_1_2::easy": { | |
| "count": 117, | |
| "acc_norm,none": 0.1794871794871795, | |
| "acc_norm_stderr,none": 0.035631196604084675 | |
| }, | |
| "elementary_school_math_continuation::time::grades_2_3::easy": { | |
| "count": 55, | |
| "acc_norm,none": 0.9090909090909091, | |
| "acc_norm_stderr,none": 0.03912104390108503 | |
| }, | |
| "elementary_school_math_continuation::two_step_add_subtract::grades_2_3::medium": { | |
| "count": 46, | |
| "acc_norm,none": 0.1956521739130435, | |
| "acc_norm_stderr,none": 0.05913682829884975 | |
| }, | |
| "elementary_school_math_continuation::two_step_addition::grades_2_3::medium": { | |
| "count": 19, | |
| "acc_norm,none": 0.10526315789473684, | |
| "acc_norm_stderr,none": 0.07233518641434492 | |
| }, | |
| "elementary_school_math_continuation::two_step_subtraction::grades_2_3::medium": { | |
| "count": 32, | |
| "acc_norm,none": 0.3125, | |
| "acc_norm_stderr,none": 0.08324928557283298 | |
| } | |
| } | |
| }, | |
| "balanced_copa": { | |
| "label": "Balanced COPA", | |
| "dataset": "pkavumba/balanced-copa", | |
| "config": "default", | |
| "split": "train", | |
| "samples": 1000, | |
| "shots": 0, | |
| "primary": "acc", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.538, | |
| "stderr": 0.015773547629015002, | |
| "ci95": [ | |
| 0.5070132970793998, | |
| 0.5686958692756982 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/balanced_copa.json", | |
| "result_sha256": "62c4a30d3f5847a68094e91c764dc79ba3e9a2e2a9c67236af6d5f5266018f78", | |
| "categories": {} | |
| }, | |
| "commonsense_qa": { | |
| "label": "CommonsenseQA", | |
| "dataset": "tau/commonsense_qa", | |
| "config": "default", | |
| "split": "validation", | |
| "samples": 1221, | |
| "shots": 0, | |
| "primary": "acc", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.20065520065520065, | |
| "stderr": 0.011466011466011467, | |
| "ci95": [ | |
| 0.17914588148536167, | |
| 0.2240421844184298 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/commonsense_qa.json", | |
| "result_sha256": "8c3ba38b4757c3678e717a81217c7c304c74309f0cec30ca34f28222347d0ec4", | |
| "categories": {} | |
| }, | |
| "sciq": { | |
| "label": "SciQ (with support)", | |
| "dataset": "allenai/sciq", | |
| "config": "default", | |
| "split": "test", | |
| "samples": 1000, | |
| "shots": 0, | |
| "primary": "acc_norm", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.703, | |
| "stderr": 0.01445683229480096, | |
| "ci95": [ | |
| 0.6739460359438715, | |
| 0.730500300110993 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| }, | |
| "acc_norm": { | |
| "value": 0.654, | |
| "stderr": 0.01505026612756434, | |
| "ci95": [ | |
| 0.6239780184885133, | |
| 0.6828433398979359 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/sciq.json", | |
| "result_sha256": "d2aa221220038ac172272d4f598ade4df4841839b2de315e4438bb8f3da0c031", | |
| "categories": {} | |
| }, | |
| "truthfulqa_mc2": { | |
| "label": "TruthfulQA MC2", | |
| "dataset": "truthfulqa/truthful_qa", | |
| "config": "multiple_choice", | |
| "split": "validation", | |
| "samples": 817, | |
| "shots": 0, | |
| "primary": "acc", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.4656644987453161, | |
| "stderr": 0.016000577772697092 | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/truthfulqa_mc2.json", | |
| "result_sha256": "951d4dd4cfeba277592bdc041d88516a36c8f3ee1c21ab1bb20668bbc7b6d329", | |
| "categories": {} | |
| }, | |
| "bananamind_base": { | |
| "label": "BananaMind Base 1.1", | |
| "dataset": "BananaMind/BananaMind-Base-Bench-1.1", | |
| "config": "default", | |
| "split": "test", | |
| "samples": 350, | |
| "shots": 0, | |
| "primary": "raw_accuracy", | |
| "metrics": { | |
| "raw_accuracy": { | |
| "value": 0.36, | |
| "stderr": 0.025693810923465083, | |
| "ci95": [ | |
| 0.3114835744745811, | |
| 0.41155622892601307 | |
| ], | |
| "ci_method": "Wilson 95%; item independence approximation" | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/bananamind_base.json", | |
| "result_sha256": "473694fd9a25bf2721607c4ef28aa17f19f6850af3760f1c90a42c1dfb59a071", | |
| "categories": { | |
| "code_completion": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.24, | |
| "raw_accuracy_stderr,none": 0.06101187572589321 | |
| }, | |
| "commonsense": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.36, | |
| "raw_accuracy_stderr,none": 0.06857142857142857 | |
| }, | |
| "context_tracking": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.36, | |
| "raw_accuracy_stderr,none": 0.06857142857142857 | |
| }, | |
| "language_completion": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.42, | |
| "raw_accuracy_stderr,none": 0.07050835816716038 | |
| }, | |
| "logical_reasoning": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.34, | |
| "raw_accuracy_stderr,none": 0.0676726816132972 | |
| }, | |
| "quantitative": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.18, | |
| "raw_accuracy_stderr,none": 0.054883922035138706 | |
| }, | |
| "world_knowledge": { | |
| "count": 50, | |
| "raw_accuracy,none": 0.62, | |
| "raw_accuracy_stderr,none": 0.06934092056863769 | |
| } | |
| } | |
| }, | |
| "mmlu_continuation": { | |
| "label": "MMLU continuation", | |
| "dataset": "cais/mmlu", | |
| "config": "57 subjects", | |
| "split": "test", | |
| "samples": 14042, | |
| "shots": 0, | |
| "primary": "acc", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.24932345819683804, | |
| "stderr": 0.003640204268569072 | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-001/mmlu_continuation.json", | |
| "result_sha256": "8e939afec45700c2d5f70dafa5df87b52841ace350b74969bbfcacc9fc1e54f8", | |
| "categories": {} | |
| }, | |
| "blimp": { | |
| "label": "BLiMP", | |
| "dataset": "nyu-mll/blimp", | |
| "config": "67 minimal-pair subsets", | |
| "split": "train", | |
| "samples": 67000, | |
| "shots": 0, | |
| "primary": "acc", | |
| "metrics": { | |
| "acc": { | |
| "value": 0.6988955223880597, | |
| "stderr": 0.0015386258071003309 | |
| } | |
| }, | |
| "result_path": "runs/eval/cria-backfill-20261001/t5/attempt-002/blimp.json", | |
| "result_sha256": "58646e23f3b9a6fbeca3b50b503fe9893fe83c34306a830f520761e225bb4c2f", | |
| "categories": {} | |
| } | |
| }, | |
| "batch_size": 8, | |
| "label": "T5 MoE UL2", | |
| "manifests": { | |
| "runs/eval/cria-backfill-20261001/t5/attempt-001": { | |
| "checkpoint": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/checkpoints/t5", | |
| "local_model": "moe_t5", | |
| "kind": "seq2seq", | |
| "tokenizer": "runs/moe_t5-gpu-v1/checkpoint-last", | |
| "revision": "main", | |
| "tokenizer_revision": null, | |
| "protocol": "UL2 S-mode: encoder S+context+sentinel+EOS; decoder BOS+sentinel+answer; score answer text only, full vocabulary", | |
| "harness_version": "0.4.12", | |
| "smoke_only": false, | |
| "tasks": [ | |
| { | |
| "task": "arithmark3", | |
| "shots": 0, | |
| "description": "ArithMark-3.0 train-named evaluation split; token-normalized acc_norm", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "token_continuation" | |
| }, | |
| { | |
| "task": "balanced_copa", | |
| "shots": 0, | |
| "description": "Balanced COPA mirrored 1000-item train split; acc; cRia split inferred", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| }, | |
| { | |
| "task": "commonsense_qa", | |
| "shots": 0, | |
| "description": "CommonsenseQA; acc", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| }, | |
| { | |
| "task": "sciq", | |
| "shots": 0, | |
| "description": "SciQ test with support passage; acc and acc_norm", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| }, | |
| { | |
| "task": "truthfulqa_mc2", | |
| "shots": 0, | |
| "description": "TruthfulQA multiple choice; mc2", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| }, | |
| { | |
| "task": "bananamind_base", | |
| "shots": 0, | |
| "description": "BananaMind Base Bench 1.1 test; mean-token raw accuracy; gated", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "token_continuation" | |
| }, | |
| { | |
| "task": "mmlu_continuation", | |
| "shots": 0, | |
| "description": "MMLU cloze: full answer continuations; acc", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| }, | |
| { | |
| "task": "blimp", | |
| "shots": 0, | |
| "description": "Full BLiMP group; grammatical minimal pairs", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| } | |
| ], | |
| "options": { | |
| "checkpoint": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/checkpoints/t5", | |
| "model": null, | |
| "tokenizer": "runs/moe_t5-gpu-v1/checkpoint-last", | |
| "revision": "main", | |
| "tokenizer_revision": null, | |
| "suite": "core", | |
| "tasks": "arithmark3,balanced_copa,commonsense_qa,sciq,truthfulqa_mc2,bananamind_base,mmlu_continuation,blimp", | |
| "num_fewshot": null, | |
| "device": "cuda:0", | |
| "dtype": "bfloat16", | |
| "allow_tf32": false, | |
| "batch_size": 8, | |
| "max_length": 2048, | |
| "max_gen_toks": null, | |
| "limit": null, | |
| "seed": 1234, | |
| "threads": 4, | |
| "bootstrap_iters": 1000, | |
| "trust_remote_code": false, | |
| "allow_code_execution": false, | |
| "apply_chat_template": false, | |
| "log_samples": false, | |
| "output": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/t5/attempt-001", | |
| "tb_dir": "runs/moe_t5-gpu-v1/tb", | |
| "training_step": 15000, | |
| "list": false, | |
| "dry_run": false | |
| }, | |
| "checkpoint_stamp": { | |
| "model.safetensors": [ | |
| 219483760, | |
| 1789897063000000000 | |
| ], | |
| "config.json": [ | |
| 7837, | |
| 1789897062000000000 | |
| ], | |
| "generation_config.json": [ | |
| 226, | |
| 1789897062000000000 | |
| ], | |
| "tokenizer.json": [ | |
| 2420528, | |
| 1789897063000000000 | |
| ], | |
| "tokenizer_config.json": [ | |
| 9425, | |
| 1789897063000000000 | |
| ] | |
| }, | |
| "checkpoint_sha256": { | |
| "config.json": "8031de3caf99e9d424970b69701b1077d9dff5151bafd4ce214936a50fbd3bd4", | |
| "generation_config.json": "8f1ce3b416c7f594951681c73e2b3c7933baac0a09a4f8d649aad27af21941f6", | |
| "model.safetensors": "1efcd3aeed90f8b87ff06495f29f2cffa36435db1885d34a5e237450d2404b9f", | |
| "tokenizer.json": "1d1f7409de3e53ae51d7085b3009aa5cd491dc08fbe114794c8d314f3fba2af4", | |
| "tokenizer_config.json": "879515efa0fe331af8eeb0ede00fdd13c0bd3c827bff8bb6f4b7f959c02159ec" | |
| }, | |
| "source_sha256": { | |
| "src/tiny_llm/__init__.py": "8c7f075d51006a53fe1d645ef5a84d65db58ff4251e7544074fe29ee7549c14c", | |
| "src/tiny_llm/benchmarks.py": "5d9f461366389bd1f926c6a6468192f3e2313c87caf5b1de27f407d563a858a4", | |
| "src/tiny_llm/checkpoint.py": "ac9df5f8c981a47eb1f1e59d0687ff41a1f6c38ca74ced50f71ac446a58d6ef5", | |
| "src/tiny_llm/comet_sync.py": "ea802bfb58dacad66f83b9e5ac5ccb330c190204a81c90c9e14aafe44c893c37", | |
| "src/tiny_llm/comet_views.py": "3e03f4ae69f8edc2eaf5f946f5cc90f373ca1b9cc91f9dd2e9adae6d6a672689", | |
| "src/tiny_llm/continuation.py": "4f4143c828d4f94a851bcbf8a466d3e1b3e9eaec1027970a426208cc0ff4f4db", | |
| "src/tiny_llm/data.py": "1a52e38afef9af32463779cc2f471d9b7174a7c4c73753e74f485d526d3c6733", | |
| "src/tiny_llm/diffusion_diagnostics.py": "2ae667a076b1d3f057a29d20f00559068eb73a073b3f99030e1d5b4c2f478fa5", | |
| "src/tiny_llm/evaluation/__init__.py": "e49e80f7be3ab6e200f0db14b028e465de9caa8f6f705430eabce0c41b8322ba", | |
| "src/tiny_llm/evaluation/adapters.py": "8e64e5542bbcb0f9ad448b320e712916a9559823987e298b22321a8d8e2f727f", | |
| "src/tiny_llm/evaluation/catalog.py": "2fd29084ec780539023a38bcbfd7302931fa6be48624e227875f13518a1b986f", | |
| "src/tiny_llm/evaluation/cli.py": "a322e39bf58ba52132249dae524e9772c3a8e1a5ba388f5c0267f8dc1ea18882", | |
| "src/tiny_llm/evaluation/tasks/balanced_copa.yaml": "f47baa6a713e15f7a5600ed257f316f8c54252bf4977bbc716a6da7aa580f05d", | |
| "src/tiny_llm/evaluation/tasks/balanced_copa_test.yaml": "dce21ad70b395d8d576c584be1b87ebae846e360e59878030645215a387b28e6", | |
| "src/tiny_llm/evaluation/telemetry.py": "f4dad30574753739ea10ba71ef9a7ed9d9118a89a8202dee224961be27f13a48", | |
| "src/tiny_llm/evaluation/token_continuation.py": "e8d88cd4aca51ca2fda07f2b799a54612d0e2964a98e05824a135bea1a533550", | |
| "src/tiny_llm/losses.py": "592a05287369f6e5077b3faca404a9392de4772cf06389fb8aa7e36924e74f86", | |
| "src/tiny_llm/models/__init__.py": "b94ab167b6eb9f9b455823eea382e67b5ac2cc7f9ef5f6de183a09b9650b9718", | |
| "src/tiny_llm/models/diffusion/__init__.py": "c92eaf5986d604d8aaf1894a862a217cefa32ab25359175b88110e45a71c1b25", | |
| "src/tiny_llm/models/diffusion/configuration_diffusion_lm.py": "3b303aba60f0a831c966f5cfa7f5441daa2584814e335646195bd6ceecfa6d0c", | |
| "src/tiny_llm/models/diffusion/modeling_diffusion_lm.py": "4fc21c2aed1d5a231afe6f7634d9d8c3682b4e2abfc8fecef27b52136de77b2f", | |
| "src/tiny_llm/models/looped/__init__.py": "4b46b426021c59ed313cb2f49afc1926ad2793d771ef90117e3fef0462957cff", | |
| "src/tiny_llm/models/looped/configuration_looped_lm.py": "b8d99c52b98e0a97a4cae76ca4c20affd8f4d019ec38e0a9beefd5fad864dc25", | |
| "src/tiny_llm/models/looped/modeling_looped_lm.py": "9bdb3c5e3aaae663d09e98055f95df1adbc42e0a4bf5dffb7f6c39242af41b9e", | |
| "src/tiny_llm/models/moe_t5/__init__.py": "18e33fe38b2049521f573512ac8796afa209e1b25c2052784e1c927a7160d66c", | |
| "src/tiny_llm/models/moe_t5/configuration_alicet5_moe.py": "4647b1e9ab9ce07887bffc1abc3beb3e691962a380b4cc237aae376490bb56b7", | |
| "src/tiny_llm/models/moe_t5/modeling_alicet5_moe.py": "5f899c4726b120eed6c9bf6258ad319bb1ccea734645f27e7e6ccbc5057222ff", | |
| "src/tiny_llm/models/prefixlm/__init__.py": "c94c3c861acba8d2147a3a6c46bccfc71ac8adc8d70757bede009598029d31ea", | |
| "src/tiny_llm/models/prefixlm/configuration_prefix_lm.py": "20b3a62ae3628b5a64566ef19d2947c90f10c6159daad00dd13d18fc5910c42c", | |
| "src/tiny_llm/models/prefixlm/modeling_prefix_lm.py": "f62a38211df175696f64945173d1e41d9f7fcf6152a7648020da8c91788e98fd", | |
| "src/tiny_llm/models/q50m/__init__.py": "9d0dfec6fa0cfa6967e2f17a73d8fbeca69ba93b4a93651eda4d95626d5f5eb5", | |
| "src/tiny_llm/models/q50m/configuration_q50m.py": "796806b2d5d982b633796bc7de53c3073b1b6e52d259f44ef798b63f55eaea92", | |
| "src/tiny_llm/models/q50m/modeling_q50m.py": "b68abce1f19a36deb39c4ac4e520ffaab0c879d8a7cc2abef35f014bacc011f2", | |
| "src/tiny_llm/models/qmod/__init__.py": "e8c3e10710ac1cceac72ea329db480e680e8bbaae2a92ec26c36c92a778552ad", | |
| "src/tiny_llm/models/qmod/configuration_qmod.py": "0dfc9a4540697a9b15e423b9592df9fc435dcfc437ddad20dc6b6377ebebea40", | |
| "src/tiny_llm/models/qmod/modeling_qmod.py": "3831e34a86309b91a7a0787b9471eaceb91aa4c267f8030eb2e6722a0fcab6f8", | |
| "src/tiny_llm/models/qmoe/__init__.py": "cfb2bb24bc2f3d5b72b3890a393b0386cb25f127e375f5454255478a36dec32f", | |
| "src/tiny_llm/models/qmoe/configuration_qmoe.py": "bba7c9ac16f444146a78c99e7a26b4a11864b215d2b673bf658d9930f2f4de2f", | |
| "src/tiny_llm/models/qmoe/modeling_qmoe.py": "1d713e5e287f46dbf13b7cf9c562b4539efbf6231f7ceaa261f8cea1acd2768c", | |
| "src/tiny_llm/models/qt5/__init__.py": "df943b42a9f339e006687bcd30a666838c1dec15c79751b944e5f7a2bc69186c", | |
| "src/tiny_llm/models/qt5/configuration_qt5.py": "59709fe2c9fe08ccecea6633b29b3b47170fd30abcbee37952507d676446cd44", | |
| "src/tiny_llm/models/qt5/modeling_qt5.py": "6269829cca6631824622df034ff7b544dbbe47790556c60b26e356b5279cceff", | |
| "src/tiny_llm/models/t5/__init__.py": "e4d8838253800edbcac1f9a9b6dc875bcb3ffa2962a0c9044d066448ea10488e", | |
| "src/tiny_llm/models/t5/configuration_alicet5.py": "81b1f594139f201af9afb25ba58df862432c15dedf57f6e340ee2af80ed0d44f", | |
| "src/tiny_llm/models/t5/modeling_alicet5.py": "bd52ce0022668f2defb57770df229f148e070ea73becca9a2e681325114bfc20", | |
| "src/tiny_llm/models/zarya/__init__.py": "651b19cdba422680ff16bf2b2f63be0c3e9045c44c410fb17dbf51df2a23bd55", | |
| "src/tiny_llm/models/zarya/configuration_zarya_lm.py": "42a3e7a1b11b687bac34e00f9cd3cb59d1cd573eec7492e1c55834662d37ed3f", | |
| "src/tiny_llm/models/zarya/modeling_zarya_lm.py": "73ab033860288f7367d5b912c79de908fd81a4e88f369a658b96dd2bcfcd7bfe", | |
| "src/tiny_llm/telemetry.py": "7ee8f14f4545bfa0e5ac632440d49c7f86fd6812a16429f862ebe3cfe09e3545", | |
| "src/tiny_llm/ul2.py": "f4d948d89d5e106965612971dd473a54b57719f811761f18457c6221d763b876", | |
| "src/tiny_llm/zarya.py": "8b5084ab5086c75745115d3a9ba4f2fa9cfda8ad713f65d3f8a22e77db8c0941" | |
| }, | |
| "packages": { | |
| "lm_eval": "0.4.12", | |
| "torch": "2.11.0+cu128", | |
| "transformers": "5.17.0", | |
| "datasets": "5.0.1", | |
| "accelerate": "1.15.0" | |
| }, | |
| "git_commit": "fb65467388c90a7d94d50fbcb711abfd8a568102", | |
| "git_dirty": true | |
| }, | |
| "runs/eval/cria-backfill-20261001/t5/attempt-002": { | |
| "checkpoint": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/checkpoints/t5", | |
| "local_model": "moe_t5", | |
| "kind": "seq2seq", | |
| "tokenizer": "runs/moe_t5-gpu-v1/checkpoint-last", | |
| "revision": "main", | |
| "tokenizer_revision": null, | |
| "protocol": "UL2 S-mode: encoder S+context+sentinel+EOS; decoder BOS+sentinel+answer; score answer text only, full vocabulary", | |
| "harness_version": "0.4.12", | |
| "smoke_only": false, | |
| "tasks": [ | |
| { | |
| "task": "blimp", | |
| "shots": 0, | |
| "description": "Full BLiMP group; grammatical minimal pairs", | |
| "unsafe_code": false, | |
| "rolling": false, | |
| "engine": "lm_eval" | |
| } | |
| ], | |
| "options": { | |
| "checkpoint": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/checkpoints/t5", | |
| "model": null, | |
| "tokenizer": "runs/moe_t5-gpu-v1/checkpoint-last", | |
| "revision": "main", | |
| "tokenizer_revision": null, | |
| "suite": "core", | |
| "tasks": "blimp", | |
| "num_fewshot": null, | |
| "device": "cuda:0", | |
| "dtype": "bfloat16", | |
| "allow_tf32": false, | |
| "batch_size": 8, | |
| "max_length": 2048, | |
| "max_gen_toks": null, | |
| "limit": null, | |
| "seed": 1234, | |
| "threads": 4, | |
| "bootstrap_iters": 1000, | |
| "trust_remote_code": false, | |
| "allow_code_execution": false, | |
| "apply_chat_template": false, | |
| "log_samples": false, | |
| "output": "/mnt/d/Projects/tiny_llm/runs/eval/cria-backfill-20261001/t5/attempt-002", | |
| "tb_dir": "runs/moe_t5-gpu-v1/tb", | |
| "training_step": 15000, | |
| "list": false, | |
| "dry_run": false | |
| }, | |
| "checkpoint_stamp": { | |
| "model.safetensors": [ | |
| 219483760, | |
| 1789897063000000000 | |
| ], | |
| "config.json": [ | |
| 7837, | |
| 1789897062000000000 | |
| ], | |
| "generation_config.json": [ | |
| 226, | |
| 1789897062000000000 | |
| ], | |
| "tokenizer.json": [ | |
| 2420528, | |
| 1789897063000000000 | |
| ], | |
| "tokenizer_config.json": [ | |
| 9425, | |
| 1789897063000000000 | |
| ] | |
| }, | |
| "checkpoint_sha256": { | |
| "config.json": "8031de3caf99e9d424970b69701b1077d9dff5151bafd4ce214936a50fbd3bd4", | |
| "generation_config.json": "8f1ce3b416c7f594951681c73e2b3c7933baac0a09a4f8d649aad27af21941f6", | |
| "model.safetensors": "1efcd3aeed90f8b87ff06495f29f2cffa36435db1885d34a5e237450d2404b9f", | |
| "tokenizer.json": "1d1f7409de3e53ae51d7085b3009aa5cd491dc08fbe114794c8d314f3fba2af4", | |
| "tokenizer_config.json": "879515efa0fe331af8eeb0ede00fdd13c0bd3c827bff8bb6f4b7f959c02159ec" | |
| }, | |
| "source_sha256": { | |
| "src/tiny_llm/__init__.py": "8c7f075d51006a53fe1d645ef5a84d65db58ff4251e7544074fe29ee7549c14c", | |
| "src/tiny_llm/benchmarks.py": "5d9f461366389bd1f926c6a6468192f3e2313c87caf5b1de27f407d563a858a4", | |
| "src/tiny_llm/checkpoint.py": "ac9df5f8c981a47eb1f1e59d0687ff41a1f6c38ca74ced50f71ac446a58d6ef5", | |
| "src/tiny_llm/comet_sync.py": "ea802bfb58dacad66f83b9e5ac5ccb330c190204a81c90c9e14aafe44c893c37", | |
| "src/tiny_llm/comet_views.py": "3e03f4ae69f8edc2eaf5f946f5cc90f373ca1b9cc91f9dd2e9adae6d6a672689", | |
| "src/tiny_llm/continuation.py": "4f4143c828d4f94a851bcbf8a466d3e1b3e9eaec1027970a426208cc0ff4f4db", | |
| "src/tiny_llm/data.py": "1a52e38afef9af32463779cc2f471d9b7174a7c4c73753e74f485d526d3c6733", | |
| "src/tiny_llm/diffusion_diagnostics.py": "2ae667a076b1d3f057a29d20f00559068eb73a073b3f99030e1d5b4c2f478fa5", | |
| "src/tiny_llm/evaluation/__init__.py": "e49e80f7be3ab6e200f0db14b028e465de9caa8f6f705430eabce0c41b8322ba", | |
| "src/tiny_llm/evaluation/adapters.py": "8e64e5542bbcb0f9ad448b320e712916a9559823987e298b22321a8d8e2f727f", | |
| "src/tiny_llm/evaluation/catalog.py": "2fd29084ec780539023a38bcbfd7302931fa6be48624e227875f13518a1b986f", | |
| "src/tiny_llm/evaluation/cli.py": "a322e39bf58ba52132249dae524e9772c3a8e1a5ba388f5c0267f8dc1ea18882", | |
| "src/tiny_llm/evaluation/tasks/balanced_copa.yaml": "f47baa6a713e15f7a5600ed257f316f8c54252bf4977bbc716a6da7aa580f05d", | |
| "src/tiny_llm/evaluation/tasks/balanced_copa_test.yaml": "dce21ad70b395d8d576c584be1b87ebae846e360e59878030645215a387b28e6", | |
| "src/tiny_llm/evaluation/telemetry.py": "f4dad30574753739ea10ba71ef9a7ed9d9118a89a8202dee224961be27f13a48", | |
| "src/tiny_llm/evaluation/token_continuation.py": "e8d88cd4aca51ca2fda07f2b799a54612d0e2964a98e05824a135bea1a533550", | |
| "src/tiny_llm/losses.py": "592a05287369f6e5077b3faca404a9392de4772cf06389fb8aa7e36924e74f86", | |
| "src/tiny_llm/models/__init__.py": "b94ab167b6eb9f9b455823eea382e67b5ac2cc7f9ef5f6de183a09b9650b9718", | |
| "src/tiny_llm/models/diffusion/__init__.py": "c92eaf5986d604d8aaf1894a862a217cefa32ab25359175b88110e45a71c1b25", | |
| "src/tiny_llm/models/diffusion/configuration_diffusion_lm.py": "3b303aba60f0a831c966f5cfa7f5441daa2584814e335646195bd6ceecfa6d0c", | |
| "src/tiny_llm/models/diffusion/modeling_diffusion_lm.py": "4fc21c2aed1d5a231afe6f7634d9d8c3682b4e2abfc8fecef27b52136de77b2f", | |
| "src/tiny_llm/models/looped/__init__.py": "4b46b426021c59ed313cb2f49afc1926ad2793d771ef90117e3fef0462957cff", | |
| "src/tiny_llm/models/looped/configuration_looped_lm.py": "b8d99c52b98e0a97a4cae76ca4c20affd8f4d019ec38e0a9beefd5fad864dc25", | |
| "src/tiny_llm/models/looped/modeling_looped_lm.py": "9bdb3c5e3aaae663d09e98055f95df1adbc42e0a4bf5dffb7f6c39242af41b9e", | |
| "src/tiny_llm/models/moe_t5/__init__.py": "18e33fe38b2049521f573512ac8796afa209e1b25c2052784e1c927a7160d66c", | |
| "src/tiny_llm/models/moe_t5/configuration_alicet5_moe.py": "4647b1e9ab9ce07887bffc1abc3beb3e691962a380b4cc237aae376490bb56b7", | |
| "src/tiny_llm/models/moe_t5/modeling_alicet5_moe.py": "5f899c4726b120eed6c9bf6258ad319bb1ccea734645f27e7e6ccbc5057222ff", | |
| "src/tiny_llm/models/prefixlm/__init__.py": "c94c3c861acba8d2147a3a6c46bccfc71ac8adc8d70757bede009598029d31ea", | |
| "src/tiny_llm/models/prefixlm/configuration_prefix_lm.py": "20b3a62ae3628b5a64566ef19d2947c90f10c6159daad00dd13d18fc5910c42c", | |
| "src/tiny_llm/models/prefixlm/modeling_prefix_lm.py": "f62a38211df175696f64945173d1e41d9f7fcf6152a7648020da8c91788e98fd", | |
| "src/tiny_llm/models/q50m/__init__.py": "9d0dfec6fa0cfa6967e2f17a73d8fbeca69ba93b4a93651eda4d95626d5f5eb5", | |
| "src/tiny_llm/models/q50m/configuration_q50m.py": "796806b2d5d982b633796bc7de53c3073b1b6e52d259f44ef798b63f55eaea92", | |
| "src/tiny_llm/models/q50m/modeling_q50m.py": "b68abce1f19a36deb39c4ac4e520ffaab0c879d8a7cc2abef35f014bacc011f2", | |
| "src/tiny_llm/models/qmod/__init__.py": "e8c3e10710ac1cceac72ea329db480e680e8bbaae2a92ec26c36c92a778552ad", | |
| "src/tiny_llm/models/qmod/configuration_qmod.py": "0dfc9a4540697a9b15e423b9592df9fc435dcfc437ddad20dc6b6377ebebea40", | |
| "src/tiny_llm/models/qmod/modeling_qmod.py": "3831e34a86309b91a7a0787b9471eaceb91aa4c267f8030eb2e6722a0fcab6f8", | |
| "src/tiny_llm/models/qmoe/__init__.py": "cfb2bb24bc2f3d5b72b3890a393b0386cb25f127e375f5454255478a36dec32f", | |
| "src/tiny_llm/models/qmoe/configuration_qmoe.py": "bba7c9ac16f444146a78c99e7a26b4a11864b215d2b673bf658d9930f2f4de2f", | |
| "src/tiny_llm/models/qmoe/modeling_qmoe.py": "1d713e5e287f46dbf13b7cf9c562b4539efbf6231f7ceaa261f8cea1acd2768c", | |
| "src/tiny_llm/models/qt5/__init__.py": "df943b42a9f339e006687bcd30a666838c1dec15c79751b944e5f7a2bc69186c", | |
| "src/tiny_llm/models/qt5/configuration_qt5.py": "59709fe2c9fe08ccecea6633b29b3b47170fd30abcbee37952507d676446cd44", | |
| "src/tiny_llm/models/qt5/modeling_qt5.py": "6269829cca6631824622df034ff7b544dbbe47790556c60b26e356b5279cceff", | |
| "src/tiny_llm/models/t5/__init__.py": "e4d8838253800edbcac1f9a9b6dc875bcb3ffa2962a0c9044d066448ea10488e", | |
| "src/tiny_llm/models/t5/configuration_alicet5.py": "81b1f594139f201af9afb25ba58df862432c15dedf57f6e340ee2af80ed0d44f", | |
| "src/tiny_llm/models/t5/modeling_alicet5.py": "bd52ce0022668f2defb57770df229f148e070ea73becca9a2e681325114bfc20", | |
| "src/tiny_llm/models/zarya/__init__.py": "651b19cdba422680ff16bf2b2f63be0c3e9045c44c410fb17dbf51df2a23bd55", | |
| "src/tiny_llm/models/zarya/configuration_zarya_lm.py": "42a3e7a1b11b687bac34e00f9cd3cb59d1cd573eec7492e1c55834662d37ed3f", | |
| "src/tiny_llm/models/zarya/modeling_zarya_lm.py": "73ab033860288f7367d5b912c79de908fd81a4e88f369a658b96dd2bcfcd7bfe", | |
| "src/tiny_llm/telemetry.py": "7ee8f14f4545bfa0e5ac632440d49c7f86fd6812a16429f862ebe3cfe09e3545", | |
| "src/tiny_llm/ul2.py": "f4d948d89d5e106965612971dd473a54b57719f811761f18457c6221d763b876", | |
| "src/tiny_llm/zarya.py": "8b5084ab5086c75745115d3a9ba4f2fa9cfda8ad713f65d3f8a22e77db8c0941" | |
| }, | |
| "packages": { | |
| "lm_eval": "0.4.12", | |
| "torch": "2.11.0+cu128", | |
| "transformers": "5.17.0", | |
| "datasets": "5.0.1", | |
| "accelerate": "1.15.0" | |
| }, | |
| "git_commit": "fb65467388c90a7d94d50fbcb711abfd8a568102", | |
| "git_dirty": true | |
| } | |
| }, | |
| "tokenizer_sha256": { | |
| "tokenizer.json": "1d1f7409de3e53ae51d7085b3009aa5cd491dc08fbe114794c8d314f3fba2af4", | |
| "tokenizer_config.json": "879515efa0fe331af8eeb0ede00fdd13c0bd3c827bff8bb6f4b7f959c02159ec" | |
| }, | |
| "tensorboard_event_files": { | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866065.DESKTOP-OU3A33R.18959.0.eval": "7eea85c0d3f0054571717d24651da3cff3342a3028b07c4bfb24ab4923bacff8", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866080.DESKTOP-OU3A33R.18959.1.eval": "d13ad2fcca6784100fca5121ac1a8252e7b096b15022fd1cc9d163e645f93491", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866112.DESKTOP-OU3A33R.18959.2.eval": "49704476a69178e92c19346ef026115f71eeeda9a685c670a58efd75941ef15d", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866137.DESKTOP-OU3A33R.18959.3.eval": "9d2032c4df62dbd5b81ba7e6561ea413322ff5bbc9d56d03d990f0b30e768bce", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866174.DESKTOP-OU3A33R.18959.4.eval": "209958c83b2c6b092ce334f461c0f7ee00b8ff7266fbca818cccbaa147ca5090", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866180.DESKTOP-OU3A33R.18959.5.eval": "4e77ff5d52036d37678147b767e53bf51e8fde20583934d13b28d4f2cbd8d9b3", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790866494.DESKTOP-OU3A33R.18959.6.eval": "5cbfd8342dca93812e864b3c1789c0afebcb1f6e56338ae68a003394f9d946bd", | |
| "runs/moe_t5-gpu-v1/tb/events.out.tfevents.1790868998.DESKTOP-OU3A33R.20360.0.eval": "48e3f3505d57a16c34141e7758bd6d09979ba71dde024b1b34cefa63c3dbafaa" | |
| }, | |
| "tensorboard_verified_points": 1166 | |
| } | |
| } | |