Slayer139 1.01
A small decoder-only English language model (139.3M parameters) trained from scratch by Fabryka AI, built for the Glint Tiny-ML leaderboard (≤150M parameters).
Author: Arkadiusz Słota (Fabryka AI / SlayerLab). Training, evaluation and safety gates were run by the Fabryka AI agent team (Kolektyw) under his direction.
Result
lm-evaluation-harness 0.4.13, board protocol (C2), our measurement on 1× RTX 5090. Our two evaluation machines (RTX 5090, AMD Radeon) gave identical scores, component by component, on a shared calibration checkpoint.
| measure | Slayer139 1.01 | 95% CI | pretrained base (step 90,310) | 95% CI | Δ (paired 95% CI) |
|---|---|---|---|---|---|
| BLiMP (67 tasks) | 79.13 | [78.85; 79.40] | 79.41 | [79.14; 79.69] | −0.28 |
| ARC-Easy (test, 2,376) | 63.09 | [61.15; 65.07] | 57.58 | [55.60; 59.55] | +5.51 |
| WikiText-2 score | 99.83 | 99.86 | −0.03 | ||
| Overall | 80.68 | [80.03; 81.34] | 78.95 | [78.28; 79.62] | +1.73 [+1.25; +2.21] |
Overall is the mean of the three measures. Slayer139 1.01 is the pretrained base fine-tuned on ARC-Easy train (see Training); both were evaluated on the same machine.
Confidence intervals: item-level bootstrap (ARC-Easy and each of the 67 BLiMP tasks resampled separately, 10,000 draws); the Δ column uses the same resampled items for both models (paired). They cover test-set sampling only, not seed-to-seed variation.
Train–test overlap check: 59 ARC-Easy test questions also appear in ARC-Easy train. Without them (2,317 questions) ARC-Easy is 63.19 for Slayer139 1.01 and 57.70 for the base: a gain of +5.48, against +5.51 on the full test, so the gain does not come from the overlapping questions.
Provenance (sha256 prefixes of the lm-eval results.json): Slayer139 1.01 ed9e9ba3, base fabe34ce; per-item samples 00bd9ddb / 19f78648.
Numbers are our measurements with the board's protocol, not official scores until the leaderboard lists them.
Model
- Architecture: 17 layers × 768, 12 heads (12 KV heads), SwiGLU, decoder-only; vocabulary 24,576 (BPE), context 1,024.
- Parameters: 139,279,760 unique parameters (input and output embeddings tied: one shared tensor, counted once; counting the shared tensor twice gives 158,154,128).
- Checkpoint: the pretrained model (step 90,310, the end of the run, sha256
ba223162…) fine-tuned on ARC-Easy train for 5 epochs (sha256aa4cabed…). Both choices follow rules written before any test number: the end of the run rather than a selected step; checkpoint averaging was tested on a 64M control model and not adopted (+0.22/+0.24 Overall, below the pre-registered +0.3 threshold); the fine-tune was adopted only because it passed a pre-registered acceptance rule (Overall gain ≥ +0.5 with a paired CI excluding zero, BLiMP and WikiText not degraded beyond fixed limits, ARC gain unchanged without the test questions that overlap train).
Training
- Data: ~23.67B tokens, no repeated epoch (sources and licences below). First the ARC-MIX pool, then a top-up from the same source families in the base corpus's proportions, mixed at document level. Both stages were scanned against WikiText-2 (test, validation) and ARC-Easy/Challenge (validation, test) for shared normalized 13-grams, with short ARC questions matched exactly, and against BLiMP sentences of ≥ 6 words; matching documents were removed. ARC train was not filtered.
- Optimiser and schedule: Muon (hidden matrices, LR 0.0566) + AdamW (other parameters, LR 1.6971e-3), both ×√8 from the single-node recipe; WSD schedule: 250 warm-up steps, constant to step 72,248, then 1−√ decay to 10% of peak at step 90,310; batch 256 × 1,024 tokens; seed 1337.
- Same recipe as our single-node final (batch 32, 4× RTX 5090, ~191k tokens/s), scaled to batch 256 with LR ×√8 and the identical data order — 4.4× faster. At equal tokens the 8-GPU run had lower held-out DCLM bits-per-byte at every compared checkpoint (−0.014 at 1.1B tokens, −0.032 at 19.4B tokens).
- Hardware and time: 8× NVIDIA RTX 5090 (data parallel, bf16 all-reduce), median ~832k tokens/s, about 8 hours of training.
- ARC-Easy fine-tune (after pretraining): listwise cross-entropy over the summed log-probabilities of the answer tokens, prompt
Question: {q}\nAnswer:/{a}(the lm-evaluation-harness format), AdamW LR 5e-6, betas (0.9, 0.95), weight decay 0.01, 8 questions per step, up to 5 epochs, bf16 forward. Data: ARC-Easy train, 2,174 questions after removing 77 that overlap the test set. The epoch was chosen on ARC-Easy validation only (570 questions; accuracy 0.579 before, 0.637 after epoch 5); the test split was not read during fine-tuning. Under 3 minutes on one RTX 5090. Recipe after SlayerLab/Slayer149-Balanced-ARC-E-ft-45500. - Training curves: https://track.fabryka.ai/run/5a4d174e-4a9d-4cc1-9602-838a3ffbb917
Training data
Aggregate corpus; each source keeps its upstream terms.
- ARC-MIX pool (
SlayerLab/gollem-v5-arcmix-9b, 9.39B tokens before retokenization): the 15-source baseSlayerLab/minimal-en-corpus-5b(FineWeb-Edu, DCLM-baseline, open-web-math, FineMath, StarCoderData, StackExchange, LoC public-domain books, Wikipedia, Project Gutenberg, scientific papers, UltraChat, WildChat, CC-News, tiny-textbooks, OpenSubtitles) + a FineWeb-Edu expansion + OpenStax textbooks (×4) + extra copies of FineWeb-Edu documents scored as related to ARC science topics (+2 copies for the highest score, +1 for the next). Removed before training: 4,362 documents by the WikiText-2 / ARC validation+test scan and 1,175 documents marked CC BY-NC-SA. - Top-up (11.4M documents, 57.45 GB of text, build manifest
b73f45f5…): FineWeb-Edu expansion (37.1%, ODC-By 1.0, plus the same dose of ARC-related copies as the pool), FineWeb-Edu (14.9%, ODC-By 1.0), DCLM-baseline (10.9%, CC BY 4.0), StackExchange via RedPajama (6.1%, CC BY-SA content), LoC public-domain books (5.4%, CC0), StarCoderData (5.1%, The Stack terms), Wikipedia 20231101.en (4.6%, CC BY-SA 3.0 / GFDL), Project Gutenberg (2.9%, US public domain), PMC Open Access commercial-use subset (2.8%, CC0 / CC BY / CC BY-SA per article; ND and NC excluded), open-web-math (2.7%, ODC-By), FineMath (2.6%, ODC-By), WildChat-1M (2.0%, ODC-By; English, non-toxic), CC-News (1.4%, licence unknown), UltraChat 200k (0.7%, MIT), tiny-textbooks (0.7%, Apache 2.0). The top-up scan removed 25,808 documents (0.23%), each together with its ARC-related copies. - Licences upstream include ODC-By, CC BY 4.0, CC BY-SA (Wikipedia, StackExchange, some PMC articles), CC0 / public domain, and sources with unknown or restrictive terms. Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData and model-generated chat data (UltraChat, WildChat, tiny-textbooks).
- OpenStax textbooks (55 titles) are used under CC BY 4.0; attribution: OpenStax, Rice University (see OpenStax attribution below).
- Fine-tune data: ARC-Easy train and validation from
allenai/ai2_arc(revision210d026faf9955653af8916fad021475a3f00453; Clark et al., 2018, Allen Institute for AI), CC BY-SA 4.0. Train: 2,174 questions after removing 77 that overlap the test set; validation (570) used only to choose the epoch; test not used.
Limitations
- English only; small model; one seed.
- The pretraining mix deliberately upweights web documents related to ARC science topics (decontaminated against the ARC test set); ARC-Easy is therefore an in-domain benchmark for this model.
- Fine-tuned on ARC-Easy train (2,174 questions after removing 77 that overlap the test set). ARC-Easy test is also reported without the 59 test questions that overlap train; the gain is the same within 0.03 points.
OpenStax attribution
The training data of this model includes text extracted from the following OpenStax textbooks, each licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), © Rice University. Download for free at https://openstax.org. The texts were obtained from the Hugging Face dataset crumb/openstax-text (revision 8f502ca45f9f05cb5673eae445b7a97a4e8c4349). Modified: text extracted, chunked and filtered (chunks matching benchmark test/validation sets were removed).
| Title (from source file name) | © year |
|---|---|
| APBiology | 2018 |
| APCollege Physics | 2017 |
| APMacroeconomics 2e | 2017 |
| APMicroeconomics 2e | 2017 |
| Algebra and Trigonometry 2e | 2021 |
| American Government 3e | 2021 |
| Anatomy and Physiology 2e | 2022 |
| Anatomyand Physiology | 2017 |
| Astronomy 2e | 2022 |
| Astronomy | 2018 |
| Biology 2e | 2020 |
| Business Ethics | 2018 |
| Chemistry 2e | 2019 |
| Chemistry Atoms First 2e | 2019 |
| College Algebra 2e | 2021 |
| College Algebra Corequisite Support 2e | 2021 |
| College Physics | 2020 |
| College Physics 2e | 2022 |
| College Physics for AP Courses 2e | 2022 |
| College Success | 2020 |
| College Success | 2023 |
| Concepts Biology | 2017 |
| Contemporary Mathematics | 2023 |
| Economics 2e | 2018 |
| Economics 3e | 2022 |
| Elementary Algebra 2e | 2020 |
| Entrepreneurship | 2020 |
| Intermediate Algebra 2e | 2020 |
| Introduction to Intellectual Property | n/a |
| Introduction to Philosophy | 2022 |
| Introduction to Political Science | 2022 |
| Introductionto Anthropology | 2022 |
| Introductionto Sociology 3e | 2021 |
| Introductory Business Statistics | 2018 |
| Introductory Statistics | 2018 |
| Macroeconomics 2e | 2018 |
| Macroeconomics 3e | 2022 |
| Microbiology | 2021 |
| Microeconomics 2e | 2018 |
| Microeconomics 3e | 2022 |
| Physics | n/a |
| Prealgebra 2e | 2020 |
| Precalculus 2e | 2021 |
| Preparing for College Success | 2023 |
| Principles Marketing | 2023 |
| Principlesof Finance | 2022 |
| Psychology 2e | 2020 |
| Statistics | n/a |
| USHistory | 2021 |
| University Physics Vol 1 | 2021 |
| University Physics Volume 2 | 2021 |
| University Physics Volume 3 | 2021 |
| World History Volume 1 | 2023 |
| World History Volume 2 | 2022 |
| Writing Guide | 2021 |
Titles licensed CC BY-NC-SA 4.0, non-English titles, and three CC BY 4.0 titles whose text contains elements marked ‘CC BY-NC-SA’ (Introduction to Business, Organizational Behavior, Principles of Management) were not used.
Reproduce
Evaluation: lm-evaluation-harness 0.4.13, tasks blimp, arc_easy, wikitext (board protocol C2); per-item samples are kept for paired bootstrap. Method and full run history: paper „Same recipe, 4.4× faster on 8× RTX 5090” (Fabryka AI, in preparation).
- Downloads last month
- 18