LLMPretrain — a 352M decoder-only LM trained from scratch

Every weight in this model was produced by this project. There is no from_pretrained, no base checkpoint, and no adapter: the tokenizer, the architecture code, the data pipeline and the training loop are all bespoke, and training started from a random initialisation.

Trained on a single RTX 5070 Ti (16 GB).

Architecture

Written by hand — no transformers model classes were used to define it.

Type decoder-only, pre-norm transformer
Parameters 351,650,304
Layers 18
d_model 1280
Attention heads 20 query / 4 KV (GQA)
Head dim 64
FFN SwiGLU, hidden 3456 (≈ 8/3 × d_model)
Normalisation RMSNorm (pre-norm) + QK-norm
Positions RoPE (θ=10000) — no learned position embeddings
Context 1024
Vocab 32,768 (own byte-level BPE, padded to a multiple of 128)
Biases none, in any linear layer
Embeddings tied between input and LM head

Parameter split: 41,943,040 embedding (11.9%) / 309,705,984 transformer blocks.

Training

Tokens seen 43,779,672,064
Steps 83,500
Schedule warmup-stable-decay (WSD)
Optimizer fused AdamW, betas (0.9, 0.95), wd 0.1 on 2D params only
Precision bf16 autocast, fp32 master weights
Hardware 1 × RTX 5070 Ti 16 GB

Evaluation

Bits per byte is the headline metric. This model's vocabulary (32,768) differs from GPT-2's (50,257), and cross-entropy per token is not comparable across tokenizers. BPB divides that out and is directly comparable.

Metric This model GPT-2 124M
Parameters 351,650,304 124,439,808
Validation loss 1.4474 —
Bits per byte 0.4905 —
HellaSwag (0-shot, len-norm) pending 29.55%

Validation loss is measured on a held-out slice of the training corpus and is not comparable to GPT-2's OpenWebText number. Bits per byte is, because it is tokenizer-independent.

Data provenance and licensing

Source HF id License
fineweb-edu HuggingFaceFW/fineweb-edu ODC-By 1.0

Attribution: FineWeb-Edu (HuggingFaceFW/fineweb-edu), ODC-By 1.0

ODC-By requires attribution, which is given above. Whether pretraining-data licenses flow through to model weights is legally unsettled; permissively licensed sources were chosen deliberately as the most defensible posture. Model weights are released under cc-by-nc-4.0.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("Dikshan1234/LLMPretrain")
model = AutoModelForCausalLM.from_pretrained("Dikshan1234/LLMPretrain",
                                             trust_remote_code=True)

inputs = tok("Once upon a time", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=100, do_sample=True,
                     temperature=0.8, top_p=0.95)
print(tok.decode(out[0]))

trust_remote_code=True is required because the architecture (RoPE + SwiGLU + GQA + QK-norm) is not a stock transformers model. The two files it loads, configuration_llmpretrain.py and modeling_llmpretrain.py, are in this repo and are short enough to read.

Sample generations

===== step 83500 =====

--- [greedy] 'Once upon a time'
Once upon a time, the world was a very different place. The world was a very different place. The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The world was a very different place. The world was a very different place.
The

--- [t=0.8] 'Once upon a time'
Once upon a time, I lived with a family of four – a girl and a boy, with two and a husband. They had six children but only one survived infancy; the last, who was only a year old. I was still trying to understand why this family was so poor and this child was so sick and so poor. I looked for a way to solve the problem.
In my search, I found out that there were many ways to solve the problem of poverty, and that the solution to the problem of poverty was not just to give food to the children, but also to educate them.
I

--- [greedy] 'The capital of France is'
The capital of France is divided into 100,000,000 shares of common stock, par value $0.0001 per share, and 100,000,000 shares of preferred stock, par value $0.0001 per share. As of December 31, 2019, there were approximately 1,000,000 shares of common stock issued and outstanding and no shares of preferred stock issued and outstanding.

The holders of the common stock are entitled to one vote per share on all matters to be voted upon by the stockholders of the Company. The common stock does not have cumulative voting rights.

The holders of the common stock

--- [t=0.8] 'The capital of France is'
The capital of France is divided into four provinces: the north contains Alsace and Lorraine, with the northern border forming the border between France and Germany; the south contains Poitou, Picardy, and Auvergne. In the east is the Côte d'Azur, the area between the Rhône and the Loire rivers.
France's population grew from 2.3 million in the year 1900 to 4.4 million in the year 2000, making it the second most populous country in Europe. It is the largest country in Europe and the European Union. The

--- [greedy] 'In 1969, humans first'
In 1969, humans first began to use the term “mammoth” to describe a large mammal, the mammoth. The term was first used in a scientific paper by the American paleontologist and paleontologist, John Ostrom, in which he described the mammoth as a “megafauna” (a large mammal) that “existed in the Pleistocene era, and was the largest land mammal ever to live.”
In the early 20th century, the term “mammoth” was used to describe a large mammal, the mammoth, that was the largest land mammal ever to live.

--- [t=0.8] 'In 1969, humans first'
In 1969, humans first encountered the world of the deep in the ocean floor. The first ever successful exploration of the deepest parts of the ocean went down with the recovery of the Challenger Deep on December 14, 1953. In July 1998, the last U.S. submarine, the Deepwater Horizon, sank with all of its crew.
The first deep-sea exploration was the Dutch ship Rotterdam, in 1929, which explored the Challenger Deep, the deepest part of the ocean. Other countries were involved in an initial survey of the region. In April 1931, the British government awarded rights to another British

--- [greedy] 'The three states of matter are'
The three states of matter are distinguished by their ability to conduct certain types of interactions with each other and with the environment. The states of matter are distinguished by their ability to conduct certain types of interactions with each other and with the environment.
The three states of matter are distinguished by their ability to conduct certain types of interactions with each other and with the environment.
The three states of matter are distinguishe

Limitations

This is a base model trained on a modest token budget. It has had no instruction tuning, no RLHF and no safety alignment. It will produce confidently wrong text, and it reflects whatever biases exist in its training corpus. It is a research and educational artifact, not a product.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Dikshan1234/LLMPretrain