Bielik-1.5B-v3 · matura-history SFT v2 ("Mały, ale wariat")

Full-parameter SFT of speakleash/Bielik-1.5B-v3.0-Instruct on synthetic tasks and model essays in the format of the Polish extended-level history matura (CKE). Built for the "Mały, ale wariat" track of the SlayerLab hackathon: the smallest model file that scores at least 35% on the team's history-matura evaluator.

Registered candidate: sft2-IQ4_XS.gguf, 868,041,888 bytes (0.87 GB). It is the smallest file that passes the team evaluator, when run through the matura-harness v5 retrieval setup. Read the caveats below before quoting the number. The 4.5B sibling that also passes an unseen 2026 paper is SlayerLab/bielik-4.5b-v3-matura-history-sft2.

This repo holds the bf16 safetensors (Transformers format), the llama.cpp GGUF quants made with a Polish importance matrix, the imatrix itself, and the training logs. Code, data recipe and every evaluation run: slayerlabs/hackathon → maly-ale-wariat/ (MODEL_CARD.md, RESULTS.md, REPORT.md).

Files

File Bytes Notes sha256
model.safetensors (+ config.json, tokenizer, chat_template.jinja) 3,193,073,112 bf16, Transformers format 89a2b0b4230b74dc25a2b82968c821ed90ad6a88f6bd556681eebce9bf700db6
sft2-bf16.gguf 3,195,508,608 convert_hf_to_gguf.py; the source of every quant below fd287cada04bbad137711cf98d77d799aada9e6a24ad41d2013e5147b1366de7
sft2-Q8_0.gguf 1,699,567,776 PPL 17.18 8aeef74ef9ec054f7897b68f8397f4acbebc5ec3e9f8c33963cb2325df3c7f2f
sft2-Q4_K_M.gguf 972,797,184 PPL 17.36. Passes the evaluator under both protocols 8ba1d092523b3af347a278bd99ba1ef5cf72805af4db22e3d817d286c46ada2f
sft2-IQ4_XS.gguf 868,041,888 PPL 17.22. Registered candidate 5fbe1f31d32b0a46704acfc4ab2bffe2aba46c23007ed9a48e57cc62723ef5ee
sft2-Q3_K_M.gguf 782,735,520 PPL 17.85. Fails the past-exam track 6697a8c44c65b62e0e7e5d43fc7b24ada260927ca6bae2c6266391ffcf2916c5
sft2-MIX-IQ4XS-FFNUG-IQ3S.gguf 778,585,248 Experiment: IQ4_XS with ffn_up/ffn_gate at IQ3_S. PPL 20.48 bf175696273ede1dcdac76b3f18af8acc02b8e0fbf41e535d771c649b51853e9
sft2-IQ3_M.gguf 728,017,056 PPL 22.50. Fails 74f8d360c64a11e4d5b3ed4dc5d09cff58add910402cc6c7c8cb75713c0d063e
sft2-Q3_K_S.gguf 709,007,520 PPL 21.11 b9e5526ec017378baa63389ae2f862f63f25ddcd94c62f32dda001e7a15133fc
sft2-IQ3_XXS.gguf 632,585,376 PPL 22.26 ed8553f24ce1c24d0fc716cf4a209fc559a87c838af76de0c2be8cb6e15b4e2f
sft2.imatrix.gguf 2,361,056 Polish importance matrix: SFT conversations + plwiki passages (train/make_calib.py, calib-v1.txt) 18c2bc1871db633630041feacdcbde69b36c1439dea0180abc79a83438a51aa2
training/ train_args.json, metrics.json, log.jsonl, data_stats.json, gpu.txt, out-pip-freeze.txt
SHA256SUMS sha256 of every file in this repo (sha256sum -c SHA256SUMS)

PPL = perplexity on held-out Polish text (calib/ppl-heldout.txt, 40 × 1024 tokens); lower is better. IQ4_XS with the Polish imatrix is effectively lossless against Q8_0; every 3-bit variant failed the evaluator. Quants were made with llama.cpp build 11146 (commit 7fe450e19) via train/quantize.sh.

How to run

The configuration that was scored (llama.cpp llama-server, OpenAI-compatible API on port 8891):

hf download SlayerLab/bielik-1.5b-v3-matura-history-sft2 sft2-IQ4_XS.gguf --local-dir .
llama-server -m sft2-IQ4_XS.gguf --port 8891 --ctx-size 8192 -ngl 99 --flash-attn on --jinja \
  --chat-template-kwargs '{"enable_thinking":false}'

llama-server -hf SlayerLab/bielik-1.5b-v3-matura-history-sft2:IQ4_XS and ollama run hf.co/SlayerLab/bielik-1.5b-v3-matura-history-sft2:IQ4_XS also work.

The reported scores used matura-harness in front of that server with the flags in maly-ale-wariat/eval/harness_args_v5.txt (evaluator-style prompts for source-based tasks, plwiki BM25 retrieval of 2 × 700 characters only for tasks of 600 characters or less and for essays, strict closed-answer normalization, repetition trimming, DRY sampling and a 380-word floor for essays). The 1.7 GB retrieval index is not part of the model; the steps to rebuild it are in the repo README.

Transformers (bf16 safetensors):

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "SlayerLab/bielik-1.5b-v3-matura-history-sft2"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")

messages = [
    {"role": "system", "content": "Jesteś zdającym maturę rozszerzoną z historii. Odpowiadasz po polsku, zwięźle i konkretnie. "
                                  "Najważniejsze są źródła podane w zadaniu; materiały pomocnicze mogą być nietrafne. "
                                  "Wykonaj dokładnie polecenie i podaj wymaganą liczbę elementów."},
    {"role": "user", "content": "Podaj nazwę bitwy z 1410 r., w której wojska polsko-litewskie pokonały zakon krzyżacki, "
                                "i wymień jednego z dowódców strony polsko-litewskiej."},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=200, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

The system prompt above is the one used in training. Chat template: Bielik ChatML (<s><|im_start|>role\n…<|im_end|>), no thinking mode. Context: 8192 tokens.

Evaluation

Team evaluator history-matura-eval: two tracks of 55 points each, 34 items adapted from the CKE May 2023 extended-level history paper ("past exam") and 34 Wikipedia-based items, graded by a deterministic closed-answer scorer plus an LLM judge (google/gemini-3.1-flash-lite). Pass = at least 20/55 (35%) on each track. Two protocols:

  • raw: the evaluator's bare-model runner (its system prompt, the task text only, the answer submitted verbatim);
  • harness v5: matura-harness with eval/harness_args_v5.txt, as described above.

A third, truly unseen paper, CKE May 2026 (39 items, 60 points, pass = 21), was converted to the evaluator's item format and scored with the same scorer and judge. Scores below are strict re-scores (eval/strict_rescore.py); the value the evaluator reported is in brackets where it differs.

GGUF Size Protocol Past exam (/55) Wikipedia (/55) CKE May 2026 (/60)
sft2-IQ4_XS 0.87 GB harness v5 23 (25) ✅ 31 ✅ 19 ✗
sft2-IQ4_XS 0.87 GB raw 18 ✗ 26 ✅ 19 ✗
sft2-Q4_K_M 0.97 GB harness v5 22 ✅ 28 ✅ 14 ✗
sft2-Q4_K_M 0.97 GB raw 25 ✅ 28 ✅ 15 ✗
sft2-Q3_K_M 0.78 GB harness v5 incomplete¹ 25 ✅ –
sft2-IQ3_M 0.73 GB harness v5 17 ✗ 24 ✅ –
base Bielik-1.5B-v3.0-Instruct Q8_0 1.70 GB harness v5 17 ✗ 35 ✅ 19 ✗
base Bielik-1.5B-v3.0-Instruct Q8_0 1.70 GB raw incomplete¹ 26 ✅ 20 ✗

¹ The evaluator reports no total when the judge errs or grades an item "review".

What the SFT changed: bare Bielik-1.5B scores 15–17/55 on the past exam, just under the threshold. SFT on about 4k synthetic and CKE-style tasks lifts past-exam open answers from 7–8 to 14–16 points, which is what makes the model pass.

Caveats

  • Judge glitch (+2) in the headline run. On one 3-statement true/false item the model answered only "P"; the judge wrote "incomplete" but filled the full correct key, and the scorer counted 2 points. The corrected past-exam score is 23/55 (41.8%), still a pass with a 3-point margin.
  • The 0.87 GB file passes only through the harness. Under the evaluator's bare runner it scores 18/55 on the past exam. The harness adds points through retrieval on short tasks, essay handling and strict closed-answer formatting. A dashboard that runs the bare GGUF will show this file failing; the 0.97 GB Q4_K_M passes under both protocols.
  • A single evaluator run varies by about ±3 points; the judge is an LLM and a small model's answer can flip between near-identical prompts.
  • On the unseen CKE May 2026 paper every 1.5B variant, the untouched base included, lands at 19–21/60, right at the 35% line. The fine-tuning gain shows on the evaluator's paper, not on this one.
  • Known failure modes: true/false statements at chance level, open answers naming the wrong person/place/date, essays that hallucinate facts.

Training

Base speakleash/Bielik-1.5B-v3.0-Instruct (Apache-2.0), loaded from the ungated mirror cpral/Bielik-1.5B-v3.0-Instruct-ungated (revision a3a660b1). Its model.safetensors sha256 (3c337d1d…c28978) equals the LFS sha256 of the gated original, so the base is verified identical.
Method Full-parameter SFT (train/sft.py, plain transformers 4.57.1), loss only on assistant tokens.
Data 4,144 chat rows (sft/v2, prompt version compact-v2) + 116 validation rows: 2,714 synthetic syllabus tasks, 300 model essays (×2), 800 supplementary items, and 82 items converted from the CKE May 2024/2025 papers (×2; not redistributed, CKE copyright). 66% compact prompt with 0–2 retrieved plwiki passages, 20% evaluator-style bare prompt. 2.80M training tokens, 576k of them supervised.
Hyper-parameters 3 epochs, lr 1e-5 cosine to 0.1×, warmup 3%, weight decay 0, batch 8 × grad-accum 4 (effective 32), max length 4096, bf16, gradient checkpointing, seed 42.
Compute 1 × NVIDIA L40S (RunPod), 390 steps, 1,020 s (17 min).
Loss train 1.344, eval 1.413 (final; curves in training/log.jsonl).

Things that did not help (see RESULTS.md): continued pretraining on 39M plwiki tokens, more real CKE papers, 4 epochs at lr 2e-5, a uniform model soup, reason-first prompting or self-consistency voting for closed questions.

Data provenance and licensing

  • Weights: Apache-2.0, same as the base model.
  • The synthetic tasks and essays were written by Claude (Anthropic) via Claude Code workflows following data/gen/STYLE.md, fact-checked by a second agent pass and validated by data/gen/validate_items.py. They are CC BY-SA 4.0-compatible.
  • Retrieved evidence comes from Polish Wikipedia (CirrusSearch dump 2026-09-20, CC BY-SA 4.0); every record keeps its source URL.
  • The CKE 2024/2025 items are not redistributed (CKE copyright). The evaluator's questions and keys were never used for training; build_sft.py drops any record sharing word 8-grams with them, plus essays on the held-out themes.

Intended use and limitations

A hackathon research artifact: a small Polish model that answers exam-style history questions in the CKE matura format. It hallucinates names, dates and places, answers true/false items near chance, and was tuned for a specific evaluator. Do not rely on it for factual information or grading.

Acknowledgements

SpeakLeash for Bielik-1.5B-v3.0-Instruct; llama.cpp for quantization and inference; Polish Wikipedia authors (CC BY-SA 4.0).

Downloads last month
960
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SlayerLab/bielik-1.5b-v3-matura-history-sft2

Finetuned
(8)
this model