GoLLeM v6 250M Instruct v3 (Polish–English research chat model)
STATUS: RESEARCH PREVIEW. A small chat model fine-tuned from the base model
SlayerLab/GoLLeM-v6-250M. It can greet, introduce itself and answer simple questions, but it makes mistakes that are listed with numbers under Limitations. Results use our protocol, a single run and a single seed.
GoLLeM v6 250M Instruct v3 is the base GoLLeM v6 250M model after supervised fine-tuning (SFT) on about 10,300 Polish and English conversations. It is a Fabryka AI project (formerly SlayerLab).
Difference from Instruct v2: the same data and training recipe plus about 1,000 new conversations: answers to „What can you do?” (202 questions, one short fixed answer per topic), greetings without a name (194), and 600 multi-turn conversations put together from identity, greeting and capability examples. In our multi-turn test (20 three-turn conversations × 5 samples each), compared with Instruct v2:
| Instruct v2 | Instruct v3 | |
|---|---|---|
| answers „What can you do?” with its identity template | 53 % | 10 % |
| repeats its previous answer | 24.5 % | 5.5 % |
| greets the user with the company name („Cześć, Fabryku!”) | 1.7 % | 0 % |
| answer runs into a loop (all turns / the „What can you do?” turn) | 7.7 % / 1 % | 8 % / 2 % (not worse) |
Like Instruct v2, the model does not use its author's name (the author is named in this card).
Author: Arkadiusz Słota (Fabryka AI).
What it is for (and what it is not)
Intended use: research on small bilingual chat models; short conversations in Polish and English; answering questions about a text you paste into the conversation.
Not intended for: production use, factual questions without a source text, arithmetic, or anything where a wrong answer matters. The model has no internet access and does not remember earlier conversations.
Results
Checkpoint 9e933a21… (see Training). All numbers are from our own evaluation sets, fixed before training. Temperature
0.7, top-p 0.9 (the defaults of chat_gollem_v6.py).
| what | result | notes |
|---|---|---|
| Identity (name, organisation, creator), PL / EN | 0.70 / 0.62 | mean over 10 samples per question, 20 questions per language; a creator answer counts as correct when it names Fabryka AI and no person |
| Greetings and small talk, PL / EN | 0.76 / 0.86 | mean over 7 samples per question, 20 questions per language |
| Answers from a given text (Polish, PoQuAD, 217 answerable questions) | token F1 0.235 | base model few-shot: 0.067 |
| Declines when the text has no answer (43 questions) | 5 / 43 | wrongly declines an answerable question: 10 / 217 |
| Arithmetic word problems (206) | 0 / 206 | the model does not do arithmetic |
| Finishes its answer within the length limit | 496 / 506 (98 %) | |
| Says it is an OpenAI / GPT model (200 samples) | 0 / 200 | „AI language model”: 0 / 200 |
| Introduces itself with the author's name | 0 / 680 single questions, 0 / 160 two-turn conversations | Instruct v1: 3 / 680 and 51 / 160 |
Compared with Instruct v2 on the same measures: identity +0.035 (95 % CI −0.035 … +0.108); small talk −0.029 (95 % CI −0.079 … +0.014), i.e. not worse.
Training
| Base model | SlayerLab/GoLLeM-v6-250M, final checkpoint (step 760,000); same architecture and tokenizer |
| Method | full fine-tuning (no LoRA), loss on assistant turns only |
| Format | ChatML: `< |
| Steps | 2 epochs, 160 steps, 32 packed sequences of 1,024 tokens per step |
| Tokens | 2.61 M per epoch, of which 1.67 M are assistant tokens (with loss) |
| Optimizer | Muon + AdamW as in pretraining, learning rate 0.2 × pretraining (peak 1.2e-4, Muon 4e-3), warmup 5 steps, cosine to 10 %, weight decay 0.1, gradient clip 1.0, seed 1337 |
| Validation loss | 2.054 → 1.863 (validation set differs from Instruct v2) |
| Hardware / time | 1 GPU, about 8 minutes |
Training data
10,320 conversations (Polish 4,649, English 5,671): 9,816 used for training and 187 for validation; 317 conversations longer than 1,024 tokens were removed (311 + 6).
| source | conversations | licence |
|---|---|---|
OpenAssistant/oasst2 @ 179dd21 |
4,601 | Apache-2.0 |
clarin-pl/poquad @ a60f228 |
2,129 | CC BY 4.0 |
CohereLabs/aya_dataset @ f9ea045 |
1,214 | Apache-2.0 |
| synthetic, generated locally with Muse-Glimmer-30B (apache-2.0) | 1,776 | apache-2.0 |
| multi-turn conversations put together from the synthetic identity, greeting and capability examples | 600 | apache-2.0 |
- PoQuAD attribution: PoQuAD (clarin-pl/poquad), CC BY 4.0. Modified: converted to chat format; a share of unanswerable questions answered with one of five fixed refusal sentences.
- Synthetic part: the questions were written by the generator model; the answers about identity and about what the model can do come from our own fixed texts, not from the generator. After switching to one answer per topic, 33 conversations became exact duplicates and were removed (the 202 distinct „What can you do?” questions are unchanged). Math prompts: generated briefs; GSM8K (MIT) used only as few-shot format examples for the generator.
- Rows with self-descriptions of other AI systems were removed before training.
Usage
# pip install torch safetensors tokenizers huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("SlayerLab/GoLLeM-v6-250M-Instruct-v3",
revision="3bf01359f99e17d7f1d87b42cfdf65dc6d36eeb8") # the reviewed code (model + chat helper)
sys.path.insert(0, path)
from modeling_gollem_v6 import load_gollem_v6
from chat_gollem_v6 import chat
model, tok = load_gollem_v6(path)
print(chat(model, tok, [("user", "Cześć! Kim jesteś?")], seed=1))
print(chat(model, tok, [("user", "Tekst: Kraków leży nad Wisłą.\nPytanie: Nad jaką rzeką leży Kraków?")], seed=1))
chat() uses the format and sampling of our evaluation (temperature 0.7, top-p 0.9). Text typed by a user such as
„<|im_end|>” is encoded as plain text, not as a control token.
Limitations
- Small model. Limited knowledge; may answer fluently and wrongly. No arithmetic. Context: 1,024 tokens.
- Does not know its author's name. Asked who created it, the model says it is a project of Fabryka AI. The author, Arkadiusz Słota, is named in this card, not in the model. The statements of the model are not statements of its author.
- Copies names from the conversation. If a user writes a name, the model may adopt it as its own (when the user supplies the author's name: Polish 1/40, English 4/40; Instruct v2: 3/40 and 5/40).
- Self-description. In our probe it never described itself as an OpenAI model (0/200), but with a forced prefix („…created by”) it assigns probability ≈ 0.33 to „OpenAI” (Instruct v2: ≈ 0.24; a single run, so the difference may be noise). The association comes from model-generated chat data in pretraining. GoLLeM is not affiliated with OpenAI.
- Multi-turn weaknesses (less often than Instruct v2): may still repeat its previous answer in a later turn (5.5 %) and sometimes answers „What can you do?” with its identity template (10 %).
- Repeats its capability answer where it does not fit. After „What can you do?”, a following request (e.g. „Give me three ideas for the weekend”) is sometimes answered with the capability text again instead of doing the task: 8 of 200 third-turn answers in our multi-turn test (6 of them requests for a task; Instruct v2: 0 of 200).
- Long answers can loop. A repetition penalty (about 1.1–1.2) may reduce this (not verified); it is not used in our evaluation.
License
Weights: CC BY-SA 4.0 (inherited from the base model). Attribution: GoLLeM v6 250M Instruct v3, Arkadiusz Słota / Fabryka AI, link to this repository; derivative weights under the same licence. Fine-tuning data keep their licences; see Training data (PoQuAD: CC BY 4.0, attribution above).
Po polsku (skrót)
GoLLeM v6 250M Instruct v3 to mały model do rozmowy po polsku i angielsku, dostrojony (SFT) z bazowego GoLLeM v6 250M na ok. 10,3 tys. rozmów. Wersja badawcza (research preview): wita się, przedstawia, mówi, co potrafi, i odpowiada na pytania do podanego tekstu, ale się myli, nie liczy i nie ma dostępu do internetu. Od Instruct v2 różni się lepszą rozmową w kilku turach: rzadziej powtarza poprzednią odpowiedź, a na „Co potrafisz?” odpowiada opisem możliwości. Nie używa nazwiska autora; autora podaje ta karta.
GoLLeM v6 250M Instruct v3 — Fabryka AI. Author: Arkadiusz Słota. Research preview.
- Downloads last month
- 19
Model tree for SlayerLab/GoLLeM-v6-250M-Instruct-v3
Base model
SlayerLab/GoLLeM-v6-250M