daviddrzik commited on
Commit
d5dfc8c
·
verified ·
1 Parent(s): ce32ffd

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +103 -1
README.md CHANGED
@@ -2,4 +2,106 @@
2
  license: mit
3
  language:
4
  - sk
5
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  license: mit
3
  language:
4
  - sk
5
+ datasets:
6
+ - oscar-corpus/OSCAR-2109
7
+ pipeline_tag: fill-mask
8
+ library_name: transformers
9
+ ---
10
+
11
+ # Slovak BPE Baby Language Model (SK_BPE_BLM)
12
+
13
+ **SK_BPE_BLM** is a pretrained small language model for the Slovak language, based on the RoBERTa architecture. The model utilizes standard Byte-Pair Encoding (BPE) tokenization and is case-insensitive, meaning it operates in lowercase. While the pretrained model can be used for masked language modeling, it is primarily intended for fine-tuning on downstream NLP tasks.
14
+
15
+ ## How to Use the Model
16
+
17
+ To use the SK_BPE_BLM model, follow these steps:
18
+
19
+ ```python
20
+ from transformers import pipeline, RobertaTokenizer, AutoModelForMaskedLM
21
+
22
+ # Load the custom tokenizer and model
23
+ tokenizer = RobertaTokenizer.from_pretrained("daviddrzik/SK_BPE_BLM")
24
+ model = AutoModelForMaskedLM.from_pretrained("daviddrzik/SK_BPE_BLM")
25
+
26
+ # Create a pipeline with the custom model and tokenizer
27
+ unmasker = pipeline('fill-mask', model=model, tokenizer=tokenizer)
28
+
29
+ # Use the pipeline
30
+ result = unmasker("včera večer sme <mask> nový film v kine, ktorý mal premiéru iba pred týždňom.")
31
+ print(result)
32
+
33
+ [{'score': 0.2665567100048065,
34
+ 'token': 18599,
35
+ 'token_str': ' pozreli',
36
+ 'sequence': 'včera večer sme pozreli nový film v kine, ktorý mal premiéru iba pred týždňom.'},
37
+ {'score': 0.23860174417495728,
38
+ 'token': 1056,
39
+ 'token_str': ' mali',
40
+ 'sequence': 'včera večer sme mali nový film v kine, ktorý mal premiéru iba pred týždňom.'},
41
+ {'score': 0.1962040513753891,
42
+ 'token': 6915,
43
+ 'token_str': ' videli',
44
+ 'sequence': 'včera večer sme videli nový film v kine, ktorý mal premiéru iba pred týždňom.'},
45
+ {'score': 0.03656836599111557,
46
+ 'token': 26996,
47
+ 'token_str': ' pozerali',
48
+ 'sequence': 'včera večer sme pozerali nový film v kine, ktorý mal premiéru iba pred týždňom.'},
49
+ {'score': 0.030735589563846588,
50
+ 'token': 9058,
51
+ 'token_str': ' objavili',
52
+ 'sequence': 'včera večer sme objavili nový film v kine, ktorý mal premiéru iba pred týždňom.'}]
53
+ ```
54
+
55
+ ## Training Data
56
+
57
+ The `SK_BPE_BLM` model was pretrained using a subset of the OSCAR 2019 corpus, specifically focusing on the Slovak language. The corpus underwent comprehensive preprocessing to ensure the quality and relevance of the data:
58
+
59
+ - **Language Filtering:** Non-Slovak text was removed to focus solely on the Slovak language.
60
+ - **Character Normalization:** Various types of spaces, quotes, dashes, and separators were standardized (e.g., replacing different types of spaces with a single space, or dashes with hyphens). Emoticons were replaced with spaces.
61
+ - **Symbol and Unwanted Text Removal:** Sentences containing mathematical symbols, pictograms, or characters from Asian and African languages were deleted. Duplicates of punctuation, special characters, and spaces were also removed.
62
+ - **URL and Text Normalization:** All web addresses were removed, and the text was converted to lowercase to simplify tokenization.
63
+ - **Content Cleanup:** Text that included irrelevant content from web crawling, such as keywords and HTML tags, was identified and removed.
64
+
65
+ Additionally, the preprocessing included further refinement steps to create the final dataset:
66
+
67
+ - **Parentheses Content Removal:** All content within parentheses was removed to reduce noise.
68
+ - **Selection of Text Segments:** Medium-length text paragraphs were selected to maintain consistency.
69
+ - **Similarity Filtering:** Paragraphs with at least 50% similarity to previous ones were removed to minimize redundancy.
70
+ - **Random Sampling:** Finally, 20% of the remaining paragraphs were randomly selected.
71
+
72
+ After preprocessing, the training corpus consisted of:
73
+ - **455 MB of text**
74
+ - **895,125 paragraphs**
75
+ - **64.6 million words**
76
+ - **1.13 million unique words**
77
+ - **119 unique characters**
78
+
79
+ ## Pretraining
80
+
81
+ The `SK_BPE_BLM` model was trained with the following key parameters:
82
+
83
+ - **Architecture:** Based on RoBERTa, with 6 hidden layers and 12 attention heads.
84
+ - **Hidden size:** 576
85
+ - **Vocabulary size:** 50,264 tokens
86
+ - **Sequence length:** 256 tokens
87
+ - **Dropout:** 0.1
88
+ - **Number of parameters:** 58 million
89
+ - **Optimizer:** AdamW, learning rate 1×10^(-4), weight decay 0.01
90
+ - **Training:** 30 epochs, divided into 3 phases:
91
+ - **Phase 1:** 10 epochs on CPU (4x AMD EPYC 7542), batch size 64, 50 hours per epoch, 139,870 steps total.
92
+ - **Phase 2:** 5 epochs on GPU (1x Nvidia A100 40GB), batch size 64, 100 minutes per epoch, 69,935 steps total.
93
+ - **Phase 3:** 15 epochs on GPU (2x Nvidia A100 40GB), batch size 128, 60 minutes per epoch, 104,910 steps total.
94
+
95
+ The model was trained using the Hugging Face library, but without using the `Trainer` class—native PyTorch was used instead.
96
+
97
+ ## Fine-Tuned Versions of the SK_BPE_BLM Model
98
+
99
+ Here are the fine-tuned versions of the `SK_BPE_BLM` model based on the folders provided:
100
+
101
+ - [`SK_BPE_BLM-ner`](https://huggingface.co/daviddrzik/SK_BPE_BLM-ner): Fine-tuned for Named Entity Recognition (NER) tasks.
102
+ - [`SK_BPE_BLM-pos`](https://huggingface.co/daviddrzik/SK_BPE_BLM-pos): Fine-tuned for Part-of-Speech (POS) tagging.
103
+ - [`SK_BPE_BLM-qa`](https://huggingface.co/daviddrzik/SK_BPE_BLM-qa): Fine-tuned for Question Answering tasks.
104
+ - [`SK_BPE_BLM-sentiment-csfd`](https://huggingface.co/daviddrzik/SK_BPE_BLM-sentiment-csfd): Fine-tuned for sentiment analysis on the CSFD (movie review) dataset.
105
+ - [`SK_BPE_BLM-sentiment-multidomain`](https://huggingface.co/daviddrzik/SK_BPE_BLM-sentiment-multidomain): Fine-tuned for sentiment analysis across multiple domains.
106
+ - [`SK_BPE_BLM-sentiment-reviews`](https://huggingface.co/daviddrzik/SK_BPE_BLM-sentiment-reviews): Fine-tuned for sentiment analysis on general review datasets.
107
+ - [`SK_BPE_BLM-topic-news`](https://huggingface.co/daviddrzik/SK_BPE_BLM-topic-news): Fine-tuned for topic classification in news articles.