Instructions to use daviddrzik/SK_BPE_BLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use daviddrzik/SK_BPE_BLM with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="daviddrzik/SK_BPE_BLM")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("daviddrzik/SK_BPE_BLM") model = AutoModelForMaskedLM.from_pretrained("daviddrzik/SK_BPE_BLM", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -2,4 +2,106 @@
|
|
| 2 |
license: mit
|
| 3 |
language:
|
| 4 |
- sk
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
license: mit
|
| 3 |
language:
|
| 4 |
- sk
|
| 5 |
+
datasets:
|
| 6 |
+
- oscar-corpus/OSCAR-2109
|
| 7 |
+
pipeline_tag: fill-mask
|
| 8 |
+
library_name: transformers
|
| 9 |
+
---
|
| 10 |
+
|
| 11 |
+
# Slovak BPE Baby Language Model (SK_BPE_BLM)
|
| 12 |
+
|
| 13 |
+
**SK_BPE_BLM** is a pretrained small language model for the Slovak language, based on the RoBERTa architecture. The model utilizes standard Byte-Pair Encoding (BPE) tokenization and is case-insensitive, meaning it operates in lowercase. While the pretrained model can be used for masked language modeling, it is primarily intended for fine-tuning on downstream NLP tasks.
|
| 14 |
+
|
| 15 |
+
## How to Use the Model
|
| 16 |
+
|
| 17 |
+
To use the SK_BPE_BLM model, follow these steps:
|
| 18 |
+
|
| 19 |
+
```python
|
| 20 |
+
from transformers import pipeline, RobertaTokenizer, AutoModelForMaskedLM
|
| 21 |
+
|
| 22 |
+
# Load the custom tokenizer and model
|
| 23 |
+
tokenizer = RobertaTokenizer.from_pretrained("daviddrzik/SK_BPE_BLM")
|
| 24 |
+
model = AutoModelForMaskedLM.from_pretrained("daviddrzik/SK_BPE_BLM")
|
| 25 |
+
|
| 26 |
+
# Create a pipeline with the custom model and tokenizer
|
| 27 |
+
unmasker = pipeline('fill-mask', model=model, tokenizer=tokenizer)
|
| 28 |
+
|
| 29 |
+
# Use the pipeline
|
| 30 |
+
result = unmasker("včera večer sme <mask> nový film v kine, ktorý mal premiéru iba pred týždňom.")
|
| 31 |
+
print(result)
|
| 32 |
+
|
| 33 |
+
[{'score': 0.2665567100048065,
|
| 34 |
+
'token': 18599,
|
| 35 |
+
'token_str': ' pozreli',
|
| 36 |
+
'sequence': 'včera večer sme pozreli nový film v kine, ktorý mal premiéru iba pred týždňom.'},
|
| 37 |
+
{'score': 0.23860174417495728,
|
| 38 |
+
'token': 1056,
|
| 39 |
+
'token_str': ' mali',
|
| 40 |
+
'sequence': 'včera večer sme mali nový film v kine, ktorý mal premiéru iba pred týždňom.'},
|
| 41 |
+
{'score': 0.1962040513753891,
|
| 42 |
+
'token': 6915,
|
| 43 |
+
'token_str': ' videli',
|
| 44 |
+
'sequence': 'včera večer sme videli nový film v kine, ktorý mal premiéru iba pred týždňom.'},
|
| 45 |
+
{'score': 0.03656836599111557,
|
| 46 |
+
'token': 26996,
|
| 47 |
+
'token_str': ' pozerali',
|
| 48 |
+
'sequence': 'včera večer sme pozerali nový film v kine, ktorý mal premiéru iba pred týždňom.'},
|
| 49 |
+
{'score': 0.030735589563846588,
|
| 50 |
+
'token': 9058,
|
| 51 |
+
'token_str': ' objavili',
|
| 52 |
+
'sequence': 'včera večer sme objavili nový film v kine, ktorý mal premiéru iba pred týždňom.'}]
|
| 53 |
+
```
|
| 54 |
+
|
| 55 |
+
## Training Data
|
| 56 |
+
|
| 57 |
+
The `SK_BPE_BLM` model was pretrained using a subset of the OSCAR 2019 corpus, specifically focusing on the Slovak language. The corpus underwent comprehensive preprocessing to ensure the quality and relevance of the data:
|
| 58 |
+
|
| 59 |
+
- **Language Filtering:** Non-Slovak text was removed to focus solely on the Slovak language.
|
| 60 |
+
- **Character Normalization:** Various types of spaces, quotes, dashes, and separators were standardized (e.g., replacing different types of spaces with a single space, or dashes with hyphens). Emoticons were replaced with spaces.
|
| 61 |
+
- **Symbol and Unwanted Text Removal:** Sentences containing mathematical symbols, pictograms, or characters from Asian and African languages were deleted. Duplicates of punctuation, special characters, and spaces were also removed.
|
| 62 |
+
- **URL and Text Normalization:** All web addresses were removed, and the text was converted to lowercase to simplify tokenization.
|
| 63 |
+
- **Content Cleanup:** Text that included irrelevant content from web crawling, such as keywords and HTML tags, was identified and removed.
|
| 64 |
+
|
| 65 |
+
Additionally, the preprocessing included further refinement steps to create the final dataset:
|
| 66 |
+
|
| 67 |
+
- **Parentheses Content Removal:** All content within parentheses was removed to reduce noise.
|
| 68 |
+
- **Selection of Text Segments:** Medium-length text paragraphs were selected to maintain consistency.
|
| 69 |
+
- **Similarity Filtering:** Paragraphs with at least 50% similarity to previous ones were removed to minimize redundancy.
|
| 70 |
+
- **Random Sampling:** Finally, 20% of the remaining paragraphs were randomly selected.
|
| 71 |
+
|
| 72 |
+
After preprocessing, the training corpus consisted of:
|
| 73 |
+
- **455 MB of text**
|
| 74 |
+
- **895,125 paragraphs**
|
| 75 |
+
- **64.6 million words**
|
| 76 |
+
- **1.13 million unique words**
|
| 77 |
+
- **119 unique characters**
|
| 78 |
+
|
| 79 |
+
## Pretraining
|
| 80 |
+
|
| 81 |
+
The `SK_BPE_BLM` model was trained with the following key parameters:
|
| 82 |
+
|
| 83 |
+
- **Architecture:** Based on RoBERTa, with 6 hidden layers and 12 attention heads.
|
| 84 |
+
- **Hidden size:** 576
|
| 85 |
+
- **Vocabulary size:** 50,264 tokens
|
| 86 |
+
- **Sequence length:** 256 tokens
|
| 87 |
+
- **Dropout:** 0.1
|
| 88 |
+
- **Number of parameters:** 58 million
|
| 89 |
+
- **Optimizer:** AdamW, learning rate 1×10^(-4), weight decay 0.01
|
| 90 |
+
- **Training:** 30 epochs, divided into 3 phases:
|
| 91 |
+
- **Phase 1:** 10 epochs on CPU (4x AMD EPYC 7542), batch size 64, 50 hours per epoch, 139,870 steps total.
|
| 92 |
+
- **Phase 2:** 5 epochs on GPU (1x Nvidia A100 40GB), batch size 64, 100 minutes per epoch, 69,935 steps total.
|
| 93 |
+
- **Phase 3:** 15 epochs on GPU (2x Nvidia A100 40GB), batch size 128, 60 minutes per epoch, 104,910 steps total.
|
| 94 |
+
|
| 95 |
+
The model was trained using the Hugging Face library, but without using the `Trainer` class—native PyTorch was used instead.
|
| 96 |
+
|
| 97 |
+
## Fine-Tuned Versions of the SK_BPE_BLM Model
|
| 98 |
+
|
| 99 |
+
Here are the fine-tuned versions of the `SK_BPE_BLM` model based on the folders provided:
|
| 100 |
+
|
| 101 |
+
- [`SK_BPE_BLM-ner`](https://huggingface.co/daviddrzik/SK_BPE_BLM-ner): Fine-tuned for Named Entity Recognition (NER) tasks.
|
| 102 |
+
- [`SK_BPE_BLM-pos`](https://huggingface.co/daviddrzik/SK_BPE_BLM-pos): Fine-tuned for Part-of-Speech (POS) tagging.
|
| 103 |
+
- [`SK_BPE_BLM-qa`](https://huggingface.co/daviddrzik/SK_BPE_BLM-qa): Fine-tuned for Question Answering tasks.
|
| 104 |
+
- [`SK_BPE_BLM-sentiment-csfd`](https://huggingface.co/daviddrzik/SK_BPE_BLM-sentiment-csfd): Fine-tuned for sentiment analysis on the CSFD (movie review) dataset.
|
| 105 |
+
- [`SK_BPE_BLM-sentiment-multidomain`](https://huggingface.co/daviddrzik/SK_BPE_BLM-sentiment-multidomain): Fine-tuned for sentiment analysis across multiple domains.
|
| 106 |
+
- [`SK_BPE_BLM-sentiment-reviews`](https://huggingface.co/daviddrzik/SK_BPE_BLM-sentiment-reviews): Fine-tuned for sentiment analysis on general review datasets.
|
| 107 |
+
- [`SK_BPE_BLM-topic-news`](https://huggingface.co/daviddrzik/SK_BPE_BLM-topic-news): Fine-tuned for topic classification in news articles.
|