You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Polish → Silesian Translator

  • Paper: TBA

A Polish-to-Silesian (POL → SZL) translation model based on TranslateGemma 4B IT, fine-tuned with QLoRA on the custom Polish–Silesian parallel corpus.

Model Details

  • Task: Polish → Silesian machine translation
  • Base model: google/translategemma-4b-it
  • Parameters: 4B
  • Fine-tuning: QLoRA
  • Metrics: BLEU and chrF (sacreBLEU 2.6.0)

Usage

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

base_model_id = "google/translategemma-4b-it"
adapter_path = "NASK-PIB/translategemma-4b-it-pol-szl-qlora"

tokenizer = AutoTokenizer.from_pretrained(base_model_id)

base_model = AutoModelForCausalLM.from_pretrained(
    base_model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

model = PeftModel.from_pretrained(
    base_model,
    adapter_path,
)

model.eval()

SOURCE_LANG = "Polish"
SOURCE_CODE = "pol"

TARGET_LANG = "Silesian"
TARGET_CODE = "szl"

text = "Jak się masz?"

# We need manually construct the prompt because original TranslateGemma model chat template does not support Silesian language.
prompt = (
    f"<start_of_turn>user\n"
    f"You are a professional {SOURCE_LANG} ({SOURCE_CODE}) to "
    f"{TARGET_LANG} ({TARGET_CODE}) translator. Your goal is to accurately "
    f"convey the meaning and nuances of the original {SOURCE_LANG} text while "
    f"adhering to {TARGET_LANG} grammar, vocabulary, and cultural sensitivities.\n"
    f"{text}"
    f"<end_of_turn>\n"
    f"<start_of_turn>model\n"
)

inputs = tokenizer(
    prompt,
    return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        max_new_tokens=512,
        do_sample=False,
        use_cache=True,
    )

response = tokenizer.decode(
    outputs[0][inputs["input_ids"].shape[-1] :],
    skip_special_tokens=True,
)

print(response)

Results

We evaluate on SiLTT and the BOUQuET benchmark.

Model SiLTT BLEU SiLTT chrF BOUQuET BLEU BOUQuET chrF
PLLuM-12B-nc-chat 1.7 22.3 3.4 28.4
Bielik-PL-11B-v3.0-IT 3.4 26.3 7.8 35.5
MADLAD-400-10B-MT 2.1 21.3 7.5 32.0
NLLB 3.3B 3.8 26.8 12.2 39.2
NLLB 54B MOE 3.2 26.0 9.9 38.3
TranslateGemma 4B 2.5 23.6 7.4 30.9
GPT-5.4 7.4 31.7 20.0 48.6
Google Translate 4.2 29.3 21.8 49.4
TranslateGemma-FT (REALESED) 8.0 31.9 26.3 52.8

Authors

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NASK-PIB/translategemma-4b-it-pol-szl-qlora

Finetuned
(25)
this model

Collection including NASK-PIB/translategemma-4b-it-pol-szl-qlora