Instructions to use coriollon/whisper-large-v3-turbo-russian with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use coriollon/whisper-large-v3-turbo-russian with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="coriollon/whisper-large-v3-turbo-russian")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("coriollon/whisper-large-v3-turbo-russian") model = AutoModelForSpeechSeq2Seq.from_pretrained("coriollon/whisper-large-v3-turbo-russian", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Whisper-large-v3-turbo Russian (Stage AW β long-form)
A Russian-specialized openai/whisper-large-v3-turbo
fine-tune with English preserved. This release (Stage AW) continues
Stage AS v2 with a targeted second pass that fixes long-form
punctuation collapse β the model's tendency to stop emitting .,
,, ?, ! in the second half of 20-30 s meeting-style segments.
Aggregate WER / Cap F1 / Sentence F1 stay at Stage AS v2 levels; punctuation density in long turns grows ~3Γ and n_valid rises from 53/150 to 140/150.
Headline metrics
General ASR (6 RU test sets, N=500 each, augmented refs):
| Metric | Stage AS v2 (prev) | Stage AW (this) | Ξ |
|---|---|---|---|
| WER (aggregate) | 9.02 % | 9.07 % | +0.05 pp |
| PER F1 (punct macro F1) | 0.398 | 0.388 | β0.010 |
| Cap F1 (word-level case-match) | 0.822 | 0.821 | β0.001 |
| Sent F1 (sentence-boundary F1) | 0.811 | 0.818 | +0.007 |
Long-form punct density (150 held-out concat-turns, 12-29 s each):
| Metric | Stage AS v2 | Stage AW | Ξ |
|---|---|---|---|
| first-half density mean | 0.074 | 0.214 | Γ2.9 |
| second-half density mean | 0.090 | 0.238 | Γ2.7 |
| ratio_med (d2/d1) | 1.14 | 1.00 | parity |
| n_valid (both halves punctuated) | 53/150 | 140/150 | Γ2.6 |
Per-dataset WER / PER F1
| Dataset | AS v2 WER | AW WER | AS v2 PER | AW PER |
|---|---|---|---|---|
| Common Voice 21 RU | 5.01 % | 4.96 % | 0.387 | 0.393 |
| RuLibriSpeech | 8.05 % | 7.41 % | 0.278 | 0.335 |
| Sberdevices Golos far-field | 9.42 % | 9.42 % | 0.596 | 0.487 |
| Sberdevices Golos crowd | 8.59 % | 8.35 % | 0.498 | 0.406 |
| SOVA RuDevices | 13.46 % | 13.97 % | 0.244 | 0.212 |
| Podlodka Speech | 9.61 % | 9.27 % | 0.383 | 0.494 |
WER improves on 4/6 datasets; PER regresses on short read-speech (Golos / SOVA) but improves on long meeting-style speech (Podlodka +0.11, RuLibri +0.06) β the corpus AW was tuned for.
What Stage AW fixed
Stage AS v2 was already an accurate Russian punctuator, but on long same-speaker turns (>15 s) the model would front-load punctuation: periods and commas in the first half, then trail off into punctuation-free run-on text. Empirically 65 % of 20-30 s turns in production meeting logs showed this pattern; users saw walls of text that were correct word-for-word but unreadable.
Stage AW retrains on 611 same-speaker concatenated turns (12-29 s each, 3.88 h total) from Pavel's own meeting corpus, with LLM-restored punctuation targets from gemma3:12b. The recipe continues from Stage AS v2 weights (not from Stage AI) so no short-form quality is sacrificed.
Recipe
- Init: Stage AS v2 merged weights
- Data: 611 rows of same-speaker concat, 12-29 s, gemma3:12b punct-restored targets, WER β€ 1 % word-preservation verified
- LoRA: r=48, target q/k/v/o + fc1/fc2 (same as AS v2)
- Schedule: 3 epochs, LR 5e-6 cosine, weight decay 0.01
- Compute: ~4 GPU-hours on RTX 4090 mobile
Usage
from transformers import WhisperProcessor, WhisperForConditionalGeneration
import torch
processor = WhisperProcessor.from_pretrained("coriollon/whisper-large-v3-turbo-russian")
model = WhisperForConditionalGeneration.from_pretrained(
"coriollon/whisper-large-v3-turbo-russian",
torch_dtype=torch.float16,
).to("cuda")
inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
input_features = inputs.input_features.to("cuda", dtype=torch.float16)
predicted_ids = model.generate(input_features, language="ru", task="transcribe")
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
Pre-quantized faster-whisper variants
| Folder | Quantization | Size |
|---|---|---|
ct2_int8_float16/ |
int8 weights + fp16 compute | 782 MB |
from huggingface_hub import snapshot_download
from faster_whisper import WhisperModel
ct2_path = snapshot_download(
repo_id="coriollon/whisper-large-v3-turbo-russian",
allow_patterns="ct2_int8_float16/*",
)
model = WhisperModel(f"{ct2_path}/ct2_int8_float16",
device="cuda", compute_type="int8_float16")
segments, _ = model.transcribe("audio.wav", language="ru", beam_size=5)
Reverting to Stage AS v2
The Stage AS v2 revision remains accessible via HF revision pinning:
processor = WhisperProcessor.from_pretrained(
"coriollon/whisper-large-v3-turbo-russian",
revision="stage-as-v2", # git tag pointing to pre-AW commit
)
Limitations
- On short read-speech (Golos / SOVA) PER F1 drops 0.02-0.11 vs AS v2 because AW was tuned for long conversational turns, not short clean utterances. WER is unaffected.
- English-only audio quality inherited from AS v2 baseline β degraded vs vanilla turbo.
- Same over-punctuation risk as AS v2 on very short (< 3 s) inputs.
License
Apache 2.0 (inherited from base Whisper model).
Changelog
- Stage AW (2026-09-07) β long-form density fine-tune. Continue-train from AS v2 on 611 same-speaker concat-turns (12-29 s, gemma3:12b targets, 3.88 h). Long-form density Γ2.7-2.9, n_valid 53β140/150, ratio_med 1.14β1.00. Aggregate WER +0.05 pp (in noise); 4/6 datasets improve.
- Metric fix (2026-07-19) β Evaluation harness now counts leading
em-dash (Russian dialogue) and standalone
.../β¦tokens. No change to the released weights, only to how their punctuation quality is reported. - Stage AS v2 (2026-07-14) β LoRA r=48, 3 epochs, rare-mark Γ2 sample weighting. WER 9.83 β 8.93 %. PER 0.39 β 0.41, Cap and Sent essentially preserved. Golos WER drops β2 pp on both splits.
- Stage AS v1 (2026-07-12) β punctuation & capitalization LoRA fine-tune. Cap F1 0.48 β 0.83, Sent F1 0.45 β 0.81, PER F1 0.20 β 0.39. WER 9.43 β 9.83 %.
- Stage AI (previous) β GigaAM-teacher relabel, WER 13.25 β 9.43 %.
- Downloads last month
- 768
Model tree for coriollon/whisper-large-v3-turbo-russian
Base model
openai/whisper-large-v3