Whisper-large-v3-turbo Russian (Stage AW β€” long-form)

A Russian-specialized openai/whisper-large-v3-turbo fine-tune with English preserved. This release (Stage AW) continues Stage AS v2 with a targeted second pass that fixes long-form punctuation collapse β€” the model's tendency to stop emitting ., ,, ?, ! in the second half of 20-30 s meeting-style segments.

Aggregate WER / Cap F1 / Sentence F1 stay at Stage AS v2 levels; punctuation density in long turns grows ~3Γ— and n_valid rises from 53/150 to 140/150.

Headline metrics

General ASR (6 RU test sets, N=500 each, augmented refs):

Metric Stage AS v2 (prev) Stage AW (this) Ξ”
WER (aggregate) 9.02 % 9.07 % +0.05 pp
PER F1 (punct macro F1) 0.398 0.388 βˆ’0.010
Cap F1 (word-level case-match) 0.822 0.821 βˆ’0.001
Sent F1 (sentence-boundary F1) 0.811 0.818 +0.007

Long-form punct density (150 held-out concat-turns, 12-29 s each):

Metric Stage AS v2 Stage AW Ξ”
first-half density mean 0.074 0.214 Γ—2.9
second-half density mean 0.090 0.238 Γ—2.7
ratio_med (d2/d1) 1.14 1.00 parity
n_valid (both halves punctuated) 53/150 140/150 Γ—2.6

Per-dataset WER / PER F1

Dataset AS v2 WER AW WER AS v2 PER AW PER
Common Voice 21 RU 5.01 % 4.96 % 0.387 0.393
RuLibriSpeech 8.05 % 7.41 % 0.278 0.335
Sberdevices Golos far-field 9.42 % 9.42 % 0.596 0.487
Sberdevices Golos crowd 8.59 % 8.35 % 0.498 0.406
SOVA RuDevices 13.46 % 13.97 % 0.244 0.212
Podlodka Speech 9.61 % 9.27 % 0.383 0.494

WER improves on 4/6 datasets; PER regresses on short read-speech (Golos / SOVA) but improves on long meeting-style speech (Podlodka +0.11, RuLibri +0.06) β€” the corpus AW was tuned for.

What Stage AW fixed

Stage AS v2 was already an accurate Russian punctuator, but on long same-speaker turns (>15 s) the model would front-load punctuation: periods and commas in the first half, then trail off into punctuation-free run-on text. Empirically 65 % of 20-30 s turns in production meeting logs showed this pattern; users saw walls of text that were correct word-for-word but unreadable.

Stage AW retrains on 611 same-speaker concatenated turns (12-29 s each, 3.88 h total) from Pavel's own meeting corpus, with LLM-restored punctuation targets from gemma3:12b. The recipe continues from Stage AS v2 weights (not from Stage AI) so no short-form quality is sacrificed.

Recipe

  • Init: Stage AS v2 merged weights
  • Data: 611 rows of same-speaker concat, 12-29 s, gemma3:12b punct-restored targets, WER ≀ 1 % word-preservation verified
  • LoRA: r=48, target q/k/v/o + fc1/fc2 (same as AS v2)
  • Schedule: 3 epochs, LR 5e-6 cosine, weight decay 0.01
  • Compute: ~4 GPU-hours on RTX 4090 mobile

Usage

from transformers import WhisperProcessor, WhisperForConditionalGeneration
import torch

processor = WhisperProcessor.from_pretrained("coriollon/whisper-large-v3-turbo-russian")
model = WhisperForConditionalGeneration.from_pretrained(
    "coriollon/whisper-large-v3-turbo-russian",
    torch_dtype=torch.float16,
).to("cuda")

inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
input_features = inputs.input_features.to("cuda", dtype=torch.float16)
predicted_ids = model.generate(input_features, language="ru", task="transcribe")
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]

Pre-quantized faster-whisper variants

Folder Quantization Size
ct2_int8_float16/ int8 weights + fp16 compute 782 MB
from huggingface_hub import snapshot_download
from faster_whisper import WhisperModel

ct2_path = snapshot_download(
    repo_id="coriollon/whisper-large-v3-turbo-russian",
    allow_patterns="ct2_int8_float16/*",
)
model = WhisperModel(f"{ct2_path}/ct2_int8_float16",
                    device="cuda", compute_type="int8_float16")
segments, _ = model.transcribe("audio.wav", language="ru", beam_size=5)

Reverting to Stage AS v2

The Stage AS v2 revision remains accessible via HF revision pinning:

processor = WhisperProcessor.from_pretrained(
    "coriollon/whisper-large-v3-turbo-russian",
    revision="stage-as-v2",   # git tag pointing to pre-AW commit
)

Limitations

  • On short read-speech (Golos / SOVA) PER F1 drops 0.02-0.11 vs AS v2 because AW was tuned for long conversational turns, not short clean utterances. WER is unaffected.
  • English-only audio quality inherited from AS v2 baseline β€” degraded vs vanilla turbo.
  • Same over-punctuation risk as AS v2 on very short (< 3 s) inputs.

License

Apache 2.0 (inherited from base Whisper model).

Changelog

  • Stage AW (2026-09-07) β€” long-form density fine-tune. Continue-train from AS v2 on 611 same-speaker concat-turns (12-29 s, gemma3:12b targets, 3.88 h). Long-form density Γ—2.7-2.9, n_valid 53β†’140/150, ratio_med 1.14β†’1.00. Aggregate WER +0.05 pp (in noise); 4/6 datasets improve.
  • Metric fix (2026-07-19) β€” Evaluation harness now counts leading em-dash (Russian dialogue) and standalone ... / … tokens. No change to the released weights, only to how their punctuation quality is reported.
  • Stage AS v2 (2026-07-14) β€” LoRA r=48, 3 epochs, rare-mark Γ—2 sample weighting. WER 9.83 β†’ 8.93 %. PER 0.39 β†’ 0.41, Cap and Sent essentially preserved. Golos WER drops βˆ’2 pp on both splits.
  • Stage AS v1 (2026-07-12) β€” punctuation & capitalization LoRA fine-tune. Cap F1 0.48 β†’ 0.83, Sent F1 0.45 β†’ 0.81, PER F1 0.20 β†’ 0.39. WER 9.43 β†’ 9.83 %.
  • Stage AI (previous) β€” GigaAM-teacher relabel, WER 13.25 β†’ 9.43 %.
Downloads last month
768
Safetensors
Model size
0.8B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for coriollon/whisper-large-v3-turbo-russian

Finetuned
(648)
this model
Finetunes
2 models
Quantizations
1 model