Instructions to use thewh1teagle/parakeet-tdt-0.6b-v3-he with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thewh1teagle/parakeet-tdt-0.6b-v3-he with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="thewh1teagle/parakeet-tdt-0.6b-v3-he")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("thewh1teagle/parakeet-tdt-0.6b-v3-he", device_map="auto") - Notebooks
- Google Colab
- Kaggle
parakeet-tdt-0.6b-v3-he
Hebrew speech recognition, fine-tuned from nvidia/parakeet-tdt-0.6b-v3.
Usage
pip install "transformers>=5.18" torch librosa
import librosa
from transformers import AutoModelForTDT, AutoProcessor
repo = "thewh1teagle/parakeet-tdt-0.6b-v3-he"
processor = AutoProcessor.from_pretrained(repo)
model = AutoModelForTDT.from_pretrained(repo)
audio, _ = librosa.load("audio.wav", sr=16000, mono=True)
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
out = model.generate(**inputs)
print(processor.batch_decode(out.sequences, skip_special_tokens=True)[0])
Also included: parakeet-tdt-0.6b-v3-he.q8_0.gguf for NeMo-Speech.cpp.
For the Vibe transcription app, use vibe-app/parakeet-tdt-0.6b-v3-he-gguf.
Results
WER (%, lower is better) on ivrit.ai's human-transcribed test sets, scored with the leaderboard normalization. Other models' numbers are from the leaderboard.
| Model | eval-d1 β | eval-whatsapp β |
|---|---|---|
| Soniox | 4.8 | 9.0 |
| ivrit-ai/whisper-large-v3-turbo | 5.3 | 7.1 |
| π¦ parakeet-tdt-0.6b-v3-he (this model) | 6.1 | 8.8 |
| Amazon Transcribe | 6.6 | 10.4 |
| Deepgram nova-3 | 6.7 | 12.0 |
| OpenAI gpt-4o-transcribe | 7.3 | 12.6 |
| openai/whisper-large-v3-turbo | 8.4 | 12.8 |
| openai/whisper-large-v3 | 9.8 | 13.2 |
Limitations
- Hebrew only: English words come out in Hebrew letters.
- Weaker on read encyclopedic text (FLEURS: 76.8 vs ~82 for the best models): rare names, numbers, Latin terms.
- Numbers are written as words or digits inconsistently.
Fine-tuning & training notes
Fine-tune this model
The .nemo includes the extended tokenizer β load it and train as usual, no change_vocabulary().
from nemo.collections.asr.models import ASRModel
model = ASRModel.from_pretrained("thewh1teagle/parakeet-tdt-0.6b-v3-he")
# then NeMo's standard ASR fine-tuning (manifest: audio_filepath, text, duration)
- Audio 16 kHz mono; normalize text like training: strip nikud and bidi marks, geresh/gershayim β
'/". - Low LR (~1e-5 β 3e-5), short warmup, bf16, ~600 s of audio per optimizer step (use grad accumulation).
- Set
decoding.greedy.use_cuda_graph_decoder: falsefor in-training validation β otherwise it decodes with stale weights and shows empty output.
How it was trained (new language on Parakeet)
Extend the tokenizer, don't replace it. Train a Hebrew BPE on ~250 h of normalized transcripts, then append its new pieces after v3's 8,192 β old ids and merges stay intact (vocab 9,206):
spm.SentencePieceTrainer.train( sentence_iterator=iter(texts), model_prefix="heb_1024", model_type="bpe", vocab_size=1024, character_coverage=1.0, # every Hebrew letter its own piece normalization_rule_name="nmt_nfkc", split_digits=True, # match v3 byte_fallback=False, unk_id=0, bos_id=-1, eos_id=-1, pad_id=-1, ) # append its NORMAL pieces not already in v3, scored below every v3 piece (keeps BPE merge order)Keep every pretrained weight.
change_vocabulary()re-initializes the whole decoder and joint; copy every old weight back by token id (prediction LSTM, joint hidden layers, old token rows, blank and TDT duration rows). Only the new rows start fresh: prediction-embedding and joint-output rows at the mean of the old rows (+1% noise), joint bias at the mean old bias.10Γ LR on decoder + joint (2e-3 vs 2e-4 encoder). New rows start identical and Adam moves them ~lr per step, so at a normal LR they stay indistinguishable for thousands of steps β the long "empty output" phase. With 10Γ it ends in ~100 steps. Keep warmup short.
Data: ~9k h of ivrit.ai audio pseudo-labeled with
ivrit-ai/whisper-large-v3-turbo, filtered withcompression_ratio β€ 2.2, one pass, duration-bucketed batches.More detail: NeMo #14140.
License
CC-BY-4.0. Base model by NVIDIA (CC-BY-4.0). Trained on Hebrew audio from ivrit.ai (license), labeled with ivrit-ai/whisper-large-v3-turbo.
- Downloads last month
- 208