You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Koyal Hindi ASR 120M

Koyal is a family of open speech recognition models for Indian languages, built by Adalat AI for document-ready dictation. Koyal models transcribe in rich orthography (RO): the output carries punctuation and formatted numerals as they appear in written documents, rather than a normalised lexical stream. This is the monolingual Hindi model. For Kannada, Malayalam, Telugu, or a multilingual model with streaming support, see Related models.

At a glance

Field Value
Task Automatic speech recognition, rich orthography
Language Hindi
Parameters ~120M (checkpoint ~0.5 GB)
Mode Offline
Base model MahaDhwani (AI4Bharat)
Framework NVIDIA NeMo, no custom code
WER_SCRIBE 9.03%
License CC-BY-4.0

Quickstart

Install NVIDIA NeMo:

pip install "nemo-toolkit[asr]"

Load and transcribe:

import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.from_pretrained("adalat-ai/koyal-hi-120m-1.0")
print(model.transcribe(["audio.wav"])[0].text)  # TDT greedy decoding

Input audio should be 16 kHz mono WAV. TDT beam search (maes) gives a small additional accuracy gain over greedy decoding.

Intended use

Document-ready Hindi speech transcription: domains where the transcript is the deliverable and punctuation and numeral formatting must match written convention, such as legal and courtroom dictation.

Limitations

  • Offline only.
  • Trained and evaluated on 16 kHz audio; performance on other sampling rates is not characterised.
  • Not trained for code-switched speech.
  • Error rates vary widely across corpora. FLEURS-RO reflects read speech; conversational corpora such as IndicVoices and MUCS are materially harder. See the evaluation tables for the observed spread.

Model architecture

Field Value
Architecture Conformer Hybrid TDT-CTC (EncDecHybridRNNTCTCBPEModel)
Parameters ~120M
Encoder 17-layer ConformerEncoder, d_model=512, ×4 subsampling, warm-started from MahaDhwani
Decoder TDT (durations 0–4) with auxiliary CTC head
Tokenizer SentencePiece BPE, 512 tokens
Input 16 kHz mono audio
Output Document-ready Hindi text with punctuation and formatted numerals

The encoder is warm-started from MahaDhwani [1], AI4Bharat's self-supervised Conformer encoder pretrained on 279K hours of raw audio across 22 Indian languages, then fine-tuned end-to-end on Hindi dictation data. The released weights are an average of the final converged checkpoints. This is a standard NeMo Conformer hybrid transducer-CTC model — it loads, fine-tunes and exports with standard NeMo ASR tooling.

MahaDhwani is a Conformer encoder rather than a FastConformer, so the cache-aware streaming recipes do not apply; this model is offline only. For streaming, see koyal-indic-600m-1.0.

Evaluation

Results are reported on two test-set families:

  • rich-orthography, with references curated to carry punctuation and written numerals, scored with WER_SCRIBE; and
  • verbatim, the public benchmarks as published, scored with ER_LEX.

Numbers are not comparable across the two settings. Several corpora appear in both tables — those are the same benchmarks scored against different references, so their ER_lex values differ. That is expected, not a discrepancy.

  • Why the two settings exist

    Most public test sets are not punctuated. Of the Hindi benchmarks used here, only FLEURS-RO (100%), IndicTTS (97%) and RESPIN (58%) carry punctuation in their references; IndicVoices, Kathbath, MUCS and Common Voice carry none.

    Scoring a rich-orthography model on WER_SCRIBE against an unpunctuated reference charges every emitted comma as an insertion — it measures the reference's annotation convention, not the model. The verbatim setting therefore uses ER_LEX, which scores word identity alone and is the task all these systems share.

    Rich-orthography Verbatim
    References Curated to rich orthography Public benchmarks as published
    Metric WER_SCRIBE, with ER_lex / ER_num / ER_punc and jiwer WER/CER ER_LEX only
    Question Does the model produce the correct document-ready transcript? Does the model recognise the words?

Rich-orthography results

Evaluated with scribe-eval, which aligns hypothesis and reference at the token level and decomposes errors by category. ER_lex — lexical tokens. ER_num — number tokens. ER_punc — punctuation tokens. WER_SCRIBE — over all tokens (WER_S in the SCRIBE paper). WER / CER — computed with jiwer on the same unnormalised text.

No text normalisation is applied to references or hypotheses. Lower is better throughout. Lahaja is held out entirely for evaluation and contributes no training data.

*koyal-hi-120m-1.0 on curated RO test sets.*

Dataset Clips ER_lex (%) ER_num (%) ER_punc (%) WER_SCRIBE (%) WER (%) CER (%)
FLEURS-RO 418 7.09 0.30 1.64 9.03 10.63 3.79
IndicTTS 293 4.57 0.15 2.96 7.69 10.26 2.73
Kathbath 1,572 5.54 0.10 0.71 6.36 8.06 2.66
Mann ki Baat 479 6.41 0.18 1.74 8.34 10.53 3.97
OpenSLR 118 735 29.77 0.38 5.35 35.51 38.06 25.53
Shrutilipi 2,118 2.98 0.10 0.65 3.73 4.86 1.90
Mozilla Common Voice 401 5.95 0.05 4.55 10.54 12.65 4.79
IndicVoices 1,705 10.27 0.16 4.33 14.76 17.68 7.53
Lahaja (held out) 5,893 11.27 0.21 2.94 14.42 16.40 6.28

FLEURS-RO is the headline benchmark. For reference, the SCRIBE paper [2] reports the following systems on it:

Model ER_lex (%) ER_num (%) ER_punc (%) WER_SCRIBE (%) WER (%)
IndicWhisper 23.80 1.06 6.87 31.73 35.20
IndicConformer 10.16 1.35 6.99 18.50 21.70
SCRIBE-ASR (Whisper-small) 11.68 0.31 3.30 15.29 17.57

Koyal is about half the size of the SCRIBE Whisper-small and scores 9.03 against its 15.29.

Why WER_SCRIBE and WER differ

Alignment in scribe-eval is sandhi-tolerant: agglutination makes word boundaries unstable in Indic text, and the same speech can be validly written as one word or two — for example Hindi उस में and उसमें (a postposition merge). scribe-eval detects such merges and splits at alignment time and scores them as matches, where plain word-level WER counts each as an error. This accounts for much of the gap between the two columns above.

Because the model emits punctuation and written number forms, error rates on unnormalised rich-orthography references are not directly comparable to WER figures computed on normalised text. See the SCRIBE paper [2] for metric definitions and the FLEURS-RO annotation protocol.

Verbatim results

Public Hindi benchmarks as published, scored with ER_LEX. Datasets marked * have punctuated references; the rest do not.

Numerals are normalised on both sides with Indic Num2Words before scoring. Without it, a model that writes 15 is charged an error against a reference that spells the number out, which measures formatting convention rather than recognition. Unnormalised figures are on the benchmarks page.

The systems below differ in size, training objective and output format, and saaras-v4 is a commercial API rather than an open model. Benchmark coverage differs by language. This is a range, not a ranking.

ER_LEX (%). Lower is better. Macro is the unweighted mean across the seven benchmarks.

Model Macro IndicVoices FLEURS-RO* Kathbath RESPIN* CommonVoice MUCS IndicTTS*
clips (n) 5,325 418 3,151 2,288 1,727 3,897 100
indic-conformer-600m-ctc 10.26 13.31 9.52 8.23 11.31 11.63 10.33 7.50
indic-conformer-600m-rnnt 9.71 12.62 9.49 7.76 11.00 10.97 9.75 6.35
koyal-hi-120m-1.0 9.59 14.76 7.45 7.42 9.48 6.85 14.46 6.73
koyal-indic-600m-1.0 12.52 17.58 10.42 9.91 11.06 13.25 16.81 8.60
sravaani-1.0 9.70 12.54 8.60 8.60 7.85 12.47 11.93 5.92
saaras-v4 (commercial) 7.95 10.72 7.27 5.62 8.93 7.68 8.65 6.81

Training data

Fine-tuned on ~2,400 hours of Hindi speech assembled from public sources, with transcripts curated to rich orthographic form (punctuation and written numerals preserved) following the transcription curation pipeline described in the SCRIBE paper [2]. Sources:

  • IndicVoices
  • Shrutilipi
  • IndicVoices-ST
  • Kathbath
  • OpenSLR 118
  • Mann ki Baat
  • Mozilla Common Voice
  • IndicTTS
  • Google FLEURS
  • Festvox IIITH

Lahaja is held out entirely for evaluation and contributes no training data.

Related models

Model Language Params Modes
koyal-hi-120m-1.0 Hindi 120M Offline
koyal-kn-120m-1.0 Kannada 120M Offline
koyal-ml-120m-1.0 Malayalam 120M Offline
koyal-te-120m-1.0 Telugu 120M Offline
koyal-indic-600m-1.0 hi · kn · ml · te 600M Offline + streaming

References

  1. Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages (MahaDhwani), ICASSP 2025
  2. SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

Citation

@inproceedings{bhogale2025mahadhwani,
  author    = {Bhogale, Kaushal Santosh and Mehendale, Deovrat and Javed, Tahir and Anuragi, Devbrat and Joshi, Sakshi and Sundaresan, Sai and Ananthanarayanan, Aparna and Dey, Sharmistha and G, Sathish Kumar Reddy and Srinivasan, Anusha and Raman, Abhigyan and Kumar, Pratyush and Khapra, Mitesh M.},
  title     = {Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages},
  booktitle = {ICASSP 2025 -- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2025},
  pages     = {1--5},
  doi       = {10.1109/ICASSP49660.2025.10888018}
}
@article{manohar2026scribe,
  title   = {SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR},
  author  = {Manohar, Kavya and Bhattacharya, Arghya and Juvekar, Kush and Nethil, Kumarmanas},
  journal = {arXiv preprint arXiv:2605.20712},
  year    = {2026}
}

License

Released under CC-BY-4.0. The encoder is initialised from the MIT-licensed MahaDhwani pretrained Conformer checkpoint by AI4Bharat.

Contact

Questions and feedback: the Community tab of this repository.

Downloads last month
58
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adalat-ai/koyal-hi-120m-1.0

Finetuned
(4)
this model

Dataset used to train adalat-ai/koyal-hi-120m-1.0

Collection including adalat-ai/koyal-hi-120m-1.0

Paper for adalat-ai/koyal-hi-120m-1.0

Evaluation results