You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Koyal Telugu ASR 120M

Koyal is a family of open speech recognition models for Indian languages, built by Adalat AI for document-ready dictation. Koyal models transcribe in rich orthography (RO): the output carries punctuation and formatted numerals as they appear in written documents, rather than a normalised lexical stream. This is the monolingual Telugu model. For Hindi, Kannada, Malayalam, or a multilingual model with streaming support, see Related models.

At a glance

Field Value
Task Automatic speech recognition, rich orthography
Language Telugu
Parameters ~120M (checkpoint ~0.5 GB)
Mode Offline
Base model MahaDhwani (AI4Bharat)
Framework NVIDIA NeMo, no custom code
WER_SCRIBE 17.81%
License CC-BY-4.0

Quickstart

Install NVIDIA NeMo:

pip install "nemo-toolkit[asr]"

Load and transcribe:

import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.from_pretrained("adalat-ai/koyal-te-120m-1.0")
print(model.transcribe(["audio.wav"])[0].text)  # TDT greedy decoding

Input audio should be 16 kHz mono WAV. TDT beam search (maes) gives a small additional accuracy gain over greedy decoding.

Intended use

Document-ready Telugu speech transcription: domains where the transcript is the deliverable and punctuation and numeral formatting must match written convention, such as legal and courtroom dictation.

Limitations

  • Offline only.
  • Trained and evaluated on 16 kHz audio; performance on other sampling rates is not characterised.
  • Not trained for code-switched speech.
  • Error rates vary widely across corpora. FLEURS-RO reflects read speech; conversational corpora such as IndicVoices and SeamlessAlign are materially harder. See the evaluation tables for the observed spread.

Model architecture

Field Value
Architecture Conformer Hybrid TDT-CTC (EncDecHybridRNNTCTCBPEModel)
Parameters ~120M
Encoder 17-layer ConformerEncoder, d_model=512, ×4 subsampling, warm-started from MahaDhwani
Decoder TDT (durations 0–4) with auxiliary CTC head
Tokenizer SentencePiece BPE, 512 tokens
Input 16 kHz mono audio
Output Document-ready Telugu text with punctuation and formatted numerals

The encoder is warm-started from MahaDhwani [1], AI4Bharat's self-supervised Conformer encoder pretrained on 279K hours of raw audio across 22 Indian languages, then fine-tuned end-to-end on Telugu dictation data. The released weights are an average of the final converged checkpoints. This is a standard NeMo Conformer hybrid transducer-CTC model — it loads, fine-tunes and exports with standard NeMo ASR tooling.

MahaDhwani is a Conformer encoder rather than a FastConformer, so the cache-aware streaming recipes do not apply; this model is offline only. For streaming, see koyal-indic-600m-1.0.

Evaluation

Results are reported on two test-set families:

  • rich-orthography, with references curated to carry punctuation and written numerals, scored with WER_SCRIBE; and
  • verbatim, the public benchmarks as published, scored with ER_LEX.

Numbers are not comparable across the two settings. Several corpora appear in both tables — those are the same benchmarks scored against different references, so their ER_lex values differ. That is expected, not a discrepancy.

  • Why the two settings exist

    Most public test sets are not punctuated. Of the Telugu benchmarks used here, only FLEURS-RO (100%), IndicTTS (100%) and RESPIN (56%) carry punctuation in their references; IndicVoices, Kathbath and MUCS carry none.

    Scoring a rich-orthography model on WER_SCRIBE against an unpunctuated reference charges every emitted comma as an insertion — it measures the reference's annotation convention, not the model. The verbatim setting therefore uses ER_LEX, which scores word identity alone and is the task all these systems share.

    Rich-orthography Verbatim
    References Curated to rich orthography Public benchmarks as published
    Metric WER_SCRIBE, with ER_lex / ER_num / ER_punc and jiwer WER/CER ER_LEX only
    Question Does the model produce the correct document-ready transcript? Does the model recognise the words?

Rich-orthography results

Evaluated with scribe-eval, which aligns hypothesis and reference at the token level and decomposes errors by category. ER_lex — lexical tokens. ER_num — number tokens. ER_punc — punctuation tokens. WER_SCRIBE — over all tokens (WER_S in the SCRIBE paper). WER / CER — computed with jiwer on the same unnormalised text.

No text normalisation is applied to references or hypotheses. Lower is better throughout.

*koyal-te-120m-1.0 on curated RO test sets.*

Dataset Clips ER_lex (%) ER_num (%) ER_punc (%) WER_SCRIBE (%) WER (%) CER (%)
FLEURS-RO 466 10.63 0.57 6.61 17.81 24.49 5.12
IndicTTS 643 8.31 0.15 5.96 14.42 25.58 4.84
IndicVoices 1,648 16.00 0.20 5.49 21.69 29.29 9.21
Kathbath 1,190 11.39 0.09 2.45 13.93 20.84 3.92
Mann ki Baat 657 10.29 0.20 2.10 12.59 18.16 6.10
OpenSLR 66 333 6.81 0.05 1.57 8.43 15.73 2.73
SeamlessAlign 1,507 17.57 0.23 6.77 24.58 30.79 12.66
Shrutilipi 1,402 9.80 0.14 1.34 11.28 18.01 3.84

FLEURS-RO is the headline benchmark. The SCRIBE paper [2] does not cover Telugu, so there is no directly comparable prior system on this benchmark.

Why WER_SCRIBE and WER differ

Alignment in scribe-eval is sandhi-tolerant: agglutination makes word boundaries unstable in Indic text, and the same speech can be validly written as one word or two. scribe-eval detects such merges and splits at alignment time and scores them as matches, where plain word-level WER counts each as an error. This accounts for much of the gap between the two columns above — on FLEURS-RO, 17.81 WER_SCRIBE against 24.49 plain WER on identical output.

Because the model emits punctuation and written number forms, error rates on unnormalised rich-orthography references are not directly comparable to WER figures computed on normalised text. See the SCRIBE paper [2] for metric definitions and the FLEURS-RO annotation protocol.

Verbatim results

Public Telugu benchmarks as published, scored with ER_LEX. Datasets marked * have punctuated references; the rest do not.

Numerals are normalised on both sides with Indic Num2Words before scoring. Without it, a model that writes 15 is charged an error against a reference that spells the number out, which measures formatting convention rather than recognition. Unnormalised figures are on the benchmarks page.

The systems below differ in size, training objective and output format, and saaras-v4 is a commercial API rather than an open model. Benchmark coverage differs by language — Telugu has no Common Voice coverage, so this macro is over six benchmarks rather than the seven used for Hindi. This is a range, not a ranking.

ER_LEX (%). Lower is better. Macro is the unweighted mean across the six benchmarks.

Model Macro IndicVoices FLEURS-RO* Kathbath RESPIN* MUCS IndicTTS*
clips (n) 3,235 466 2,379 2,226 2,549 100
indic-conformer-600m-ctc 17.36 22.85 14.64 14.82 18.31 22.12 11.40
indic-conformer-600m-rnnt 16.56 21.53 13.94 14.38 18.83 19.83 10.85
koyal-te-120m-1.0 15.81 21.90 11.06 13.79 18.29 19.72 10.12
koyal-indic-600m-1.0 20.45 27.84 14.15 18.60 21.57 27.16 13.36
sravaani-1.0 16.55 21.67 13.84 17.16 16.47 21.80 8.37
saaras-v4 (commercial) 13.71 17.90 11.70 11.69 15.97 14.94 10.03

Training data

Fine-tuned on ~1,720 hours of Telugu speech assembled from public sources, with transcripts curated to rich orthographic form (punctuation and written numerals preserved) following the transcription curation pipeline described in the SCRIBE paper [2].

Sources:

  • SeamlessAlign
  • IndicVoices
  • Kathbath
  • Shrutilipi
  • Mann ki Baat
  • IndicTTS
  • Google FLEURS
  • OpenSLR 66
  • Festvox IIITH

Related models

Model Language Params Modes
koyal-hi-120m-1.0 Hindi 120M Offline
koyal-kn-120m-1.0 Kannada 120M Offline
koyal-ml-120m-1.0 Malayalam 120M Offline
koyal-te-120m-1.0 Telugu 120M Offline
koyal-indic-600m-1.0 hi · kn · ml · te 600M Offline + streaming

References

  1. Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages (MahaDhwani), ICASSP 2025
  2. SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

Citation

@inproceedings{bhogale2025mahadhwani,
  author    = {Bhogale, Kaushal Santosh and Mehendale, Deovrat and Javed, Tahir and Anuragi, Devbrat and Joshi, Sakshi and Sundaresan, Sai and Ananthanarayanan, Aparna and Dey, Sharmistha and G, Sathish Kumar Reddy and Srinivasan, Anusha and Raman, Abhigyan and Kumar, Pratyush and Khapra, Mitesh M.},
  title     = {Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages},
  booktitle = {ICASSP 2025 -- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2025},
  pages     = {1--5},
  doi       = {10.1109/ICASSP49660.2025.10888018}
}
@article{manohar2026scribe,
  title   = {SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR},
  author  = {Manohar, Kavya and Bhattacharya, Arghya and Juvekar, Kush and Nethil, Kumarmanas},
  journal = {arXiv preprint arXiv:2605.20712},
  year    = {2026}
}

License

Released under CC-BY-4.0. The encoder is initialised from the MIT-licensed MahaDhwani pretrained Conformer checkpoint by AI4Bharat.

Contact

Questions and feedback: the Community tab of this repository.

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adalat-ai/koyal-te-120m-1.0

Finetuned
(4)
this model

Dataset used to train adalat-ai/koyal-te-120m-1.0

Collection including adalat-ai/koyal-te-120m-1.0

Paper for adalat-ai/koyal-te-120m-1.0

Evaluation results