Instructions to use adalat-ai/koyal-ml-120m-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use adalat-ai/koyal-ml-120m-1.0 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("adalat-ai/koyal-ml-120m-1.0") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Koyal Malayalam ASR 120M
Koyal is a family of open speech recognition models for Indian languages, built by Adalat AI for document-ready dictation. Koyal models transcribe in rich orthography (RO): the output carries punctuation and formatted numerals as they appear in written documents, rather than a normalised lexical stream. This is the monolingual Malayalam model. For Hindi, Kannada, Telugu, or a multilingual model with streaming support, see Related models.
At a glance
| Field | Value |
|---|---|
| Task | Automatic speech recognition, rich orthography |
| Language | Malayalam |
| Parameters | ~120M (checkpoint ~0.5 GB) |
| Mode | Offline |
| Base model | MahaDhwani (AI4Bharat) |
| Framework | NVIDIA NeMo, no custom code |
| WER_SCRIBE | 14.74% |
| License | CC-BY-4.0 |
Quickstart
Install NVIDIA NeMo:
pip install "nemo-toolkit[asr]"
Load and transcribe:
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.from_pretrained("adalat-ai/koyal-ml-120m-1.0")
print(model.transcribe(["audio.wav"])[0].text) # TDT greedy decoding
Input audio should be 16 kHz mono WAV. TDT beam search (maes) gives a small additional accuracy gain over greedy decoding.
Intended use
Document-ready Malayalam speech transcription: domains where the transcript is the deliverable and punctuation and numeral formatting must match written convention, such as legal and courtroom dictation.
Limitations
- Offline only.
- Trained and evaluated on 16 kHz audio; performance on other sampling rates is not characterised.
- Not trained for code-switched speech.
- Error rates vary widely across corpora. FLEURS-RO reflects read speech; conversational corpora such as IndicVoices are materially harder. See the evaluation tables for the observed spread.
Model architecture
| Field | Value |
|---|---|
| Architecture | Conformer Hybrid TDT-CTC (EncDecHybridRNNTCTCBPEModel) |
| Parameters | ~120M |
| Encoder | 17-layer ConformerEncoder, d_model=512, ×4 subsampling, warm-started from MahaDhwani |
| Decoder | TDT (durations 0–4) with auxiliary CTC head |
| Tokenizer | SentencePiece BPE, 512 tokens |
| Input | 16 kHz mono audio |
| Output | Document-ready Malayalam text with punctuation and formatted numerals |
The encoder is warm-started from MahaDhwani [1], AI4Bharat's self-supervised Conformer encoder pretrained on 279K hours of raw audio across 22 Indian languages, then fine-tuned end-to-end on Malayalam dictation data. The released weights are an average of the final converged checkpoints. This is a standard NeMo Conformer hybrid transducer-CTC model — it loads, fine-tunes and exports with standard NeMo ASR tooling.
MahaDhwani is a Conformer encoder rather than a FastConformer, so the cache-aware streaming recipes do not apply; this model is offline only. For streaming, see koyal-indic-600m-1.0.
Evaluation
Results are reported on two test-set families:
- rich-orthography, with references curated to carry punctuation and written numerals, scored with WER_SCRIBE; and
- verbatim, the public benchmarks as published, scored with ER_LEX.
Numbers are not comparable across the two settings. Several corpora appear in both tables — those are the same benchmarks scored against different references, so their ER_lex values differ. That is expected, not a discrepancy.
Why the two settings exist
Most public test sets are not punctuated. Of the Malayalam benchmarks used here, only FLEURS-RO (100%) carries punctuation in its references; IndicVoices, Kathbath and Common Voice carry none, and Malayalam IndicTTS is only 2% punctuated — unlike Hindi, Kannada and Telugu IndicTTS, which are heavily punctuated.
Scoring a rich-orthography model on WER_SCRIBE against an unpunctuated reference charges every emitted comma as an insertion — it measures the reference's annotation convention, not the model. The verbatim setting therefore uses ER_LEX, which scores word identity alone and is the task all these systems share.
Rich-orthography Verbatim References Curated to rich orthography Public benchmarks as published Metric WER_SCRIBE, with ER_lex / ER_num / ER_punc and jiwer WER/CER ER_LEX only Question Does the model produce the correct document-ready transcript? Does the model recognise the words?
Rich-orthography results
Evaluated with scribe-eval, which aligns hypothesis and reference at the token level and decomposes errors by category. ER_lex — lexical tokens. ER_num — number tokens. ER_punc — punctuation tokens. WER_SCRIBE — over all tokens (WER_S in the SCRIBE paper). WER / CER — computed with jiwer on the same unnormalised text.
No text normalisation is applied to references or hypotheses. Lower is better throughout.
*koyal-ml-120m-1.0 on curated RO test sets.*
| Dataset | Clips | ER_lex (%) | ER_num (%) | ER_punc (%) | WER_SCRIBE (%) | WER (%) | CER (%) |
|---|---|---|---|---|---|---|---|
| FLEURS-RO | 957 | 7.48 | 0.41 | 6.85 | 14.74 | 22.69 | 3.47 |
| IndicTTS | 213 | 6.91 | 0.18 | 1.91 | 9.00 | 18.97 | 2.30 |
| IndicVoices | 1,587 | 17.09 | 0.16 | 6.69 | 23.94 | 35.62 | 9.69 |
| Kathbath | 883 | 9.46 | 0.17 | 0.99 | 10.61 | 18.34 | 2.80 |
| Mozilla Common Voice 17.0 | 668 | 11.06 | 0.18 | 4.55 | 15.78 | 30.19 | 4.30 |
| OpenSLR 63 | 334 | 4.38 | 0.27 | 0.31 | 4.96 | 10.33 | 1.51 |
| Shrutilipi | 3,077 | 8.40 | 0.19 | 0.91 | 9.50 | 17.15 | 3.01 |
FLEURS-RO is the headline benchmark. For reference, the SCRIBE paper [2] reports the following systems on it:
| Model | ER_lex (%) | ER_num (%) | ER_punc (%) | WER_SCRIBE (%) | WER (%) |
|---|---|---|---|---|---|
| IndicWhisper | 14.65 | 1.74 | 15.41 | 31.80 | 41.77 |
| IndicConformer | 13.58 | 2.39 | 15.40 | 31.37 | 41.00 |
| SCRIBE-ASR (Whisper-small) | 14.77 | 0.59 | 14.03 | 29.39 | 36.65 |
Koyal is about half the size of the SCRIBE Whisper-small and scores 14.74 against its 29.39.
Why WER_SCRIBE and WER differ
Alignment in scribe-eval is sandhi-tolerant: agglutination makes word boundaries unstable in Indic text, and the same speech can be validly written as one word or two. scribe-eval detects such merges and splits at alignment time and scores them as matches, where plain word-level WER counts each as an error.
In Malayalam this happens through vowel elision — എനിക്ക് അറിയാം and എനിക്കറിയാം are both valid writings of the same speech. The effect is largest in Malayalam, the most agglutinative of the four languages: on identical FLEURS-RO output this model scores 14.74 WER_SCRIBE against 22.69 plain WER, a gap of nearly eight points arising entirely from token-boundary treatment.
Because the model emits punctuation and written number forms, error rates on unnormalised rich-orthography references are not directly comparable to WER figures computed on normalised text. See the SCRIBE paper [2] for metric definitions and the FLEURS-RO annotation protocol.
Verbatim results
Public Malayalam benchmarks as published, scored with ER_LEX. Datasets marked * have punctuated references; the rest do not.
Numerals are normalised on both sides with Indic Num2Words before scoring. Without it, a model that writes 15 is charged an error against a reference that spells the number out, which measures formatting convention rather than recognition. Unnormalised figures are on the benchmarks page.
The systems below differ in size, training objective and output format, and saaras-v4 is a commercial API rather than an open model. Benchmark coverage differs by language — Malayalam has no RESPIN or MUCS coverage, so this macro is over five benchmarks rather than the seven used for Hindi. This is a range, not a ranking.
ER_LEX (%). Lower is better. Macro is the unweighted mean across the five benchmarks.
| Model | Macro | IndicVoices | FLEURS-RO* | Kathbath | CommonVoice | IndicTTS |
|---|---|---|---|---|---|---|
| clips (n) | 4,413 | 957 | 1,767 | 69 | 100 | |
indic-conformer-600m-ctc |
18.03 | 31.34 | 13.31 | 16.93 | 17.11 | 11.47 |
indic-conformer-600m-rnnt |
16.32 | 27.39 | 13.54 | 15.62 | 16.22 | 8.85 |
koyal-ml-120m-1.0 |
14.19 | 26.36 | 7.21 | 13.49 | 13.57 | 10.30 |
koyal-indic-600m-1.0 |
19.25 | 35.25 | 10.22 | 18.16 | 17.40 | 15.24 |
sravaani-1.0 |
14.95 | 26.61 | 12.44 | 15.01 | 14.75 | 5.95 |
saaras-v4 (commercial) |
12.82 | 22.42 | 11.77 | 11.88 | 11.50 | 6.53 |
Training data
Fine-tuned on ~1,350 hours of Malayalam speech assembled from public sources, with transcripts curated to rich orthographic form (punctuation and written numerals preserved) following the transcription curation pipeline described in the SCRIBE paper [2].
Sources:
- IndicVoices
- Shrutilipi
- SPRING-INX
- Kathbath
- Vaani
- IMaSC
- IndicTTS
- Google FLEURS
- ULCA
- OpenSLR 63
- Mozilla Common Voice
- Festvox IIITH
Related models
| Model | Language | Params | Modes |
|---|---|---|---|
| koyal-hi-120m-1.0 | Hindi | 120M | Offline |
| koyal-kn-120m-1.0 | Kannada | 120M | Offline |
| koyal-ml-120m-1.0 | Malayalam | 120M | Offline |
| koyal-te-120m-1.0 | Telugu | 120M | Offline |
| koyal-indic-600m-1.0 | hi · kn · ml · te | 600M | Offline + streaming |
References
- Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages (MahaDhwani), ICASSP 2025
- SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR
Citation
@inproceedings{bhogale2025mahadhwani,
author = {Bhogale, Kaushal Santosh and Mehendale, Deovrat and Javed, Tahir and Anuragi, Devbrat and Joshi, Sakshi and Sundaresan, Sai and Ananthanarayanan, Aparna and Dey, Sharmistha and G, Sathish Kumar Reddy and Srinivasan, Anusha and Raman, Abhigyan and Kumar, Pratyush and Khapra, Mitesh M.},
title = {Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages},
booktitle = {ICASSP 2025 -- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2025},
pages = {1--5},
doi = {10.1109/ICASSP49660.2025.10888018}
}
@article{manohar2026scribe,
title = {SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR},
author = {Manohar, Kavya and Bhattacharya, Arghya and Juvekar, Kush and Nethil, Kumarmanas},
journal = {arXiv preprint arXiv:2605.20712},
year = {2026}
}
License
Released under CC-BY-4.0. The encoder is initialised from the MIT-licensed MahaDhwani pretrained Conformer checkpoint by AI4Bharat.
Contact
Questions and feedback: the Community tab of this repository.
- Downloads last month
- 12
Model tree for adalat-ai/koyal-ml-120m-1.0
Base model
ai4bharat/MahaDhwani_pretrained_conformerDataset used to train adalat-ai/koyal-ml-120m-1.0
Space using adalat-ai/koyal-ml-120m-1.0 1
Collection including adalat-ai/koyal-ml-120m-1.0
Paper for adalat-ai/koyal-ml-120m-1.0
Evaluation results
- WER_SCRIBE on FLEURS-RO (Malayalam)test set self-reported14.740