You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Koyal Indic-4 Dual-Mode ASR 600M

Koyal is a family of open speech recognition models for Indian languages, built by Adalat AI for document-ready dictation. Koyal models transcribe in rich orthography (RO): the output carries punctuation and formatted numerals as they appear in written documents, rather than a normalised lexical stream. This is the multilingual model, covering Hindi, Kannada, Malayalam and Telugu from a single checkpoint, in both offline and streaming modes. For higher-accuracy single-language offline transcription, see Related models.

At a glance

Field Value
Task Automatic speech recognition, rich orthography
Languages Hindi, Kannada, Malayalam, Telugu
Parameters ~600M
Modes Offline + cache-aware streaming, one checkpoint, operating point set at inference
Base model Nemotron-3.5-ASR-Streaming-0.6B (NVIDIA)
Framework NVIDIA NeMo ≥ 3.0.0, no custom code
WER_SCRIBE (offline) 12.33 hi · 12.75 kn · 18.58 ml · 20.65 te
License OpenMDW-1.1

Quickstart

Requires NVIDIA NeMo ≥ 3.0.0 — the first release with the prompt-conditioned RNN-T model class this checkpoint targets.

pip install "nemo_toolkit[asr]>=3.0.0"

The target language is required and is passed through the input manifest — each entry carries lang and "prompt_mode": "langID". A bare list of file paths will not work.

import json, tempfile
import soundfile as sf
import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.from_pretrained("adalat-ai/koyal-indic-600m-1.0", map_location="cuda")
model.eval()

# Operating point; [-1, -1] = offline, [56, 3] = 240 ms streaming, ...
model.encoder.set_default_att_context_size([-1, -1])

def transcribe(model, wav_paths, lang, batch_size=32):
    """lang: 'hi-IN' | 'kn-IN' | 'ml-IN' | 'te-IN'."""
    with tempfile.NamedTemporaryFile("w", suffix=".jsonl", delete=False) as f:
        for p in wav_paths:
            f.write(json.dumps({
                "audio_filepath": p,
                "duration": sf.info(p).duration,
                "text": "",
                "lang": lang,
                "prompt_mode": "langID",
            }) + "\n")
    return [h.text for h in model.transcribe(f.name, batch_size=batch_size)]

print(transcribe(model, ["clip.wav"], lang="hi-IN"))

Input audio should be 16 kHz mono WAV. Greedy decoding reproduces the reported numbers.

Cache-aware streaming inference

For chunk-wise streaming (the deployment mode of the base encoder), use NeMo's cache-aware streaming script, exactly as for Nemotron-3.5-ASR-Streaming-0.6B. The script ships with the NeMo source tree, not the pip wheel.

git clone https://github.com/NVIDIA-NeMo/NeMo.git && cd NeMo
hf download adalat-ai/koyal-indic-600m-1.0 koyal-indic-600m-1.0.nemo --local-dir models

python examples/asr/asr_cache_aware_streaming/speech_to_text_cache_aware_streaming_infer.py \
    model_path=models/koyal-indic-600m-1.0.nemo \
    dataset_manifest=manifest.json \   # entries need audio_filepath (+ text for WER)
    target_lang=hi-IN \                # hi-IN | kn-IN | ml-IN | te-IN
    att_context_size="[56,13]" \       # streaming points only: [56,0] [56,3] [56,6] [56,13]
    batch_size=32 \
    output_path=out/

target_lang sets the language prompt for every chunk (model.set_inference_prompt). This script only runs the streaming operating points; [-1, -1] (offline) is not a chunked mode — use transcribe() above.

Intended use

Document-ready speech transcription in Hindi, Kannada, Malayalam and Telugu: domains where the transcript is the deliverable and punctuation and numeral formatting must match written convention, such as legal and courtroom dictation. The dual-mode design targets deployments that need both live streaming dictation and offline batch transcription from a single model.

For single-language offline transcription the monolingual Koyal models are more accurate and five times smaller. This model earns its place when four languages, or live streaming, must come from one artifact.

Limitations

  • Reported streaming operating points are streaming-masked full-utterance decodes, not chunk-wise cache-aware inference. Production streaming latency and accuracy characteristics may differ.
  • Trained and evaluated on 16 kHz audio; performance on other sampling rates is not characterised.
  • Not trained for code-switched speech; each utterance is transcribed in a single target language.
  • The target language must be supplied at inference. There is no language identification step.
  • Rich-orthography evaluation covers FLEURS-RO only, which reflects read speech.

Model architecture

Field Value
Architecture Cache-aware FastConformer encoder with prompt-conditioned RNN-T decoder
Parameters ~600M
Encoder Nemotron-3.5-ASR-Streaming-0.6B cache-aware FastConformer, extended during fine-tuning with a full-context offline attention mode alongside its native streaming modes
Decoder Prompt-conditioned RNN-T, trained from scratch
Tokenizer SentencePiece, 6,000 tokens, unified across hi/kn/ml/te (replaces the base model's tokenizer)
Language conditioning 128-dim one-hot language prompt fused per encoder frame
Input 16 kHz mono audio + target language
Output Document-ready text in the target language with punctuation and formatted numerals

The base encoder is NVIDIA's cache-aware streaming FastConformer. Its tokenizer covers Devanagari and Latin but returns unknown tokens for Dravidian scripts (NVIDIA-NeMo/Speech issue #15793), so we replaced it with a 6,000-token SentencePiece vocabulary built for the four target languages and paired the encoder with a fresh prompt-conditioned RNN-T decoder. The released weights are an average of the final training checkpoints.

Dual-mode operation

During fine-tuning the encoder's trained attention-context set is extended with a full-context offline point, training jointly on five configurations so one checkpoint serves both workloads.

Operating point att_context_size Lookahead
Offline [-1, -1] Full utterance
Streaming [56, 13] 1040 ms
Streaming [56, 6] 480 ms
Streaming [56, 3] 240 ms
Streaming [56, 0] 0 ms

The operating point is selected at inference time; no re-training or separate checkpoints are needed to move along the latency–accuracy curve.

Evaluation

Results are reported on two test-set families:

  • rich-orthography, with references curated to carry punctuation and written numerals, scored with WER_SCRIBE; and
  • verbatim, the public benchmarks as published, scored with ER_LEX.

Numbers are not comparable across the two settings. FLEURS-RO appears in both tables — that is the same benchmark scored against different references, so its ER_lex values differ. That is expected, not a discrepancy.

Why the two settings exist

Most public test sets are not punctuated. Across the four languages, only FLEURS-RO (100% in every language), RESPIN (56–59%) and IndicTTS (97% hi, 76% kn, 100% te, but 2% ml) carry punctuation in their references; IndicVoices, Kathbath, MUCS and Common Voice carry none.

Scoring a rich-orthography model on WER_SCRIBE against an unpunctuated reference charges every emitted comma as an insertion — it measures the reference's annotation convention, not the model. The verbatim setting therefore uses ER_LEX, which scores word identity alone and is the task all these systems share.

Rich-orthography Verbatim
References Curated to rich orthography Public benchmarks as published
Metric WER_SCRIBE, with ER_lex / ER_num / ER_punc and jiwer WER ER_LEX only
Question Does the model produce the correct document-ready transcript? Does the model recognise the words?

Rich-orthography results

Evaluated with scribe-eval on FLEURS-RO, the rich-orthography re-annotation of the FLEURS test sets introduced in the SCRIBE paper [1]. ER_lex — lexical tokens. ER_num — number tokens. ER_punc — punctuation tokens. WER_SCRIBE — over all tokens (WER_S in the SCRIBE paper). WER — computed with jiwer on the same unnormalised text.

No text normalisation is applied to references or hypotheses. Lower is better throughout.

WER_SCRIBE (%) on FLEURS-RO by operating point.

Language Offline 1040 ms 480 ms 240 ms 0 ms
Hindi 12.33 13.22 13.53 13.60 14.54
Kannada 12.75 13.39 13.77 14.01 14.80
Malayalam 18.58 19.23 19.62 19.78 20.51
Telugu 20.65 21.63 21.76 22.03 22.87

At zero lookahead the model stays within 2.2 WER_SCRIBE points of full-context offline decoding in every language. The degradation is gradual rather than abrupt, which is what makes a single checkpoint practical for both workloads.

Full results — per-category decomposition, all operating points
Operating point Language ER_lex (%) ER_num (%) ER_punc (%) WER_SCRIBE (%) WER (%)
Offline [-1,-1] hi 9.94 0.34 2.06 12.33 14.56
kn 9.40 0.45 2.90 12.75 18.95
ml 9.95 0.44 8.18 18.58 26.43
te 13.20 0.66 6.79 20.65 26.89
1040 ms [56,13] hi 10.67 0.36 2.19 13.22 15.62
kn 9.94 0.43 3.02 13.39 19.50
ml 10.43 0.49 8.32 19.23 27.17
te 13.87 0.73 7.03 21.63 27.68
480 ms [56,6] hi 10.92 0.34 2.27 13.53 15.92
kn 10.21 0.50 3.07 13.77 19.82
ml 10.75 0.51 8.36 19.62 27.46
te 13.88 0.76 7.12 21.76 28.06
240 ms [56,3] hi 10.97 0.39 2.24 13.60 15.99
kn 10.35 0.48 3.18 14.01 20.14
ml 10.73 0.56 8.49 19.78 27.81
te 14.01 0.78 7.24 22.03 28.15
0 ms [56,0] hi 11.75 0.45 2.35 14.54 16.94
kn 10.93 0.54 3.33 14.80 20.99
ml 11.32 0.61 8.58 20.51 28.72
te 14.47 0.87 7.53 22.87 29.21

Why WER_SCRIBE and WER differ

Alignment in scribe-eval is sandhi-tolerant: agglutination makes word boundaries unstable in Indic text, and the same speech can be validly written as one word or two. scribe-eval detects such merges and splits at alignment time and scores them as matches, where plain word-level WER counts each as an error. The gap is widest in Malayalam, the most agglutinative of the four — 18.58 WER_SCRIBE against 26.43 plain WER on identical offline output.

Because the model emits punctuation and written number forms, error rates on unnormalised rich-orthography references are not directly comparable to WER figures computed on normalised text. See the SCRIBE paper [1] for metric definitions and the FLEURS-RO annotation protocol.

Verbatim results

Public benchmarks as published, scored with ER_LEX and macro-averaged per language. Benchmark coverage differs by language — Hindi 7, Telugu 6, Kannada 5, Malayalam 5.

Numerals are normalised on both sides with Indic Num2Words before scoring. Without it, a model that writes 15 is charged an error against a reference that spells the number out, which measures formatting convention rather than recognition.

The systems below differ in size, training objective and output format, and saaras-v4 is a commercial API rather than an open model. This is a range, not a ranking.

Macro ER_LEX (%) per language. Lower is better.

Model Hindi Kannada Malayalam Telugu Mean
benchmarks (n) 7 5 5 6
indic-conformer-600m-ctc 10.26 19.28 18.03 17.36 16.23
indic-conformer-600m-rnnt 9.71 17.96 16.32 16.56 15.14
Koyal monolingual (120M each) 9.59 18.20 14.19 15.81 14.45
koyal-indic-600m-1.0 12.52 21.15 19.25 20.45 18.34
saaras-v4 (commercial) 7.95 16.67 12.82 13.71 12.79
sravaani-1.0 9.70 16.66 14.95 16.55 14.47

Training data

Fine-tuned on ~6,900 hours of speech across the four target languages, assembled from public sources with transcripts curated to rich orthographic form (punctuation and written numerals preserved) following the transcription curation pipeline described in the SCRIBE paper [1] — the same corpora used to train the monolingual Koyal models.

Language Hours
Hindi ~2,400
Telugu ~1,720
Kannada ~1,440
Malayalam ~1,350

Sources across languages include IndicVoices and IndicVoices-R, Shrutilipi, SeamlessAlign, Kathbath, SPRING-INX, OpenSLR (63, 66, 79, 118, 126), Mann ki Baat, Vaani, IMaSC, IndicTTS, Google FLEURS, ULCA, Mozilla Common Voice and Festvox IIITH. Per-language source breakdowns are on the monolingual model cards. Lahaja is held out entirely for evaluation and contributes no training data.

Related models

Model Language Params Modes
koyal-hi-120m-1.0 Hindi 120M Offline
koyal-kn-120m-1.0 Kannada 120M Offline
koyal-ml-120m-1.0 Malayalam 120M Offline
koyal-te-120m-1.0 Telugu 120M Offline
koyal-indic-600m-1.0 hi · kn · ml · te 600M Offline + streaming

References

  1. SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR
  2. Nemotron-3.5-ASR-Streaming-0.6B — base encoder
  3. Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition — cache-aware streaming architecture

Citation

@article{manohar2026scribe,
  title   = {SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR},
  author  = {Manohar, Kavya and Bhattacharya, Arghya and Juvekar, Kush and Nethil, Kumarmanas},
  journal = {arXiv preprint arXiv:2605.20712},
  year    = {2026}
}

License

Use of this model is governed by the OpenMDW-1.1 license, the same terms as the base encoder Nemotron-3.5-ASR-Streaming-0.6B.

Note that the monolingual Koyal models are CC-BY-4.0; this model carries a different licence because it inherits the base encoder's terms.

Contact

Questions and feedback: the Community tab of this repository.

Downloads last month
144
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adalat-ai/koyal-indic-600m-1.0

Finetuned
(58)
this model

Dataset used to train adalat-ai/koyal-indic-600m-1.0

Collection including adalat-ai/koyal-indic-600m-1.0

Papers for adalat-ai/koyal-indic-600m-1.0

Evaluation results