Gemma 4 E4B · Kannada speech recognition (GRPO)

Training reward and held-out CER/WER over 575 GRPO steps

A LoRA adapter for google/gemma-4-E4B-it, trained with GRPO against an OpenEnv speech-recognition environment on all 2,282 Kannada training clips in FLEURS. On the 838 held-out test clips, character error falls 45% and the model stops writing other Indic scripts into Kannada transcripts.

Part of the Multilingual Multimodal Envs collection. The environment is live at FineEnvs/fleurs-asr-env.

Results

FLEURS kn_in test, all 838 clips, greedy decoding through vLLM. Each clip is scored by the environment that produced the training reward, and the change is paired clip by clip against the untuned base model.

base this adapter change (95% CI)
CER, per clip, capped at 1 0.1047 0.0571 −0.048 (−0.057, −0.039)
CER, mean 0.1335 0.0828 −0.051 (−0.062, −0.039)
CER, median 0.0676 0.0399
WER, per clip, capped at 1 0.312 0.243
WER, mean 0.358 0.288
exact transcripts 3.3% 8.0%
transcripts containing another Indic script 174 (20.8%) 13 (1.6%)
clips that loop (CER above 1) 10 1

565 clips get better, 139 get worse and 134 are unchanged. The capped CER is the headline because a single transcript that loops can add five units of error to the mean. Most of the gain arrives in the first 150 steps, and the curve is flat from about step 250.

What it learned

It learned to stay in Kannada. One transcript in five from the untuned model switches script mid-word, writing a Gujarati or Malayalam letter where the Kannada one belongs. After training it is one in sixty. Three clips, picked at the 25th, 50th and 75th percentile of improvement, not by hand:

text CER
reference ಭಾನುವಾರ ತಡರಾತ್ರಿಯಲ್ಲಿ … ಡೊನಾಲ್ಡ್ ಟ್ರಂಪ್ ಅವರ ಪತ್ರಿಕಾ ಕಾರ್ಯದರ್ಶಿ ಮೂಲಕ ನೀಡಿದ ಹೇಳಿಕೆಯಲ್ಲಿ …
base ಬಾನುವಾರ ತಡರಾತ್ರಿಯಲ್ಲಿ … ಡೊನಾಲ್ಡ್ ಟ್ರಂಪ್ ಅವರ ಪત્રಕಾರ್ಯದರ್ಶಿ ಮೂಲಕ … ಸಿರಿಯಾವನ್ನು ಸೋಲಿಸುವುdagī ಘೋಷಿಸಿತು 0.106
tuned ಬಾನುವಾರ ತಡರಾತ್ರಿಯಲ್ಲಿ … ಡೊನಾಲ್ಡ್ ಟ್ರಂಪ್ ಅವರ ಪತ್ರಿಕಾ ಕಾರ್ಯದರ್ಶಿ ಮೂಲಕ … ಸಿರಿಯಾವನ್ನು ಸೋರುವುದಾಗಿ ಘೋಷಿಸಿತು 0.031
reference ತಾಂತ್ರಿಕ ನಿರ್ಣಾಯಕತೆಯ ಹೆಚ್ಚಿನ ವ್ಯಾಖ್ಯಾನಗಳು … ಅವುಗಳೆಂದರೆ ತಂತ್ರಜ್ಞಾನದ …
base ತಾന്ത്രിಕ ನಿರ್ಣಾಯಕತೀಯ ಹೆಚ್ಚಿನ ವ್ಯಾಖ್ಯಾನಗಳು … ಅವು ಎಂದರೆ ತಂತ್ರಜ್ಞಾನದ … 0.047
tuned ತಾಂತ್ರಿಕ ನಿರ್ಣಾಯಕತೆಯೇ ಹೆಚ್ಚಿನ ವ್ಯಾಖ್ಯಾನಗಳು … ಅವುಗಳೆಂದರೆ ತಂತ್ರಜ್ಞಾನದ … 0.010
reference … ಮೂರು ಭಾಗಗಳನ್ನಾಗಿ ವಿಭಾಗಿಸಿದಾಗ ಮೂಡುವ ರೇಖೆಗಳ ಪರಸ್ಪರ ಹಾಯುವಿಕೆಯ ಜಾಗವಾಗಿದೆ
base … ಮೂರು ಭಾಗಗಳನ್ನಾಗಿ ವಿಭಾಗಿಸಿದಾಗ ಮೂರು ವರೆಕೆಗಳ ಪರಸ್ಪರ ಹಾಯುವಿಕೆಯ ಜಾಗವಾಗಿದೆ 0.032
tuned … ಮೂರು ಭಾಗಗಳನ್ನಾಗಿ ವಿಭಾಗಿಸಿದಾಗ ಮೂಡುವ ರೇಖೆಗಳ ಪರಸ್ಪರ ಹಾಯುವಿಕೆಯ ಜಾಗವಾಗಿದೆ 0.013

Use it

With vLLM, serving the adapter over the base model:

vllm serve google/gemma-4-E4B-it --enable-lora --max-lora-rank 16 \
  --limit-mm-per-prompt '{"audio": 1}' \
  --lora-modules kannada-asr=FineEnvs/gemma-4-E4B-it-kannada-asr-grpo
import base64, requests

wav = base64.b64encode(open("clip.wav", "rb").read()).decode()  # 16 kHz mono
reply = requests.post("http://localhost:8000/v1/chat/completions", json={
    "model": "kannada-asr",
    "temperature": 0,
    "messages": [{"role": "user", "content": [
        {"type": "input_audio", "input_audio": {"data": wav, "format": "wav"}},
        {"type": "text", "text": "Transcribe this Kannada speech. Return only the transcription, "
                                 "in Kannada, lowercased and without punctuation."},
    ]}],
}).json()
print(reply["choices"][0]["message"]["content"])

With transformers and PEFT:

import soundfile as sf
import torch
from peft import PeftModel
from transformers import AutoProcessor, Gemma4ForConditionalGeneration

processor = AutoProcessor.from_pretrained("google/gemma-4-E4B-it")
model = Gemma4ForConditionalGeneration.from_pretrained(
    "google/gemma-4-E4B-it", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, "FineEnvs/gemma-4-E4B-it-kannada-asr-grpo")

audio, rate = sf.read("clip.wav", dtype="float32")  # 16 kHz mono
inputs = processor.apply_chat_template([{"role": "user", "content": [
    {"type": "audio", "audio": audio},
    {"type": "text", "text": "Transcribe this Kannada speech. Return only the transcription, "
                             "in Kannada, lowercased and without punctuation."},
]}], add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=448, do_sample=False)
print(processor.batch_decode(output[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])

Use the prompt above: it is the one the environment serves, and the one the adapter was trained on. To hear a clip and score your own transcript, try the playground.

How it was trained

One GRPO step, stage by stage, beside the code that runs it
base model google/gemma-4-E4B-it
method GRPO (TRL 1.13), LoRA r=16, α=32 on the attention q/v projections of the language model (66 modules)
data all 2,282 FLEURS kn_in training clips, one epoch (575 steps × 4 clips)
rollouts 16 transcripts per clip, temperature 0.9, up to 448 tokens
reward 0.8 × max(0, 1 − CER) + 0.2 × exact match, computed by the environment server
optimiser lr 5e-5, 29 warmup steps, cosine decay, no KL penalty
hardware one A100 80GB on HF Jobs, 4.7 hours
selection the final checkpoint: best on capped CER of all 23 scored checkpoints

The audio has to reach the loss. TRL 1.13 passes image features into the forward pass its loss is computed from, but drops audio features. A stock GRPO run therefore trains on how likely each transcript is without the clip, which teaches a language model the training sentences rather than how to listen. The gap is large: the model's own transcripts score −0.38 nats per token with the clip and −6.22 without it. This run uses AudioGRPOTrainer, which carries the clip into every log-prob pass and checks on the first batch that it arrives. Earlier Kannada runs without it raised training reward and barely moved held-out scores.

The trainer, the environment and both job launchers are in 06-multilingual/asr/ on GitHub. Every checkpoint's held-out scores, the curves and the exact job scripts are in FineEnvs/multilingual-multimodal-rl-runs, and the training and evaluation runs are in the Trackio dashboard.

Limits

One language and one dataset. FLEURS is read speech, mostly Wikipedia sentences, so this says little about conversational or noisy Kannada. Word error stays high: Kannada words are long and inflected, and most remaining errors are one or two letters inside a word.

Citation

@misc{fineenvs,
  author = {Kolavi, Adithya S},
  title  = {FineEnvs: Open Source RL Environments for LLM Agents},
  year   = {2026},
  url    = {https://github.com/adithya-s-k/FineEnvs}
}

@inproceedings{conneau2023fleurs,
  title     = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
  author    = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and
               Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
  booktitle = {IEEE Spoken Language Technology Workshop (SLT)},
  year      = {2023}
}
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FineEnvs/gemma-4-E4B-it-kannada-asr-grpo

Adapter
(379)
this model

Dataset used to train FineEnvs/gemma-4-E4B-it-kannada-asr-grpo

Collection including FineEnvs/gemma-4-E4B-it-kannada-asr-grpo