Instructions to use FineEnvs/gemma-4-E4B-it-kannada-asr-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use FineEnvs/gemma-4-E4B-it-kannada-asr-grpo with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-E4B-it") model = PeftModel.from_pretrained(base_model, "FineEnvs/gemma-4-E4B-it-kannada-asr-grpo") - Notebooks
- Google Colab
- Kaggle
Gemma 4 E4B · Kannada speech recognition (GRPO)
A LoRA adapter for google/gemma-4-E4B-it, trained
with GRPO against an OpenEnv speech-recognition environment
on all 2,282 Kannada training clips in FLEURS.
On the 838 held-out test clips, character error falls 45% and the model stops writing other
Indic scripts into Kannada transcripts.
Part of the Multilingual Multimodal Envs collection. The environment is live at FineEnvs/fleurs-asr-env.
Results
FLEURS kn_in test, all 838 clips, greedy decoding through vLLM. Each clip is scored by the
environment that produced the training reward, and the change is paired clip by clip against the
untuned base model.
| base | this adapter | change (95% CI) | |
|---|---|---|---|
| CER, per clip, capped at 1 | 0.1047 | 0.0571 | −0.048 (−0.057, −0.039) |
| CER, mean | 0.1335 | 0.0828 | −0.051 (−0.062, −0.039) |
| CER, median | 0.0676 | 0.0399 | |
| WER, per clip, capped at 1 | 0.312 | 0.243 | |
| WER, mean | 0.358 | 0.288 | |
| exact transcripts | 3.3% | 8.0% | |
| transcripts containing another Indic script | 174 (20.8%) | 13 (1.6%) | |
| clips that loop (CER above 1) | 10 | 1 |
565 clips get better, 139 get worse and 134 are unchanged. The capped CER is the headline because a single transcript that loops can add five units of error to the mean. Most of the gain arrives in the first 150 steps, and the curve is flat from about step 250.
What it learned
It learned to stay in Kannada. One transcript in five from the untuned model switches script mid-word, writing a Gujarati or Malayalam letter where the Kannada one belongs. After training it is one in sixty. Three clips, picked at the 25th, 50th and 75th percentile of improvement, not by hand:
| text | CER | |
|---|---|---|
| reference | ಭಾನುವಾರ ತಡರಾತ್ರಿಯಲ್ಲಿ … ಡೊನಾಲ್ಡ್ ಟ್ರಂಪ್ ಅವರ ಪತ್ರಿಕಾ ಕಾರ್ಯದರ್ಶಿ ಮೂಲಕ ನೀಡಿದ ಹೇಳಿಕೆಯಲ್ಲಿ … | |
| base | ಬಾನುವಾರ ತಡರಾತ್ರಿಯಲ್ಲಿ … ಡೊನಾಲ್ಡ್ ಟ್ರಂಪ್ ಅವರ ಪત્રಕಾರ್ಯದರ್ಶಿ ಮೂಲಕ … ಸಿರಿಯಾವನ್ನು ಸೋಲಿಸುವುdagī ಘೋಷಿಸಿತು | 0.106 |
| tuned | ಬಾನುವಾರ ತಡರಾತ್ರಿಯಲ್ಲಿ … ಡೊನಾಲ್ಡ್ ಟ್ರಂಪ್ ಅವರ ಪತ್ರಿಕಾ ಕಾರ್ಯದರ್ಶಿ ಮೂಲಕ … ಸಿರಿಯಾವನ್ನು ಸೋರುವುದಾಗಿ ಘೋಷಿಸಿತು | 0.031 |
| reference | ತಾಂತ್ರಿಕ ನಿರ್ಣಾಯಕತೆಯ ಹೆಚ್ಚಿನ ವ್ಯಾಖ್ಯಾನಗಳು … ಅವುಗಳೆಂದರೆ ತಂತ್ರಜ್ಞಾನದ … | |
| base | ತಾന്ത്രിಕ ನಿರ್ಣಾಯಕತೀಯ ಹೆಚ್ಚಿನ ವ್ಯಾಖ್ಯಾನಗಳು … ಅವು ಎಂದರೆ ತಂತ್ರಜ್ಞಾನದ … | 0.047 |
| tuned | ತಾಂತ್ರಿಕ ನಿರ್ಣಾಯಕತೆಯೇ ಹೆಚ್ಚಿನ ವ್ಯಾಖ್ಯಾನಗಳು … ಅವುಗಳೆಂದರೆ ತಂತ್ರಜ್ಞಾನದ … | 0.010 |
| reference | … ಮೂರು ಭಾಗಗಳನ್ನಾಗಿ ವಿಭಾಗಿಸಿದಾಗ ಮೂಡುವ ರೇಖೆಗಳ ಪರಸ್ಪರ ಹಾಯುವಿಕೆಯ ಜಾಗವಾಗಿದೆ | |
| base | … ಮೂರು ಭಾಗಗಳನ್ನಾಗಿ ವಿಭಾಗಿಸಿದಾಗ ಮೂರು ವರೆಕೆಗಳ ಪರಸ್ಪರ ಹಾಯುವಿಕೆಯ ಜಾಗವಾಗಿದೆ | 0.032 |
| tuned | … ಮೂರು ಭಾಗಗಳನ್ನಾಗಿ ವಿಭಾಗಿಸಿದಾಗ ಮೂಡುವ ರೇಖೆಗಳ ಪರಸ್ಪರ ಹಾಯುವಿಕೆಯ ಜಾಗವಾಗಿದೆ | 0.013 |
Use it
With vLLM, serving the adapter over the base model:
vllm serve google/gemma-4-E4B-it --enable-lora --max-lora-rank 16 \
--limit-mm-per-prompt '{"audio": 1}' \
--lora-modules kannada-asr=FineEnvs/gemma-4-E4B-it-kannada-asr-grpo
import base64, requests
wav = base64.b64encode(open("clip.wav", "rb").read()).decode() # 16 kHz mono
reply = requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "kannada-asr",
"temperature": 0,
"messages": [{"role": "user", "content": [
{"type": "input_audio", "input_audio": {"data": wav, "format": "wav"}},
{"type": "text", "text": "Transcribe this Kannada speech. Return only the transcription, "
"in Kannada, lowercased and without punctuation."},
]}],
}).json()
print(reply["choices"][0]["message"]["content"])
With transformers and PEFT:
import soundfile as sf
import torch
from peft import PeftModel
from transformers import AutoProcessor, Gemma4ForConditionalGeneration
processor = AutoProcessor.from_pretrained("google/gemma-4-E4B-it")
model = Gemma4ForConditionalGeneration.from_pretrained(
"google/gemma-4-E4B-it", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, "FineEnvs/gemma-4-E4B-it-kannada-asr-grpo")
audio, rate = sf.read("clip.wav", dtype="float32") # 16 kHz mono
inputs = processor.apply_chat_template([{"role": "user", "content": [
{"type": "audio", "audio": audio},
{"type": "text", "text": "Transcribe this Kannada speech. Return only the transcription, "
"in Kannada, lowercased and without punctuation."},
]}], add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=448, do_sample=False)
print(processor.batch_decode(output[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
Use the prompt above: it is the one the environment serves, and the one the adapter was trained on. To hear a clip and score your own transcript, try the playground.
How it was trained
| base model | google/gemma-4-E4B-it |
| method | GRPO (TRL 1.13), LoRA r=16, α=32 on the attention q/v projections of the language model (66 modules) |
| data | all 2,282 FLEURS kn_in training clips, one epoch (575 steps × 4 clips) |
| rollouts | 16 transcripts per clip, temperature 0.9, up to 448 tokens |
| reward | 0.8 × max(0, 1 − CER) + 0.2 × exact match, computed by the environment server |
| optimiser | lr 5e-5, 29 warmup steps, cosine decay, no KL penalty |
| hardware | one A100 80GB on HF Jobs, 4.7 hours |
| selection | the final checkpoint: best on capped CER of all 23 scored checkpoints |
The audio has to reach the loss. TRL 1.13 passes image features into the forward pass its loss
is computed from, but drops audio features. A stock GRPO run therefore trains on how likely each
transcript is without the clip, which teaches a language model the training sentences rather
than how to listen. The gap is large: the model's own transcripts score −0.38 nats per token with
the clip and −6.22 without it. This run uses AudioGRPOTrainer, which carries the clip into every
log-prob pass and checks on the first batch that it arrives. Earlier Kannada runs without it raised
training reward and barely moved held-out scores.
The trainer, the environment and both job launchers are in
06-multilingual/asr/
on GitHub. Every checkpoint's held-out scores, the curves and the exact job scripts are in
FineEnvs/multilingual-multimodal-rl-runs,
and the training and evaluation runs are in the
Trackio dashboard.
Limits
One language and one dataset. FLEURS is read speech, mostly Wikipedia sentences, so this says little about conversational or noisy Kannada. Word error stays high: Kannada words are long and inflected, and most remaining errors are one or two letters inside a word.
Citation
@misc{fineenvs,
author = {Kolavi, Adithya S},
title = {FineEnvs: Open Source RL Environments for LLM Agents},
year = {2026},
url = {https://github.com/adithya-s-k/FineEnvs}
}
@inproceedings{conneau2023fleurs,
title = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and
Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
booktitle = {IEEE Spoken Language Technology Workshop (SLT)},
year = {2023}
}
- Downloads last month
- 17
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-E4B-it") model = PeftModel.from_pretrained(base_model, "FineEnvs/gemma-4-E4B-it-kannada-asr-grpo")