AndrewMcDowell
/

wav2vec2-xls-r-300m-japanese

Automatic Speech Recognition

Generated from Trainer

hf-asr-leaderboard

mozilla-foundation/common_voice_8_0

robust-speech-event

Model card Files Files and versions

AndrewMcDowell commited on Feb 4, 2022

Commit

fe7dd5e

·

1 Parent(s): 700a297

Update README.md

Add eval results on dev data.

Files changed (1) hide show

README.md +27 -2

README.md CHANGED Viewed

@@ -27,7 +27,20 @@ model-index:
        - name: Test CER
          type: cer
          value: 23.64
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
@@ -37,7 +50,13 @@ should probably proofread and complete it, then remove this comment. -->
 This model is a fine-tuned version of [facebook/wav2vec2-xls-r-300m](https://huggingface.co/facebook/wav2vec2-xls-r-300m) on the MOZILLA-FOUNDATION/COMMON_VOICE_8_0 - JA dataset.
-Kanji are converted into Hiragana using the [pykakasi](https://pykakasi.readthedocs.io/en/latest/index.html) library during training and evaluation. The model can output both Hiragana and Katakana characters.
 It achieves the following results on the evaluation set:
 - Loss: 0.5212
@@ -98,4 +117,10 @@ The following hyperparameters were used during training:
 ```bash
 python ./eval.py --model_id AndrewMcDowell/wav2vec2-xls-r-300m-japanese --dataset mozilla-foundation/common_voice_8_0 --config ja --split test --log_outputs
 ```

        - name: Test CER
          type: cer
          value: 23.64
+  - task:
+      name: Automatic Speech Recognition
+      type: automatic-speech-recognition
+    dataset:
+      name: Robust Speech Event - Dev Data
+      type: speech-recognition-community-v2/dev_data
+      args: de
+    metrics:
+       - name: Test WER
+         type: wer
+         value: 1.0
+       - name: Test CER
+         type: cer
+         value: 30.99
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 This model is a fine-tuned version of [facebook/wav2vec2-xls-r-300m](https://huggingface.co/facebook/wav2vec2-xls-r-300m) on the MOZILLA-FOUNDATION/COMMON_VOICE_8_0 - JA dataset.
+Kanji are converted into Hiragana using the [pykakasi](https://pykakasi.readthedocs.io/en/latest/index.html) library during training and evaluation. The model can output both Hiragana and Katakana characters. Since there is no spacing, WER is not a suitable metric for evaluating performance and CER is more suitable.
+On mozilla-foundation/common_voice_8_0 it achieved:
+- cer: 23.64%
+On speech-recognition-community-v2/dev_data it achieved:
+- cer: 30.99%
 It achieves the following results on the evaluation set:
 - Loss: 0.5212
 ```bash
 python ./eval.py --model_id AndrewMcDowell/wav2vec2-xls-r-300m-japanese --dataset mozilla-foundation/common_voice_8_0 --config ja --split test --log_outputs
+```
+2. To evaluate on `mozilla-foundation/common_voice_8_0` with split `test`
+```bash
+python ./eval.py --model_id AndrewMcDowell/wav2vec2-xls-r-300m-japanese --dataset speech-recognition-community-v2/dev_data --config de --split validation --chunk_length_s 5.0 --stride_length_s 1.0
 ```