ViVQA-BEiT3 โ Expert C (YOLO-augmented)
BEiT-3 base (patch16, 224) finetuned on ViVQA (Vietnamese VQA, 218 answer classes). This is the YOLO-counting-augmented expert from the question-type routing study.
Results (test, 3001 samples)
| Total | what/other | counting | color | location |
|---|---|---|---|---|
| 69.58% | 72.09 | 56.53 | 76.96 | 66.72 |
Training
- Init:
beit3_base_indomain_patch16_224, XLM-R (beit3.spm) tokenization - 30 epochs, lr 5e-5, layer_decay 0.8, batch 64, seed 42,
--clip_grad 1.0,--randaug --augmented yolo (4,228 YOLOv8m counting QA, loss_weight 0.5)
Files
checkpoint-best.pthโ best-val checkpoint (modelkey in state_dict)
Load
from huggingface_hub import hf_hub_download
ckpt = hf_hub_download("ThucPD/vivqa-beit3-routing-C-yolo", "checkpoint-best.pth")
# use with BEiT-3 run_beit3_finetuning.py --model beit3_base_patch16_224 --finetune <ckpt>
See the routing study: per-type augmentation effects reproduce but are sub-noise; this clean
baseline is the robust optimum. Paired with the YOLO-augmented expert ThucPD/vivqa-beit3-routing-A-clean.