Qwen3-VL-4B Multi-Image Wikipedia Screenshot QA LoRA — 4x compression (top-3 retrieval)

LoRA adapter for Qwen3-VL-4B-Instruct fine-tuned as the reader in a retrieval-augmented QA pipeline: given a question and 3 candidate Wikipedia screenshots (one of which contains the answer, two are hard distractors), the model must locate the right image and extract the answer — all under 4x pixel compression.

Performance (GPT-4.1 LLM-judge on 500 test examples)

Setup LLM-judge
Uncompressed (0x) single-image, base Qwen3-VL-4B 0.958
Uncompressed (0x) multi-image (3), base Qwen3-VL-4B 0.892
This adapter @ 4x multi-image (3) 0.868

Key observation: At 4x compression the per-image pixel budget is ~¼ the original. Even with SFT, accuracy falls slightly below the un-SFTed multi-image 0x base (0.892). This checkpoint documents the compression–quality frontier at the extreme end; use 2x/3x for production.

Gain over multi-image base @ 0x: -0.024 (-2.7% relative).

Training setup

  • Method: LoRA (r=256, 2ep, lr 1e-5, LLM+ViT LoRA, cutoff_len 3072, eff-batch 32 on 8× H100)
  • Base model: Qwen/Qwen3-VL-4B-Instruct
  • Framework: LLaMA-Factory (fork)
  • Hardware: 8× H100 80GB, DeepSpeed ZeRO-2, bf16
  • Checkpoint: step 6502 (2 epochs)

Data — retrieval-augmented multi-image

Training data was built from the Chrisyichuan screenshot-training-natural-filtered-v2 QA dataset:

  1. For each query, retrieve top-6 screenshots from a Qwen3-VL-2B embedding index (dora-ls005 checkpoint) over 28M Wikipedia tiles.
  2. Construct a 3-image set: always include the gold + up to 2 non-gold retrieved distractors.
  3. Randomize gold position among the 3 to avoid positional shortcuts.
  4. Apply 4x compression (each dimension scaled by 1/sqrt(4) via PIL LANCZOS).
  5. Train the reader to answer the query given the 3 compressed images.

~104k training examples, 5.8k validation, 5.8k test. Gold-retrieval rate at top-6 across splits: ~75%.

Usage

from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from peft import PeftModel
import torch

base = "Qwen/Qwen3-VL-4B-Instruct"
adapter = "Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-4x-lora"

model = Qwen3VLForConditionalGeneration.from_pretrained(base, torch_dtype=torch.bfloat16).cuda()
model = PeftModel.from_pretrained(model, adapter).merge_and_unload()
processor = AutoProcessor.from_pretrained(base)

# Three 4x-compressed images (one gold + two distractors, any order)
messages = [{"role": "user", "content": [
    {"type": "image", "image": img1},
    {"type": "image", "image": img2},
    {"type": "image", "image": img3},
    {"type": "text",  "text": your_question},
]}]
# ... standard Qwen3-VL inference

Notes / limitations

  • Training distribution always includes the gold in the 3-image set (by construction). If your retriever misses the gold, the model has not seen that distribution — expect degradation on those queries.
  • Compression level is fixed at 4x. Use the adapter that matches your deployment pixel budget.
  • Sister adapters at other compression levels: Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-2x-lora, Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora.
Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-4x-lora

Adapter
(214)
this model