Instructions to use Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-VL-4B-Instruct") model = PeftModel.from_pretrained(base_model, "Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora") - Notebooks
- Google Colab
- Kaggle
Qwen3-VL-4B Multi-Image Wikipedia Screenshot QA LoRA β 3x compression (top-3 retrieval)
LoRA adapter for Qwen3-VL-4B-Instruct fine-tuned as the reader in a retrieval-augmented QA pipeline: given a question and 3 candidate Wikipedia screenshots (one of which contains the answer, two are hard distractors), the model must locate the right image and extract the answer β all under 3x pixel compression.
Performance (GPT-4.1 LLM-judge on 500 test examples)
| Setup | LLM-judge |
|---|---|
| Uncompressed (0x) single-image, base Qwen3-VL-4B | 0.958 |
| Uncompressed (0x) multi-image (3), base Qwen3-VL-4B | 0.892 |
| This adapter @ 3x multi-image (3) | 0.900 |
Key observation: Despite 3x pixel compression, this model slightly exceeds the un-SFTed multi-image uncompressed base (0.892) β fine-tuning compensates for both distractor confusion and 3x compression.
Gain over multi-image base @ 0x: +0.008 (+0.9% relative).
Training setup
- Method: LoRA (r=256, 2ep, lr 1e-5, LLM+ViT LoRA, cutoff_len 4096, eff-batch 32 on 8Γ H100)
- Base model:
Qwen/Qwen3-VL-4B-Instruct - Framework: LLaMA-Factory (fork)
- Hardware: 8Γ H100 80GB, DeepSpeed ZeRO-2, bf16
- Checkpoint: step 6502 (2 epochs)
Data β retrieval-augmented multi-image
Training data was built from the Chrisyichuan screenshot-training-natural-filtered-v2 QA dataset:
- For each query, retrieve top-6 screenshots from a Qwen3-VL-2B embedding index (dora-ls005 checkpoint) over 28M Wikipedia tiles.
- Construct a 3-image set: always include the gold + up to 2 non-gold retrieved distractors.
- Randomize gold position among the 3 to avoid positional shortcuts.
- Apply 3x compression (each dimension scaled by
1/sqrt(3)via PIL LANCZOS). - Train the reader to answer the query given the 3 compressed images.
~104k training examples, 5.8k validation, 5.8k test. Gold-retrieval rate at top-6 across splits: ~75%.
Usage
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from peft import PeftModel
import torch
base = "Qwen/Qwen3-VL-4B-Instruct"
adapter = "Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora"
model = Qwen3VLForConditionalGeneration.from_pretrained(base, torch_dtype=torch.bfloat16).cuda()
model = PeftModel.from_pretrained(model, adapter).merge_and_unload()
processor = AutoProcessor.from_pretrained(base)
# Three 3x-compressed images (one gold + two distractors, any order)
messages = [{"role": "user", "content": [
{"type": "image", "image": img1},
{"type": "image", "image": img2},
{"type": "image", "image": img3},
{"type": "text", "text": your_question},
]}]
# ... standard Qwen3-VL inference
Notes / limitations
- Training distribution always includes the gold in the 3-image set (by construction). If your retriever misses the gold, the model has not seen that distribution β expect degradation on those queries.
- Compression level is fixed at 3x. Use the adapter that matches your deployment pixel budget.
- Sister adapters at other compression levels:
Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-2x-lora,Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-4x-lora.
- Downloads last month
- 17
Model tree for Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora
Base model
Qwen/Qwen3-VL-4B-Instruct
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-VL-4B-Instruct") model = PeftModel.from_pretrained(base_model, "Chrisyichuan/qwen3vl-4b-wiki-screenshot-multi3-3x-lora")