MadriMed 1.2

MadriSight Logo

2B Multimodal Medical Vision-Language Model

2B Parameters Qwen3-VL Base GRPO Training Apple Silicon


Overview

MadriMed 1.2 is the second compact multimodal medical vision-language model by Madrisight. Building upon madriMed-VL-2B, MadriMed 1.2 delivers significant performance leaps across medical image understanding (Radiology, Pathology, CT, MRI, X-Ray) and medical reasoning.

It Delivers robust clinical reasoning capabilities, drastically reducing hallucinations and diagnostic contradictions while setting a new state-of-the-art benchmark for 2B medical vision models.

What's New in MadriMed 1.2?

  1. Massive Diagnostic Accuracy Gains:
    • Massive jump on VQA-RAD overall score and most of the other evalution benchmarks.
  2. Resistant to Generation Hallucinations:
    • Substantially reduced token corruption, repetition and positive-bias tendencies observed in version one.
  3. Structured Formating:
    • Trained with graduated correctness scoring and reasoning-integrity penalties to enforce structured.

Quick Start

Installation

pip install torch torchvision transformers

Input

prompt = """Question: Examine the mammogram image shown above. Which of the following findings is most evident?

Answer Choices:
A. Well-circumscribed round mass with benign features
B. Clustered microcalcifications within an area of irregular density
C. Fat-containing lesion consistent with lipoma
D. Diffuse bilateral breast edema

Follow this exact step-by-step reasoning structure.

### 1. Key Clinical Findings
Extract the patient's age, gender, main symptoms, duration, objective physical exam findings, and abnormal laboratory/imaging results from the provided question and the images.

### 2. Differential Diagnosis
List 2-3 highly probable conditions based on the extracted findings. Briefly state why the current presentation fits.

### 3. Option Elimination 
Analyze each provided choice. Write one sentence explaining why it is either correct or incorrect based on clinical guidelines.

### 4. Conclusion & Final Answer
State the final diagnostic/management conclusion.

### 5. Final correct choice
Provide final single correct choice.
"""

Input Image A

Input Image B

Run the model directly

import torch
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from PIL import Image

MODEL_ID = "madrisight/madrimed1.2-VL-2B"
device = "mps" if torch.backends.mps.is_available() else "cpu"

model = Qwen3VLForConditionalGeneration.from_pretrained(
    MODEL_ID,
    device_map=device,
    dtype=torch.float32,
    trust_remote_code=True,
).eval()

processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)

# Multi-image input
image_paths = ["images/MM-1-a.png", "images/MM-1-b.png"]
images = [Image.open(path).convert("RGB") for path in image_paths]

# Chat Message Payload
messages = [
    {
        "role": "system",
        "content": "You are an expert clinical medical educator and USMLE board reviewer.",
    },
    {
        "role": "user",
        "content": [{"type": "image"} for _ in images] + [{"type": "text", "text": prompt}]
    }
]

# Tokenize, Process, and Generate
with torch.inference_mode():
    text = processor.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True
    )

    inputs = processor(
        text=text,
        images=images,  # Pass the flat list of PIL Image objects
        return_tensors="pt",
        padding=True,
        truncation=True,
    ).to(device)

    generated_ids = model.generate(
        **inputs,
        max_new_tokens=1024,
        do_sample=False,
    )

    output_text = processor.batch_decode(
        generated_ids[:, inputs.input_ids.shape[1]:],
        skip_special_tokens=True
    )[0].strip()

Output

### 1. Key Clinical Findings
- **Patient's age and gender**: Not explicitly mentioned, but mammograms are typically used for women.
- **Main symptoms**: Not specified, but mammograms are often performed for breast abnormalities.
- **Duration**: Not specified, but mammograms are usually performed at the time of suspicion.
- **Objective physical exam findings**: Not specified, but mammograms are used to detect abnormalities.
- **Abnormal laboratory/imaging results**: Not specified, but mammograms are used to detect breast abnormalities.

### 2. Differential Diagnosis
- **Well-circumscribed round mass with benign features**: This could be a benign lesion such as a fibroadenoma.
- **Clustered microcalcifications within an area of irregular density**: This could indicate a malignant process, such as ductal carcinoma in situ (DCIS) or invasive breast cancer.
- **Fat-containing lesion consistent with lipoma**: This is a benign lesion and could be a lipoma.

### 3. Option Elimination
- **Well-circumscribed round mass with benign features**: This is a plausible finding, but it is less likely to be the primary concern in a mammogram without additional symptoms or findings.
- **Clustered microcalcifications within an area of irregular density**: This is a highly suspicious finding for malignancy, as it is often associated with breast cancer.
- **Fat-containing lesion consistent with lipoma**: This is a benign finding and less likely to be the primary concern in a mammogram without additional symptoms or findings.

### 4. Conclusion & Final Answer
The mammogram findings are most indicative of a malignant process, specifically breast cancer, due to the presence of clustered microcalcifications within an area of irregular density.

### 5. Final correct choice
B. Clustered microcalcifications within an area of irregular density

Benchmark Results

MadriMed 1.2 was rigorously evaluated across 7 medical image and text benchmarks.

🔬 Multimodal & Visual Question Answering

Benchmark MediX-R1 (2B) Qwen3.5-2B-MedVL (OpenMed) MadriMed 1.2 (2B)
VQA-RAD 53.9% 55.21% 64.30%
PathVQA 42.8% 47.48% 48.29%
SLAKE (Radiology) 65.4% 68.80% 69.36%

VQA Comparison with 4B

Benchmark Metric MediX-R1 2B Qwen3.5-2B-MedVL MedMO-4B-Next MedGemma 1.5 4B MedGemma 4B MadriMed 1.2 2B
SLAKE Overall 65.40% 68.80% 74.00% 59.80% Tok-F1 72.30% Tok-F1 69.36%
↳ Closed Accuracy 68.40% N/R 78.00% 82.80% 87.60% 71.83%
↳ Open — 57.60% N/R — 59.70% Tok-F1 — 68.13%
VQA-RAD Overall 53.90% 55.21% 59.60% 48.10% Tok-F1 49.90% Tok-F1 64.30%
↳ Closed Accuracy 61.30% N/R 79.70% 70.20% 69.10% 74.90%
↳ Open — 45.10% N/R — — — 51.00%
PathVQA Overall 42.80% 47.48% 73.30% — — 48.29%
MedXpertQA (MM+Text) Accuracy — — — 20.90% 18.80% 18.02%
Benchmark Metric MedMO-4B-Next MedGemma (4B) MadriMed-VL-2B (v1) MadriMed 1.2 (2B)
MedXpertQA (Multimodal) Accuracy 27.0% 24.43% 21.50% 23.50%
MedXpertQA (Text) Accuracy 16.5% 14.2% 11.02% 13.55%

📝 Text only Benchmarks

Benchmark MedGemma 1.5 (4B) MedGemma (4B) MediX-R1 (2B) MadriMed 1.2 (2B)
PubMedQA 68.2% 73.4% 47.2% 66.6%
MedMCQA 55.7% 59.8% 49.2% 43.1%
MedQA-4opts 69.1% 64.4% 49.7% 41.56%
MedXpertQA (Text) 14.2% 13.55%

SLAKE Per-Category Performance

Category EM Total VQA
Abnormality 40.4% 52
Color 81.8% 33
KG 57.8% 109
Modality 91.8% 85
Organ 71.7% 99
Plane 90.9% 33
Position 61.4% 171
Quantity 73.1% 52
Shape 42.9% 7
Size 69.2% 65

Medmcqa (Text-Only) Per-Category Performance

By Choice Type
Choice Type Accuracy Correct / Total
Multi 41.62% 569 / 1,367
Single 43.82% 1,234 / 2,816
By Subject
Subject Accuracy Correct / Total
Anaesthesia 41.18% 14 / 34
Anatomy 44.02% 103 / 234
Biochemistry 50.29% 86 / 171
Dental 39.83% 525 / 1,318
ENT 32.08% 17 / 53
Forensic Medicine 37.31% 25 / 67
Gynaecology & Obstetrics 43.30% 97 / 224
Medicine 46.44% 137 / 295
Microbiology 42.62% 52 / 122
Ophthalmology 41.38% 24 / 58
Orthopaedics 30.00% 6 / 20
Pathology 48.96% 165 / 337
Pediatrics 40.60% 95 / 234
Pharmacology 44.03% 107 / 243
Physiology 55.56% 95 / 171
Psychiatry 37.50% 6 / 16
Radiology 53.62% 37 / 69
Skin 29.41% 5 / 17
Social & Preventive Medicine 42.64% 55 / 129
Surgery 41.19% 152 / 369

Evaluation NoteBook

Open Notebook


Technical Specifications

Component Details
Base Architecture Qwen/Qwen3-VL-2B-Thinking
Context & Vision Native multimodal support with Dynamic M-RoPE
Dataset mint-medmax/medmax_data and synthetic data
Precision Native float32
Training Platform Apple Silicon chip
Training Framework trl with custom fused kernels
Primary Goal Memory-stable multimodal RL post training and high accuracy benchmarks

Benchmark Dataset Citations

  1. SLAKE: BoKelvin/SLAKE
  2. MedXpertQA: TsinghuaC3I/MedXpertQA
  3. VQA-RAD: flaviagiammarino/vqa-rad
  4. PubMedQA: qiaojin/PubMedQA
  5. MedQA (USMLE): GBaker/MedQA-USMLE-4-options
  6. MedMCQA: openlifescienceai/medmcqa
  7. Path-VQA: flaviagiammarino/path-vqa

Limitations & Safety Guidelines

Research & Educational Use Only

MadriMed 1.2 is an experimental research model. It is not certified for autonomous clinical diagnostic decision-making.

No Treatment Prescriptions

The model should not be used as an authoritative source for:

  • Pharmaceutical dosages
  • Surgical interventions
  • Direct treatment pathways
  • Other high-stakes clinical decisions

All model-generated interpretations, medical reasoning, and visual inferences should be reviewed and independently confirmed by a qualified, licensed medical professional before being used in a clinical context.


Citation

@software{madrimed1.2-VL-2B,
  title = {MadriMed 1.2: Advanced Multimodal Medical Reasoning at the 2B Scale},
  author = {Madrisight},
  year = {2026},
  url = {https://huggingface.co/madrisight/madrimed1.2-VL-2B}
}
Downloads last month
120
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for madrisight/madrimed1.2-VL-2B

Finetuned
(24)
this model
Quantizations
2 models