Image-Text-to-Text
Transformers
Safetensors
nvidia
VLM
conversational