Image-Text-to-Text
Vision-language models that answer or generate text from an image and prompt.
Classification follows the Hugging Face task taxonomy. See the Hugging Face Image-Text-to-Text task page for the ecosystem-level task definition.
Model families
| Model family | Declared recipes | Exact checkpoint examples | Runtime CLI |
|---|---|---|---|
deepseek_ocr | 3 | deepseek-ai/DeepSeek-OCR-2 | trtmc run |
internvl | 4 | OpenGVLab/InternVL3-2B-hfOpenGVLab/InternVL3-8B-hf | trtmc run |
lance | 1 | bytedance-research/Lance | trtmc run |
locateanything | 1 | nvidia/LocateAnything-3B | trtmc run |
phi4_multimodal | 1 | microsoft/Phi-4-multimodal-instruct | trtmc run |
qwen_vl | 4 | Qwen/Qwen2.5-VL-3B-InstructQwen/Qwen3-VL-2B-Instruct | trtmc run |