Skip to main content

← All model recipe tasks

Image-Text-to-Text

Vision-language models that answer or generate text from an image and prompt.

Classification follows the Hugging Face task taxonomy. See the Hugging Face Image-Text-to-Text task page for the ecosystem-level task definition.

Model families

Model familyDeclared recipesExact checkpoint examplesRuntime CLI
deepseek_ocr3
deepseek-ai/DeepSeek-OCR-2
trtmc run
internvl4
OpenGVLab/InternVL3-2B-hf
OpenGVLab/InternVL3-8B-hf
trtmc run
lance1
bytedance-research/Lance
trtmc run
locateanything1
nvidia/LocateAnything-3B
trtmc run
phi4_multimodal1
microsoft/Phi-4-multimodal-instruct
trtmc run
qwen_vl4
Qwen/Qwen2.5-VL-3B-Instruct
Qwen/Qwen3-VL-2B-Instruct
trtmc run