Skip to content

Use Multimodal Embedding with NeMo Retriever Library

Text-only NeMo Retriever embedding NIM

You can still use the NeMo Retriever text embedding NIM (OpenAI-compatible embeddings for passage and query vectors) alongside or instead of the multimodal flows on this page. Product and deployment details are in the NeMo Retriever Text Embedding NIM documentation. In library and CLI pipelines, route embedding to that NIM with your configured embed endpoint and model name (refer to the graph pipeline examples for environment-based remote inference).

This documentation describes how to use NeMo Retriever Library with the multimodal embedding model Llama Nemotron Embed VL 1B v2.

The Llama Nemotron Embed VL 1B v2 model is optimized for multimodal question-answering retrieval. The model can embed documents in the form of an image, text, or a combination of image and text. Documents can then be retrieved given a user query in text form. The model supports images that contain text, tables, charts, and infographics.

Example with Default Text-Based Embedding

When you use the multimodal model, by default, all extracted content (text, tables, charts) is treated as plain text. The following example provides a strong baseline for retrieval.

  • The embed method is called with no arguments.

For parameter details, refer to the Python API guide (create_ingestor and .embed()).

from nemo_retriever import create_ingestor

ingestor = (
    create_ingestor(run_mode="batch")
    .files("./data/*.pdf")
    .extract()
    .embed()  # Default behavior embeds all content as text
)
results = ingestor.ingest()

Text inputs that exceed the model limit

For text inputs, including text_image inputs without an image, NeMo Retriever Library checks the complete formatted input against the embedding model's token limit. Image inputs and text_image inputs with an image are outside this text-splitting policy. The check includes the model's document or query prefix and special tokens. The default configured limits are 8,192 tokens for passages and 128 tokens for queries. If the checkpoint declares a smaller supported limit, the checkpoint limit takes precedence.

To set the passage limit in the Python API, use EmbedParams(runtime=ModelRuntimeParams(max_length=N)). The limit includes the model prefix and special tokens and cannot exceed the checkpoint's supported limit. The same text-input policy applies to the hf and vllm backends for both text-only and vision-language embedding models. Oversized text is split losslessly, not truncated.

TextChunkParams.max_tokens controls earlier text chunking, not the final formatted embedding-input limit. EmbedParams.query_max_length controls the separate query limit.

For a registered revision-pinned model, an explicitly revision-pinned model, or a local checkpoint, the library loads the tokenizer and prompt configuration for that exact model version. If the text does not fit, the library splits it into deterministic contiguous token ranges that fit. This split does not truncate text and occurs before either local or remote embedding. If the exact tokenizer, selected prompt prefix, or checkpoint-supported token limit is unavailable, embedding stage setup fails before inference rather than guessing an admission policy.

Each split row preserves the source, page, element, bounding box, and existing document chunk metadata from its parent. The library adds one metadata["embedding_split"] mapping so you can identify and order the embedding-specific children:

  • parent_id identifies the parent content and provenance.
  • chunk_id identifies one deterministic child.
  • chunk_index and chunk_count describe the child order.
  • start_token and end_token describe the source token range.
  • content preserves the child's exact text, including whitespace.

The returned DataFrame can therefore contain more rows than the embedding stage received. Existing fields such as chunk_index and the physical page number keep their original meaning.

Dense LanceDB and collection writes preserve valid split children, including whitespace-only children, and store the complete embedding_split mapping in the JSON metadata field. After decoding that field, use metadata["embedding_split"]["chunk_id"] for the stable embedding child ID. A collection row's top-level chunk_id remains a storage key derived from the document, version, and row index; it is not the embedding child ID.

Local and remote text embedding use the same prepared rows. When this client-side policy is active for a remote endpoint, text requests use truncate="NONE" so the endpoint cannot silently replace the client decision. Image-bearing inputs retain truncate="END". For mixed text_image batches, the library sends text-only and image-bearing inputs in separate requests and preserves result order.

If a backend still rejects a prepared batch, the library reports a batch failure rather than guessing from an HTTP status or exception that one document is invalid. The VDB boundary refuses a mixed partial write when searchable rows are missing embeddings.

For an unpinned custom remote model, the library does not guess its tokenizer or input limit. Embedding stage setup fails with an actionable error. Use a registered model, a local checkpoint, or an immutable model revision so the library can enforce deterministic client-side admission.

The embedding stage records per-row counts in embedding_v1_counts_by_label. When a batch contains an overlength or failed row, it also logs a summary with input_rows, output_rows, overlength, split, split_children, truncated, failed, embedded, and unembedded. The deterministic split policy reports truncated=0.

Example with Embedding Structured Elements as Text + Images

It is common to process PDFs by embedding standard text as text and embed visual elements such as tables and charts as images. The following example enables the multimodal model to capture the spatial and structural information of the visual content.

  • The embed method is configured with embed_modality="text_image" to embed the extracted tables and charts as images.
  • This configuration is more accurate than text only, with a performance cost.

For parameter details, refer to the Python API guide (create_ingestor and .embed()).

from nemo_retriever import create_ingestor

ingestor = (
    create_ingestor(run_mode="batch")
    .files("./data/*.pdf")
    .extract()
    .embed(
        embed_modality="text_image",
    )
)
results = ingestor.ingest()

Example with Embedding Entire PDF Pages as Images

For documents where the entire page layout is important (such as infographics, complex diagrams, or forms), you can configure NeMo Retriever Library to treat every page as a single image. The following example extracts and embeds each page as an image.

  • Set embed_modality="image" to use the rendered page image as the embedding input.
  • Set embed_granularity="page" to create one result row for each PDF page.

These arguments work together. When you set both arguments, the pipeline enables page-image rendering during extraction, creates one row for each page, and embeds the full rendered page image. Either argument alone does not enable the complete page-as-image workflow.

For parameter details, refer to the Python API guide (create_ingestor and .embed()).

from nemo_retriever import create_ingestor

ingestor = (
    create_ingestor(run_mode="batch")
    .files("./data/*.pdf")
    .extract()
    .embed(
        embed_modality="image",
        embed_granularity="page",
    )
)
results = ingestor.ingest()