Input JSON Format#

This guide describes the input JSON format for the LLM inference tool. The format supports text-only and multimodal inputs, multi-turn conversations, LoRA adapters, and advanced features.

Basic Structure#

{
    "batch_size": 1,
    "temperature": 1.0,
    "top_p": 0.8,
    "top_k": 50,
    "sampling_seed": 123456789,
    "spec_proposal_sampling": "auto",
    "logit_bias": {"123": -100.0},
    "max_generate_length": 256,
    "num_logprobs": 0,
    "apply_chat_template": true,
    "enable_thinking": false,
    "context_cache_lookup_policy": "use_cache",
    "context_cache_commit_policy": "including_generated_tokens",
    "available_lora_weights": {},
    "requests": [
        {
            "messages": [
                {
                    "role": "user",
                    "content": "Your message here"
                }
            ],
            "lora_name": "optional_lora_name",
            "disable_spec_decode": false,
            "logit_bias": {"456": 5.0},
            "num_logprobs": 0,
            "stop": ["optional_stop_string"]
        }
    ]
}

Parameters#

Required#

  • requests: Array of conversation requests

Optional Global Parameters#

  • batch_size (default: 1): Number of requests per batch

  • temperature (default: 1.0): Sampling temperature (0.0 = deterministic)

  • top_p (default: 0.8): Nucleus sampling threshold

  • top_k (default: 50): Top-k sampling parameter

  • sampling_seed (optional): Stable unsigned 64-bit sampling seed. A request-level value overrides this global default. Speculative proposal, acceptance, residual, and bonus randomness is derived from this seed and absolute output position, so continuous-batching slot changes do not alter those random streams.

  • spec_proposal_sampling (default: "auto"): Proposal policy for speculative decoders that expose their proposal distribution. "auto" follows the target policy; "greedy" or "probabilistic" forces the draft policy while preserving correct target sampling through rejection verification.

  • logit_bias (optional): Sparse map from token ID to bias value. The top-level map is the default for all requests.

  • max_generate_length (default: 256): Maximum tokens to generate

  • apply_chat_template (default: true): Apply chat template formatting

  • add_generation_prompt (default: true): Add generation prompt token sequence

  • enable_thinking (default: false): Enable thinking mode (Qwen3+)

  • context_cache_lookup_policy (default: "use_cache"): Use "bypass" to skip lookup and publication for this invocation when the runtime was started with --enableContextReuse.

  • context_cache_commit_policy (default: "including_generated_tokens"): Select "prefill_state_only" to publish only the input prefix. See KV Cache Reuse.

  • num_logprobs (default: 0): Number of top log-probabilities to return per generated token. 0 disables logprobs. Maximum value is 50; values outside [0, 50] are rejected at input parsing. The cap applies to the per-token K of each individual request — it is not a cumulative budget across requests or batches: a batch simply computes at the maximum K its requests asked for, and every request in that batch receives that K. Each candidate carries token_id, token (decoded string), bytes (raw token bytes) and logprob, following OpenAI API semantics (log(softmax(logits))). The top-level value is the default for all requests; a request may raise it. Supported for both vanilla and EAGLE speculative decoding.

  • available_lora_weights (default: {}): Map of LoRA adapter names to file paths

Request Fields#

  • messages (required): Array of conversation messages

  • sampling_seed (optional): Overrides the global stable sampling seed for this request.

  • lora_name (optional): LoRA adapter name from available_lora_weights

  • save_system_prompt_kv_cache (optional): Legacy compatibility field for exact system-prompt caching. New deployments should use KV cache reuse. Bounded Gemma4 SWA storage does not preserve the complete contiguous history required by this legacy snapshot API, so the request continues without saving a snapshot.

  • disable_spec_decode (optional, default: false): Disable EAGLE speculative decoding for this request even if draft engine is loaded

  • logit_bias (optional): Request-specific sparse logit-bias map. When set, it overrides the top-level logit_bias default for this request.

  • num_logprobs (optional): Overrides the top-level num_logprobs default for this request. Applied batch-uniformly (like disable_spec_decode): the batch computes at the maximum value requested by any request in it.

  • guided_decoding (optional): Constrain the output to a schema, pattern, or grammar. Exactly one mode may be set. Like logit_bias, a top-level guided_decoding applies to every request unless a request overrides it. See Guided Decoding.

  • stop (optional): String or array of strings that halt generation when produced in the output. The stop string itself is excluded from the returned text. Each request in a batch may declare its own list independently. Defaults to no stop strings.

Message Fields#

  • role: "system", "user", or "assistant"

  • content: String (text-only) or Array (multimodal)

Content Array Format:

  • Text: {"type": "text", "text": "..."}

  • Image: {"type": "image", "image": "/path/to/image.jpg"}. Optional "do_resize" (default true): set to false when the image is already resized to the model’s target size — the vision runner then consumes it as-is instead of resizing internally (see Pre-resized image input).

  • Audio: {"type": "audio", "audio": "/path/to/clip.wav"} (raw .wav / .mp3 / .flac decoded in C++ via vendored miniaudio + in-tree mel extractor. Feature-extractor family — whisper / parakeet — is auto-derived from the engine’s audio/config.json, mirroring HF / vLLM where FE is pinned by the model. The HTTP server in experimental.server accepts the same audio formats via input_audio / audio_url and routes through the same C++ mel path.)

  • The C++ llm_inference CLI does not accept video content. The OpenAI-compatible server accepts video, video_url, or an explicit frame list for supported Qwen and InternVL model families and Nemotron Omni. Nemotron Omni accepts exactly one video and no images per request. The Cosmos3 policy runtime has a separate observation/action contract; see the Cosmos3 VLA guide.

Examples#

Basic Text Input#

{
    "batch_size": 1,
    "max_generate_length": 256,
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "What is machine learning?"}
            ]
        }
    ]
}

Multi-Turn Conversation#

{
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "What is the capital of France?"},
                {"role": "assistant", "content": "The capital of France is Paris."},
                {"role": "user", "content": "What is the population?"}
            ]
        }
    ]
}

Multimodal Input (Vision-Language Models)#

{
    "requests": [
        {
            "messages": [
                {
                    "role": "user",
                    "content": [
                        {"type": "image", "image": "/path/to/image.jpg"},
                        {"type": "text", "text": "Describe this image."}
                    ]
                }
            ]
        }
    ]
}

Pre-resized Image Input (do_resize: false)#

By default the vision runner resizes every input image to the model’s target size. If your pipeline already produces frames at the target size (e.g. a camera/robotics pipeline that resizes once upstream), set "do_resize": false on the content item to skip the runtime’s internal resize:

{"type": "image", "image": "/path/to/pre_resized.png", "do_resize": false}

The same field is available on the video content item, on the Python ImageData binding (image.do_resize = False), and on the C++ struct (rt::imageUtils::ImageData::doResize).

Contract for pre-resized inputs:

  • Supply raw uint8 RGB pixels. Do not rescale or normalize the pixel values yourself — mean/std normalization always runs inside the runtime.

  • Dimensions must exactly match the model’s resize target. The per-model target formulas are exposed as stateless C++ functions in cpp/multimodal/common/imageUtils.h. For example, for the Qwen family:

    auto [targetHeight, targetWidth] = rt::imageUtils::qwenSmartResize(
        origHeight, origWidth, patchSize, mergeSize, minImageTokensPerImage, maxImageTokensPerImage);
    

LoRA Adapters#

{
    "available_lora_weights": {
        "french": "/path/to/french_adapter.safetensors"
    },
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "Translate to French."}
            ],
            "lora_name": "french"
        }
    ]
}

Note: All requests in the same batch must use the same LoRA adapter.

Raw Format (No Chat Template)#

{
    "apply_chat_template": false,
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "Raw text without special tokens"}
            ]
        }
    ]
}

Disable Speculative Decoding#

When using EAGLE speculative decoding, you can disable it for specific requests:

{
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "Your question here"}
            ],
            "disable_spec_decode": true
        }
    ]
}

Use cases:

  • Quality: Some inputs may benefit from standard decoding over EAGLE

  • Switching strategies: Different batches can use different decoding strategies (one batch with EAGLE, another without)

  • Debugging: Compare performance with/without speculative decoding

Note: If any request in a batch has disable_spec_decode: true, speculative decoding will be disabled for the entire batch. Requests within one batch cannot use different decoding strategies simultaneously for now.

Logit Bias#

logit_bias accepts a sparse map of tokenizer token IDs to additive logit bias values. Positive values make a token more likely, and negative values make it less likely. Bias values must be finite and in [-100.0, 100.0]; each map may contain up to 1024 token IDs.

Top-level logit_bias applies to every request by default. A request-level logit_bias overrides the top-level default for that request.

{
    "logit_bias": {"123": -100.0},
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "Avoid token 123 by default."}
            ]
        },
        {
            "messages": [
                {"role": "user", "content": "Prefer token 456 for this request."}
            ],
            "logit_bias": {"456": 5.0}
        }
    ]
}

logit_bias remains active during speculative decoding. The runtime applies the request’s map to every verification row and to fallback sampling for EAGLE, MTP, DFlash, JetSpec, and DSpark. disable_spec_decode: true may still be used to force vanilla decoding for comparison, but it is not required for logit bias.

Token IDs use the full tokenizer vocabulary. With a reduced-vocabulary engine, a token omitted from the engine vocabulary cannot be restored by positive bias; the runtime ignores that entry and logs a warning.

Guided Decoding#

guided_decoding constrains generation token by token so the output is guaranteed to match the guide. Set exactly one of json_object, json_schema, regex, ebnf, structural_tag, or choice.

{
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "Give me a person record."}
            ],
            "guided_decoding": {
                "json_schema": {
                    "type": "object",
                    "properties": {
                        "name": {"type": "string"},
                        "age": {"type": "integer"}
                    },
                    "required": ["name", "age"]
                }
            }
        },
        {
            "messages": [
                {"role": "user", "content": "Answer yes or no."}
            ],
            "guided_decoding": {"choice": ["yes", "no"]}
        }
    ]
}

json_schema and structural_tag accept an inline object, choice an inline array, and every mode also accepts an escaped JSON string.

Speculative decoding limitation: guided requests are rejected while speculative decoding is active. Set disable_spec_decode: true to use vanilla decoding for that batch.

See Guided Decoding for all six modes, schema support, and build instructions.

Top-N Log-Probabilities#

Return the top-5 most likely tokens (with log-probabilities) at each generation step:

{
    "num_logprobs": 5,
    "max_generate_length": 64,
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "The capital of France is"}
            ]
        }
    ]
}

num_logprobs may also be set per request to override the top-level default (see Request Fields).

The output JSON will contain a logprobs field alongside output_text:

"logprobs": [
    [
        {"token_id": 3681, "token": " Paris", "bytes": [32, 80, 97, 114, 105, 115], "logprob": -0.041},
        {"token_id": 8098, "token": " paris", "bytes": [32, 112, 97, 114, 105, 115], "logprob": -3.812},
        {"token_id":  627, "token": ".",       "bytes": [46],                        "logprob": -4.201},
        {"token_id": 4892, "token": " PARIS", "bytes": [32, 80, 65, 82, 73, 83],     "logprob": -5.103},
        {"token_id": 3085, "token": " France","bytes": [32, 70, 114, 97, 110, 99, 101], "logprob": -5.447}
    ],
    ...
]

Each element of the outer array corresponds to one generated token (step); the inner array lists up to num_logprobs candidates sorted by descending probability. Each candidate carries token_id, token, bytes and logprob.

token is the token decoded as UTF-8 with invalid/partial bytes replaced by U+FFFD; bytes is the raw token bytes. For byte-level BPE tokenizers a single token may be only part of a multi-byte character (e.g. half a CJK character), so token can be lossy — concatenate bytes across tokens to reconstruct the exact text losslessly.

Notes:

  • num_logprobs can be set top-level (default for all requests) and/or per request (override); it is batch-uniform — the batch computes at the maximum value requested, and all requests in the batch share the same K.

  • Logprobs are computed at temperature=1.0 regardless of the temperature sampling setting.

Stop Strings#

Generation halts as soon as any of the specified substrings appears in the decoded output; the stop string itself is excluded from the returned text. Accepts a single string or an array. Each request carries its own independent list — requests in the same batch may stop on different strings or none at all.

{
    "requests": [
        {
            "messages": [
                {"role": "user", "content": "Write a short answer ending before '###'."}
            ],
            "stop": ["###", "\n\nUser:"]
        }
    ]
}

When a stop string triggers termination, the request’s finish reason is stop-words. Earliest position in the decoded output wins when multiple stop strings could match.

Notes#

  • System prompt: Uses provided system message, or model default from chat template

  • LoRA: All requests in same batch must use same adapter

  • Paths: Use absolute or relative paths for images/videos

  • Format: Follows OpenAI chat completion API structure