- Your code needs a contract. A pipeline or service must parse the reply without guessing. It needs the same keys, a typed payload, and a tag that names the shape. Prose breaks your parser the first time the wording changes.
- Your LLM needs readable text. An agent or an MCP tool call reads the reply as context. Keys and braces cost tokens, and a stream that arrives early helps.
response_format selects the format. All three modes carry the same data: a text reply is the JSON payload, rendered.
For programmatic use, pick JSON mode or JSON schema mode. For agents, leave
response_format unset.
Text Mode
Omitresponse_format. {"type": "text"} is equivalent. Over MCP, leave json_mode at its default false.
Python
<document> block. An image is the bare payload. A block body is one of two kinds.
The kind is fixed by the
(model, method) pair and is the same on every page of a request. Each model card lists the kind every method returns.
- Streaming:
stream: trueis honored. The stream is byte-identical to the non-streaming reply. See streaming. - Chat VLMs: the reply is the model’s own text, with no wrapper. A PDF sent to a chat VLM returns a
400.
Document input
An input PDF is one<document> block wrapping one <page> block per rasterized page.
Each attribute is a field of the same JSON record with the prefix dropped.
<page id>ispage_id.<document npages>isdocument_npages.- A failed page is self-closing, with
status="error"and no body. Page numbering stays intact.
Single image input
The block alone, with no wrapper:image_hash, image_width, or image_height. Use JSON mode when that metadata is needed.
JSON Mode
Setresponse_format={"type": "json_object"}.
Python
model, method, the input, then the payload. There is no data list. The payload types are in Response Types.
A JSON
response_format is buffered. A single valid JSON object cannot be assembled from SSE deltas. With stream: true the finished object is sent as chat.completion.chunk frames, so there is no time-to-first-token benefit.- Image reply
- Video reply
- Document reply
object.content per medium
A json payload is never a bare array. A markdown payload is never wrapped. An image
content has exactly one shape per (model, method).
Chat VLMs
A chat VLM reply is passed through in JSON mode too. The body is the model’s own JSON, with nodata wrapper and no object tag.
response_format={"type":"json_object"} means. The model emits the JSON.
JSON Schema Mode
JSON schema mode constrains generation to a schema you supply. The reply is the model’s own JSON, with the keys you defined and nomodel or method wrapper.
- Supported today: chat VLMs (
qwen/qwen3.5-0.8b,qwen/qwen3.8-27b,google/gemma-4-26b-a4b-it) and provider models (google/gemini-*,meta/muse-spark-1.2,moonshotai/kimi-k3,minimax/minimax-m3).meta/muse-glimmer-30band the diffusion engine are not supported. More models may add support. Checkcapabilities.supports_json_schemaonGET /v1/openai/models. - Unsupported models: OCR, detection, segmentation, and pose models return
400withcapability_violation. Use JSON mode there. - Streaming: the reply is buffered, as in JSON mode.
Python
Rounding
Setprecision to change the decimal places on normalized coordinates and score in text mode and JSON mode.
Related
Response Types
Every payload type: OCR, detection, segmentation, keypoints.
Methods
Select what the model computes, and pass method parameters.
MCP Tools
The
json_mode flag on the read tools.Document OCR
End-to-end recipe from model selection to response parsing.
Chat Completions
Full request and response schema.