Skip to main content
Vision models use the same POST /v1/openai/chat/completions route as chat models. Pick the output with method. Region models return a JSON string in choices[0].message.content. OCR models return text.
  • One route: the same OpenAI client, the same message shape, any model on this page.
  • Two request fields: method goes in extra_body. response_format is a standard OpenAI parameter, so pass it as a normal argument.
  • Images and video: pass an image_url or a video_url part. Video replies add per-frame records and tracks.
Live methods and default_method are on GET /v1/openai/models. A method a model does not list is a 400.

VQA, detection, and segmentation

Detect objects

Caption an image

Use caption, detailed_caption, or more_detailed_caption for a sentence or paragraph. With json_object, the string is in content. The samples below reuse client from the first example.
Python

Segment an object

Name the target in a text part next to the image. segment returns one instance per match, with a normalized box, a score, and the covered area. The pixels are one PNG label map on the container, not one mask per item.
Python

Pose estimation

An image with no person or hand returns "items": [].
Python
Each item carries these fields: The full item record is in Response Types.

Document OCR models

OCR models read a page image or a document_url PDF. Two families: For PDFs, treat the model, method, and document_dpi as one unit. See Document OCR.

Markdown models

Omit method to use the model’s default.
Python

PP-OCR models

detect returns one box and one polygon per text line. ocr and text add the recognized text.
Python

Read the result

Coordinates are normalized to the image. Set precision to change the decimal places.

Methods

Select the operation with method and method_params.

Response Formats

JSON mode for code, text mode for agents.

Models

Every vision model and its methods.

Supported Inputs

Image and video content parts, limits, and formats.

Document OCR

Page ranges, DPI, and method choice for PDFs.