POST /v1/openai/chat/completions route as chat models. Pick the output with method. Region models return a JSON string in choices[0].message.content. OCR models return text.
- One route: the same OpenAI client, the same message shape, any model on this page.
- Two request fields:
methodgoes inextra_body.response_formatis a standard OpenAI parameter, so pass it as a normal argument. - Images and video: pass an
image_urlor avideo_urlpart. Video replies add per-frame records and tracks.
methods and default_method are on GET /v1/openai/models. A method a model does not list is a 400.
VQA, detection, and segmentation
Detect objects
Caption an image
Usecaption, detailed_caption, or more_detailed_caption for a sentence or paragraph. With json_object, the string is in content. The samples below reuse client from the first example.
Python
Segment an object
Name the target in a text part next to the image.segment returns one instance per match, with a normalized box, a score, and the covered area. The pixels are one PNG label map on the container, not one mask per item.
Python
Pose estimation
An image with no person or hand returns
"items": [].
Python
The full item record is in Response Types.
Document OCR models
OCR models read a page image or adocument_url PDF. Two families:
For PDFs, treat the model,
method, and document_dpi as one unit. See Document OCR.
Markdown models
Omitmethod to use the model’s default.
Python
PP-OCR models
detect returns one box and one polygon per text line. ocr and text add the recognized text.
Python
Read the result
Coordinates are normalized to the image. Set
precision to change the decimal places.
Related
Methods
Select the operation with
method and method_params.Response Formats
JSON mode for code, text mode for agents.
Models
Every vision model and its methods.
Supported Inputs
Image and video content parts, limits, and formats.
Document OCR
Page ranges, DPI, and method choice for PDFs.