Skip to main content
JSON mode is built for code that must not guess. The reply is type-safe and deterministic: for a given (model, method), the shape is fixed before the call.
  • Type-safe: content is one of the types below, tagged by content.object. Validate it with the Pydantic or Zod model shown for each type.
  • Deterministic: the same (model, method) always returns the same type with the same keys. A reply never adds a key outside the type, and a payload is never a bare array.
  • Portable: two models that run the same task return the same type, so a model swap does not touch your parser.
The payload is content in JSON mode and a json block in text mode. Tags read <domain>.<task>.<unit>: what it describes (doc, img, vid, world), the operation, and what one items entry is. A markdown-kind method returns a plain string with no tag.

Object types

Each type below has a JSON sample, a Pydantic model, and a Zod schema.

OCR

doc.ocr.lines

One item is a recognized text line with its box and score. Produced by pp-ocrv6 ocr, deepseek-ocr-2 grounding_ocr, and florence-2 ocr_with_region.

doc.ocr.bboxes

One item is a text region with geometry and no text. Produced by pp-ocrv6 detect.

Document layout

doc.page.blocks

One item is a layout block with its text. On a document page the container tag is always doc.page.blocks, whatever method read the page. A whole-page read is one block with text and no geometry. Produced by dots.mocr parse_layout, and every document page in JSON mode.

Document block record

Every key is optional except block_id, and one type covers every document method.

doc.layout.bboxes

One item is a layout region with no text. Produced by dots.mocr parse_layout_only.

Detection

img.detect.bboxes

One item is a detected object. Produced by florence-2 od, dense_region_caption, and region_proposal.

Segmentation

img.segment.masks

One item is a segmented instance. The pixels are one PNG label map on the container, not one mask per item. Produced by sam3.1 segment and segment_box.

Label map

One 8-bit PNG holds every instance, and each pixel is an instance_id (image) or track_id (video). Pixel 0 is background, and ids run 1 to 255. Decode the PNG once, then filter items on area, bbox_xywh, or score. Set mask_format: "none" to keep area only.

vid.segment.masks

One item is a tracked instance on one frame. Produced by sam3.1 track. Each frame has its own label map, where the pixel value is track_id.

Keypoints

img.pose.kpts

One item is a person with 2D joints. Produced by vitpose-plus-large pose on an image. An image with no person returns "items": [].

vid.pose.kpts

One item is a person on one frame. Produced by vitpose-plus-large pose on a video.

world.pose.kpts

One item is a hand with 2D and 3D joints. Produced by hamer pose.

Parse a reply

Validate the JSON-mode reply with the type for your (model, method). This example parses a florence-2 od reply.
A mismatch raises at the boundary: a different object tag, a missing key, or a wrong type. A video reply has video_* keys and a PDF reply has pages. See JSON mode.

Item record

The keys an item can carry, across all types. items is a flat list: one row per thing, never nested per frame or per instance. Each model card names the tag every method returns under Output by method.

Response Formats

Text, JSON, and JSON schema modes.

Methods

Select what the model computes with method.

Visual Intelligence

Request samples for detection, segmentation, and pose.

Models

Every model and its methods.