(model, method), the shape is fixed before the call.
- Type-safe:
contentis one of the types below, tagged bycontent.object. Validate it with the Pydantic or Zod model shown for each type. - Deterministic: the same
(model, method)always returns the same type with the same keys. A reply never adds a key outside the type, and a payload is never a bare array. - Portable: two models that run the same task return the same type, so a model swap does not touch your parser.
content in JSON mode and a json block in text mode. Tags read <domain>.<task>.<unit>: what it describes (doc, img, vid, world), the operation, and what one items entry is. A markdown-kind method returns a plain string with no tag.
Object types
doc.ocr.lines: a recognized text line with its box.doc.ocr.bboxes: a text region, geometry only.doc.page.blocks: a layout block with its text. Every document page in JSON mode.doc.layout.bboxes: a layout region, no text.img.detect.bboxes: a detected object.img.segment.masks: a segmented instance, pixels on a label map.vid.segment.masks: a tracked instance on one frame.img.pose.kpts: a person with 2D keypoints.vid.pose.kpts: a person on one frame.world.pose.kpts: a hand with 2D and 3D keypoints.
OCR
doc.ocr.lines
One item is a recognized text line with its box and score. Produced by pp-ocrv6 ocr, deepseek-ocr-2 grounding_ocr, and florence-2 ocr_with_region.
doc.ocr.bboxes
One item is a text region with geometry and no text. Produced by pp-ocrv6 detect.
Document layout
doc.page.blocks
One item is a layout block with its text. On a document page the container tag is always doc.page.blocks, whatever method read the page. A whole-page read is one block with text and no geometry. Produced by dots.mocr parse_layout, and every document page in JSON mode.
Document block record
Every key is optional exceptblock_id, and one type covers every document method.
doc.layout.bboxes
One item is a layout region with no text. Produced by dots.mocr parse_layout_only.
Detection
img.detect.bboxes
One item is a detected object. Produced by florence-2 od, dense_region_caption, and region_proposal.
Segmentation
img.segment.masks
One item is a segmented instance. The pixels are one PNG label map on the container, not one mask per item. Produced by sam3.1 segment and segment_box.
Label map
One 8-bit PNG holds every instance, and each pixel is aninstance_id (image) or track_id (video). Pixel 0 is background, and ids run 1 to 255. Decode the PNG once, then filter items on area, bbox_xywh, or score. Set mask_format: "none" to keep area only.
vid.segment.masks
One item is a tracked instance on one frame. Produced by sam3.1 track. Each frame has its own label map, where the pixel value is track_id.
Keypoints
img.pose.kpts
One item is a person with 2D joints. Produced by vitpose-plus-large pose on an image. An image with no person returns "items": [].
vid.pose.kpts
One item is a person on one frame. Produced by vitpose-plus-large pose on a video.
world.pose.kpts
One item is a hand with 2D and 3D joints. Produced by hamer pose.
Parse a reply
Validate the JSON-mode reply with the type for your(model, method). This example parses a florence-2 od reply.
object tag, a missing key, or a wrong type. A video reply has video_* keys and a PDF reply has pages. See JSON mode.
Item record
The keys an item can carry, across all types.items is a flat list: one row per thing, never nested per frame or per instance.
Each model card names the tag every method returns under Output by method.
Related
Response Formats
Text, JSON, and JSON schema modes.
Methods
Select what the model computes with
method.Visual Intelligence
Request samples for detection, segmentation, and pose.
Models
Every model and its methods.