tools/list is that set.
read_document and read_image both accept a JPEG or a PNG. Use read_document when the value is printed text. It OCRs into Markdown and keeps headings, lists, and tables. Use read_image for what a picture depicts: a photo, a chart, or a screenshot.
url is an http(s) URL, a data: URI, or a bare base64 string. It is never a local path. Host the file, or inline it, before you call.
json_mode: false returns the native text: <document> blocks for documents, and plain text for images, audio, and video. json_mode: true returns the same flat object the REST API returns in JSON mode.
Nothing on this server writes. Clients can run these tools without a confirmation prompt.
Pick a method by what you need
Eachread_document method returns a different result from the same page. list_models tags each method with a capability. Pick the capability, then call the method that carries it.
rednote-hilab/dots.mocr carries document_markdown on markdown and text_citations on parse_layout. paddleocr/pp-ocrv6 carries text_citations on ocr and text_extraction on text.
read_document
OCR a PDF or image into Markdown, or into structured JSON. Use it for invoices, forms, reports, slide decks, screenshots, and scanned pages.
Page selection
[0, 2, 4]reads pages 0, 2, and 4[[0, 3]]reads pages 0-2[[0, 3], [5, 7], -1]reads pages 0, 1, 2, 5, 6, and the last page
One call reads at most 128 pages. For a longer document, call once per range and combine the results. The error names the ranges to use.
JSON mode
- json_mode: false
- json_mode: true
Left unset, the tool returns the same text blocks as the REST API. One A page that failed OCR is self-closing, with
<document> block per PDF wraps one <page> block per page, in reading order.status="error" and no body. A single image returns the Markdown alone, with no wrapper.json_mode at false for agent use. Set it to true only when code parses the reply. Both shapes are documented under Text mode and JSON mode.
read_image
Send one still image and get the model’s answer. Use it for photos, charts, diagrams, product shots, and UI screenshots. A region model returns records instead of prose: masks from facebook/sam3.1, keypoints from usyd-community/vitpose-plus-large, or boxes and captions from microsoft/florence-2-base-ft. Those records need json_mode: true.
The
read_image menu has no OCR model. read_document accepts an image URL, and a single image comes back as Markdown with no wrapper.
read_audio
Transcribe speech audio (wav, mp3, m4a, flac, ogg, and the rest) into text.
read_video
Describe or transcribe a video. The clip goes to a model that ingests video natively. There is no gateway-side video-to-image step, and a model without video support is refused. A region model returns one record per sampled frame: tracks from facebook/sam3.1 (track), or keypoints from usyd-community/vitpose-plus-large (pose). Those records need json_mode: true.
On the REST API,
encoder and encoder_params map to video_encoder and video_encoder_params.