Skip to main content
A read tool is mounted only when this gateway serves a model that fits it. tools/list is that set. read_document and read_image both accept a JPEG or a PNG. Use read_document when the value is printed text. It OCRs into Markdown and keeps headings, lists, and tables. Use read_image for what a picture depicts: a photo, a chart, or a screenshot. url is an http(s) URL, a data: URI, or a bare base64 string. It is never a local path. Host the file, or inline it, before you call. json_mode: false returns the native text: <document> blocks for documents, and plain text for images, audio, and video. json_mode: true returns the same flat object the REST API returns in JSON mode. Nothing on this server writes. Clients can run these tools without a confirmation prompt.
model is a fixed menu per tool, built at mount from the models this gateway serves.
  • read_document: OCR models that accept a document_url.
  • read_image and read_video: The chat VLMs plus the region models named below.
  • read_audio: The transcription models.
  • No menu: Routed provider models and embedding models.
  • Default: The preferred default leads its menu when that model is served.
Call list_models for the ids a tool accepts. A model id passed to the wrong tool is rejected.

Pick a method by what you need

Each read_document method returns a different result from the same page. list_models tags each method with a capability. Pick the capability, then call the method that carries it. rednote-hilab/dots.mocr carries document_markdown on markdown and text_citations on parse_layout. paddleocr/pp-ocrv6 carries text_citations on ocr and text_extraction on text.

read_document

OCR a PDF or image into Markdown, or into structured JSON. Use it for invoices, forms, reports, slide decks, screenshots, and scanned pages.

Page selection

  • [0, 2, 4] reads pages 0, 2, and 4
  • [[0, 3]] reads pages 0-2
  • [[0, 3], [5, 7], -1] reads pages 0, 1, 2, 5, 6, and the last page
One call reads at most 128 pages. For a longer document, call once per range and combine the results. The error names the ranges to use.

JSON mode

Left unset, the tool returns the same text blocks as the REST API. One <document> block per PDF wraps one <page> block per page, in reading order.
A page that failed OCR is self-closing, with status="error" and no body. A single image returns the Markdown alone, with no wrapper.
Leave json_mode at false for agent use. Set it to true only when code parses the reply. Both shapes are documented under Text mode and JSON mode.

read_image

Send one still image and get the model’s answer. Use it for photos, charts, diagrams, product shots, and UI screenshots. A region model returns records instead of prose: masks from facebook/sam3.1, keypoints from usyd-community/vitpose-plus-large, or boxes and captions from microsoft/florence-2-base-ft. Those records need json_mode: true. The read_image menu has no OCR model. read_document accepts an image URL, and a single image comes back as Markdown with no wrapper.

read_audio

Transcribe speech audio (wav, mp3, m4a, flac, ogg, and the rest) into text.

read_video

Describe or transcribe a video. The clip goes to a model that ingests video natively. There is no gateway-side video-to-image step, and a model without video support is refused. A region model returns one record per sampled frame: tracks from facebook/sam3.1 (track), or keypoints from usyd-community/vitpose-plus-large (pose). Those records need json_mode: true. On the REST API, encoder and encoder_params map to video_encoder and video_encoder_params.