Skip to main content

list_models

Lists models on an MCP tool menu, and which read tool each one fits. Call it before passing a non-default model. A REST model on no tool menu is omitted, for example qwen/qwen3-vl-embedding-2b or a routed provider model. Query GET /v1/openai/models for the full catalog. Each entry is {id, modality, tool, tools, aliases, methods, default_method, method_details}.
  • id: Pass it as model on the matching tool.
  • tools: Every read-tool menu that id sits on. A Qwen vision model is on read_image and read_video.
  • modality and tool: Name the primary menu. A modality filter matches any menu the id sits on.
  • Menu membership: Not inferred from capabilities. Every OCR model accepts a bare image, but only read_document takes an OCR model as model, so an OCR model lists read_document alone.
method_details is one record per method, including its capabilities.
  • default: true: Marks the model’s default.
  • requires_json_mode: true: The full payload arrives only when json_mode is true. A markdown method can carry that flag and still return Markdown as a plain string when json_mode is false.
  • Chat VLM: Returns the model’s own reply and has no capabilities.

get_model_info

Returns one served model: its methods, and the JSON schema of each method’s json_mode reply. Call it before you write code against those fields. The reply is {id, aliases, tools, capabilities, default_method, methods}.
  • methods: One record per method: {name, capabilities, default, requires_json_mode, responses}.
  • responses: Keyed by input medium (image, video, or document). Each entry is {payload, schema}.
  • payload: The reply tag, such as doc.ocr.lines, or text for a plain string.
  • schema: The JSON schema of the json_mode reply. Every field has a one-line description.
  • Verbatim replies: A model whose replies pass through verbatim, such as a chat VLM, has responses: {}.
  • Unserved model: A tool error that points at list_models.

get_completion

Look up cost and serving details for a prior read_* call. Pass the completion_id from that call’s ReadResult. The reply carries id, model, served_model_id, backend, cost, usage (tokens, or transcription seconds), status, and latency_ms.
  • status and latency_ms: Come from the billing record. A lookup answered from the short-lived cache leaves them null.
  • Rate limit: The lookup spends one model call from your rate-limit budget, same as GET /v1/completions/{completion_id}.
An unknown id, an expired id, or another caller’s id returns 404 (completion_not_found). A retry mints a new id. It does not resolve the missed one.

Return shape (ReadResult)

Each read_* tool returns text as the tool content and a typed payload. Output past roughly 60,000 characters is cut, and a [truncated: …] note is appended. Narrow pages, or read the document in sections.

Errors

A method the model does not advertise is refused before the call. The message lists the methods it does advertise.