list_models
Lists models on an MCP tool menu, and which read tool each one fits. Call it before passing a non-default model. A REST model on no tool menu is omitted, for example qwen/qwen3-vl-embedding-2b or a routed provider model. Query GET /v1/openai/models for the full catalog.
Each entry is
{id, modality, tool, tools, aliases, methods, default_method, method_details}.
id: Pass it asmodelon the matchingtool.tools: Every read-tool menu that id sits on. A Qwen vision model is onread_imageandread_video.modalityandtool: Name the primary menu. Amodalityfilter matches any menu the id sits on.- Menu membership: Not inferred from capabilities. Every OCR model accepts a bare image, but only
read_documenttakes an OCR model asmodel, so an OCR model listsread_documentalone.
method_details is one record per method, including its capabilities.
default: true: Marks the model’s default.requires_json_mode: true: The full payload arrives only whenjson_modeistrue. Amarkdownmethod can carry that flag and still return Markdown as a plain string whenjson_modeisfalse.- Chat VLM: Returns the model’s own reply and has no capabilities.
get_model_info
Returns one served model: its methods, and the JSON schema of each method’s json_mode reply. Call it before you write code against those fields.
The reply is
{id, aliases, tools, capabilities, default_method, methods}.
methods: One record per method:{name, capabilities, default, requires_json_mode, responses}.responses: Keyed by input medium (image,video, ordocument). Each entry is{payload, schema}.payload: The reply tag, such asdoc.ocr.lines, ortextfor a plain string.schema: The JSON schema of thejson_modereply. Every field has a one-line description.- Verbatim replies: A model whose replies pass through verbatim, such as a chat VLM, has
responses: {}. - Unserved model: A tool error that points at
list_models.
get_completion
Look up cost and serving details for a prior read_* call. Pass the completion_id from that call’s ReadResult.
The reply carries
id, model, served_model_id, backend, cost, usage (tokens, or transcription seconds), status, and latency_ms.
statusandlatency_ms: Come from the billing record. A lookup answered from the short-lived cache leaves them null.- Rate limit: The lookup spends one model call from your rate-limit budget, same as
GET /v1/completions/{completion_id}.
An unknown id, an expired id, or another caller’s id returns
404 (completion_not_found). A retry mints a new id. It does not resolve the missed one.
Return shape (ReadResult)
Each read_* tool returns text as the tool content and a typed payload.
Output past roughly 60,000 characters is cut, and a
[truncated: …] note is appended. Narrow pages, or read the document in sections.