Skip to main content
The VLM Run API is a unified platform for production-ready multimodal AI. Use it to extract structured data from documents, images, videos, and audio, or run complex multi-step workflows with visual agents.
  • Base URL: https://api.vlm.run/v1
  • Authentication: Authorization: Bearer <VLMRUN_API_KEY>
  • Agents Supported:
    • vlmrun-orion-2: Our second-generation visual agent. It adds code execution, and is faster, cheaper, and more capable, with high reliability and repeatability.
    • vlmrun-orion-1: Our first visual agent (November 2025). It sees, reasons, and acts with detection, segmentation, and specialized document OCR VLMs.
  • Variants: fast, auto, and pro, for example vlmrun-orion-2:auto.
Access your API keys in our dashboard.

Need OCR, VQA, or detection models?

Call them through the Gateway: one OpenAI-compatible API and one key for every visual model. This reference covers the Orion agents.

Chat Completions & Agent Executions

Use the Chat Completions endpoint for interactive multi-modal conversations, or the Agent Executions endpoint for batch execution workflows.

Next steps

Orion Introduction

What the agents are, and which one to pick.

Chat Completions

Send a multi-modal conversation to an agent.

Execute Agent

Run a saved agent on a file and poll for the result.

Upload File

Upload media once and reuse it by file ID.

Skills

Reusable, domain-specific skills for your agents.

Gateway

OCR, VQA, detection, and segmentation models behind one API.

Legacy API

Looking for the vlm-1 API? Click here.