model, and run them on our GPUs or yours. We orchestrate the models, keep the GPUs busy, and benchmark the VLMs, so you can ship the pipeline. Start with the quickstart.
Unified Interface for Visual Intelligence
usage.cost. See Pricing and Rate limits.
Why the Gateway
- Build on open-weight models: No vendor or frontier-model lock-in. Swap open-weight VLMs, ViT models, and other specialized vision models by changing
model, and compare output and cost side by side. See Models. - Our GPUs or yours: Use these models on our GPUs with no hardware to provision, at a published rate on the live model catalog. Or deploy the same models onto your own GPUs with our platform.
- Uncompromised quality: We are not another LLM serving provider. Each model is tuned and checked for visual quality. The aim is the best quality per dollar for vision.
- Orchestration built-in: Worker parallelism, retries, failure handling, GPU utilization, and response healing run on our side, not in your code. Multi-page PDFs are rasterized and fanned out per page, with no splitting or stitching.
- Agent ready: The MCP server exposes file-reading tools to any MCP-aware framework. System One answers typed questions in one read.
Next steps
Quickstart
Make your first VQA and document OCR requests.
Chat Completions
OpenAI-compatible chat for general visual intelligence: VQA, OCR, detection, segmentation, and pose.
System One
TypeSafe-compatible decision API for low-latency, calibrated answers over text, JSON, images, and PDFs.
MCP Server
File-reading tools for any MCP-aware agent, such as Pydantic AI, LangChain, Mastra, and Claude Code.
VLMs
Browse the open-weight VLMs and vision models you can call.
Pricing
Per-token rates and per-request
usage.cost.