Skip to main content
Orion runs real vision tools, not a chat model’s guess. Pick the toolset by the media you have. Orion-1 calls one tool per model round-trip. Orion-2 imports the same tools in code, for example vlmrun.image.detect(image, "person"). See Code Execution.

Image

Caption & Tag

Captions and tags for images.

Detection

Bounding boxes and confidence scores for objects, faces, and people.

Segmentation

Pixel-level masks for objects, regions, and features.

Pointing

Key points and structural features, with sub-pixel accuracy.

Generate & Edit

From a text prompt, a sketch, or an existing image.

UI Parsing

Buttons and interactive components in screenshots.

Image Tools

Crop, rotate, upscale, and colorize.

Documents

Layout Detection

Layout elements with bounding boxes and reading order.

Visual Grounding

Map text in a document to its location on the page.

Multi-Page Analysis

Correlate pages and documents in one request.

Video

Caption & Tag

Topic, summary, timestamped chapters, and content tags.

Generate & Edit

Generate a video from a text prompt.

Video Tools

Trim, sample, and extract video segments.