Skip to main content
Orion agents are priced in US dollars. Every response includes the run cost.

What makes up your cost

  • Fixed-price tools: GPU and API actions charged at a flat rate per unit.
    • Examples: image generation and editing, OCR and layout, and video generation and editing.
  • Token-billed model usage: Orion reasoning and LLM-backed tools charged by input and output tokens.
    • Examples: image captioning, document extraction, and video analysis.

How a run adds up

The service tier sets the multiplier.

The per-span cost ledger

Every Orion 2 run returns a per-span cost ledger. Each line is one billable action. Each action is a fixed-price tool or token-billed usage. The multiplier applies once, at the run level (effective_cost = SUM(span.cost) × multiplier).

An example ledger

Illustrative ballpark for one document on Orion 2 Auto in program mode. Actual token usage and cost vary with the document and the requested output.

Choose your tiers

Model quality and the service tier are independent. Quality selects the models. The service tier scales the bill.

Model quality

Use Fast, Auto, or Pro as a curated default. Token rates vary by model.
VLM Run hosts additional Orion-2 backbones, open-weight and frontier models, beyond these three tiers. Pass any supported model ID on chat completions or agent execute to pin a backbone.

Service tiers

The service tier uses the same models, tools, and output. It changes how quickly the result comes back.
TierMultiplierCost EffectBest For
standard (default)1.0×Base USD priceMost production workloads.
flex0.5×50% of standardBackground jobs with no one waiting: overnight batches, backfills.
priority1.8×180% of standardInteractive or blocking workflows where latency is user-visible.
A $1.00 standard run costs $0.50 at flex and $1.80 at priority. Set service_tier per request:

What a run actually costs

cost_dollars in the response is the exact figure.

Gateway

Gateway pricing

How a Gateway request is billed. Live model rates are on the catalog.