Skip to main content
The agent component provides methods for interacting with VLM Run’s Orion Agents for multi-modal chat completions.

Chat Completions

Generate responses from the agent using the chat completions:

Image Analysis

Analyze images using the agent:

Video Analysis

Analyze videos using the agent:

Structured Outputs

Get structured JSON responses using TypeScript interfaces:

Document Analysis

Analyze documents and PDFs:

Agent Execution

Beyond chat completions, client.agent.execute() runs a named agent (optionally with skills) over structured inputs and returns an execution you can poll:

Execution Mode

The response carries execution_mode, which reports whether the server replayed a skill’s cached pipeline.py or ran the full LLM agent loop. See program execution.
The Node.js SDK cannot select the mode yet. AgentExecutionConfig serializes only prompt, jsonSchema, skills, serviceTier, and orchestrationMode, so a mode key is dropped before the request is sent. Use the Python SDK or a raw cURL request to force a mode.

SDK Reference

client.agent.execute()

Execute an agent over structured inputs. Parameters: Returns: Promise<AgentExecutionResponse> with execution_mode set to the path taken.

client.agent.completions.create()

Create a chat completion with the agent. Parameters: Returns: Promise<ChatCompletionResponse>

Message Content Types

Best Practices

  1. Structured Outputs
    • Define clear JSON schemas for predictable responses
    • Use TypeScript interfaces for type safety
  2. Multi-Modal Inputs
    • Use appropriate content types (image_url, video_url, file_url)
    • Set detail level for images based on analysis needs
  3. Error Handling
    • Always wrap API calls in try-catch blocks
    • Handle rate limits and timeouts appropriately