agent component provides methods for interacting with VLM Run’s Orion Agents for multi-modal chat completions.
Chat Completions
Generate responses from the agent using the chat completions:Image Analysis
Analyze images using the agent:Video Analysis
Analyze videos using the agent:Structured Outputs
Get structured JSON responses using TypeScript interfaces:Document Analysis
Analyze documents and PDFs:Agent Execution
Beyond chat completions,client.agent.execute() runs a named agent (optionally with
skills) over structured inputs and returns an execution you can poll:
Execution Mode
The response carriesexecution_mode, which reports whether the server replayed a
skill’s cached pipeline.py or ran the full LLM agent loop. See
program execution.
The Node.js SDK cannot select the mode yet.
AgentExecutionConfig serializes only
prompt, jsonSchema, skills, serviceTier, and orchestrationMode, so a mode
key is dropped before the request is sent. Use the
Python SDK or a raw
cURL request to force a mode.SDK Reference
client.agent.execute()
Execute an agent over structured inputs.
Parameters:
Returns:
Promise<AgentExecutionResponse> with execution_mode set to the path taken.
client.agent.completions.create()
Create a chat completion with the agent.
Parameters:
Returns:
Promise<ChatCompletionResponse>
Message Content Types
Best Practices
-
Structured Outputs
- Define clear JSON schemas for predictable responses
- Use TypeScript interfaces for type safety
-
Multi-Modal Inputs
- Use appropriate content types (
image_url,video_url,file_url) - Set
detaillevel for images based on analysis needs
- Use appropriate content types (
-
Error Handling
- Always wrap API calls in try-catch blocks
- Handle rate limits and timeouts appropriately