vlmrun.image.detect(image, "person"). See Code Execution.
Image
Caption & Tag
Captions and tags for images.
Detection
Bounding boxes and confidence scores for objects, faces, and people.
Segmentation
Pixel-level masks for objects, regions, and features.
Pointing
Key points and structural features, with sub-pixel accuracy.
Generate & Edit
From a text prompt, a sketch, or an existing image.
UI Parsing
Buttons and interactive components in screenshots.
Image Tools
Crop, rotate, upscale, and colorize.
Documents
Layout Detection
Layout elements with bounding boxes and reading order.
Visual Grounding
Map text in a document to its location on the page.
Multi-Page Analysis
Correlate pages and documents in one request.
Video
Caption & Tag
Topic, summary, timestamped chapters, and content tags.
Generate & Edit
Generate a video from a text prompt.
Video Tools
Trim, sample, and extract video segments.