Skip to main content
Generate comprehensive, contextual captions for videos using state-of-the-art vision-language models. Perfect for accessibility, content management, and automated video analysis workflows.

Example video to be captioned.

Example Response

This is an example of the response from the Chat Completions API example (using the video shown above):

Usage Example

For best results, we recommend using the Structured Outputs API to get responses in a structured and validated data format.

FAQ

You can ask simply ask for a more detailed caption by providing a more detailed prompt. In most cases, you can provide the number of words you want the caption to be, and the model will generate a more detailed caption.
  • Content Types: presentation, tutorial, interview, documentary, news
  • Scenes: office, outdoor, studio, classroom, conference room
  • People: presenter, audience, speaker, interviewer
  • Objects: whiteboard, charts, graphs, computer, microphone
The video segments come in the format of a list of dictionaries with start time, end time, and description fields.
Yes, the structured output includes segments with timestamps that break down the video into different parts with descriptions for each segment.