Skip to main content
Generate comprehensive, contextual captions for images using state-of-the-art vision-language models. Perfect for accessibility, content management, and automated image analysis workflows.
Image captioning example showing detailed scene description

Example image to be captioned.

Example Response

This is an example of the response from the Chat Completions API example (using the image shown above):

Usage Example

For best results, we recommend using the Structured Outputs API to get responses in a structured and validated data format.

FAQ

You can ask simply ask for a more detailed caption by providing a more detailed prompt. In most cases, you can provide the number of words you want the caption to be, and the model will generate a more detailed caption.
  • Common Objects: person, car, truck, bus, bicycle, motorcycle
  • Scenes: street, building, park, forest, beach, etc.
  • Time-of-Day: morning, afternoon, evening, night
  • Weather: sunny, cloudy, rainy, snowing, etc.
The tags come in the format of a list of strings.