Skip to main content
POST

Introduction

Veo is a multimodal video generation model from Google Vertex AI, supporting text-to-video (T2V), first-frame constraint, and first-last frame constraint (3.1 series only) for generating coherent videos. Call through MixRoute unified video interface: first submit task to get task_id, then query task to poll status and retrieve results.

Authentication

Bearer Token, e.g., Bearer sk-xxxxxxxxxx

Supported Models

Call Flow

  1. Submit Task: POST /v1/video/generations, pass model, prompt and Veo-specific parameters.
  2. Poll Status: GET /v1/video/generations/{task_id}, until status is succeeded or failed.
  3. Get Result: When successful, the url in the response contains video data (Veo may return data:video/mp4;base64,... or OSS link).

Veo-Specific Parameters

string
required
Video generation prompt describing the scene and actions.
integer
Video duration (seconds), supported values: 4, 6, 8.
string
Aspect ratio, only supports: 16:9, 9:16.
string
Resolution: 720p, 1080p.
string
First frame reference image (URL or Base64), for image-to-video/first-frame constraint.
string
Last frame reference image (only veo-3.1 series supported), works with first frame for first-last frame constraint.
boolean
Whether to generate synchronized audio. Fast version models ignore this parameter and always include audio.
integer
Number of videos to generate per request, range 1-4.
For more parameters (such as personGeneration, addWatermark, seed), see Submit Video Task.

Request Examples

For complete request/response specifications and multi-model comparison, see Submit Video Task and Query Video Task.