MiniMax-H3 with the video generation endpoint.
Key capabilities
- Multimodal video generation - Supports text-to-video, image-to-video, first-and-last-frame, and multimodal reference workflows
- Native audio - Can generate synchronized native stereo sound
- Flexible output - Provider specifications cover 768P and 2K output, 4 to 15 seconds, at 24 fps
- Asynchronous workflow - Submit a task, then poll for its result
Resolution, duration, audio, and reference-input combinations depend on the active MixRoute route. Check the Model Marketplace before relying on a specific combination in production.
Workflow
Text-to-video example
- cURL
- Python