Model Information
Context and output sizes describe the model specification. Effective request limits, availability, and billing depend on the MixRoute route.
Capabilities and Usage
- Accepts text and image prompts and returns audio with accompanying text when applicable. It is not a speech transcription or text-to-speech model.
- Describe genre, instruments, vocals, and desired duration in the prompt. Duration is not a fixed seconds parameter.
- Read candidates[].content.parts[].inlineData.data as Base64 audio and inspect inlineData.mimeType. Default audio is MP3; generated audio is 44.1 kHz stereo.
Request Example
SetMIXROUTE_API_KEY before making the request.
Request Fields
Complete request format: gemini-native.