Low-latency speech generation with steerable prompts and expressive audio tags. gemini-3.1-flash-tts-preview by Google is available through the Flatkey unified API and is included in our Audio Generation Models collection. Review its current pricing, context, and availability before integrating it.
Find the right model for speech and soundtracks
Browse audio generation models for turning text into speech or creating music for video. Choose your task to explore available models, input requirements, and API prices.
Featured models include gemini-3.1-flash-tts-preview and sonilo-video-to-music, ranked by live weekly usage when available.
Audio Generation Models
Generate production-ready music from any video with synchronized timing, optional speech preservation, configurable ducking, segment-level direction, and up to 10 audio variants. Supports MP3, M4A, and WAV output through Flatkey's asynchronous API. sonilo-video-to-music by Sonilo is available through the Flatkey unified API and is included in our Audio Generation Models collection. Review its current pricing, context, and availability before integrating it.