qwen3-tts — AI Audio Model

qwen

qwen

qwen3-tts

Commercial useaudio

qwen/qwen3-tts is a next-generation, open-weight text-to-speech (TTS) model family developed by Qwen (Alibaba), designed to generate highly natural, expressive, and controllable human-like speech from text. It stands out for its ultra-low latency streaming, multilingual support, and advanced voice control capabilities, making it suitable for both real-time applications and high-quality audio production. Built with a unified end-to-end architecture, Qwen3-TTS eliminates traditional pipeline limitations and delivers fast, high-fidelity speech with strong contextual understanding and emotional expression. This model costing on 5000 input characters.

Model Type
coin

Pricing: 10.35 joules / generation

Input

voice

Voice to use for synthesis

OutputAUDIO

Output will be displayed here

Fill in the inputs on the left and hit RUN to preview the results.

Examples

Explore different use cases and parameter configurations