qwen3-tts — AI Audio Model
qwen
qwen3-tts
qwen/qwen3-tts is a next-generation, open-weight text-to-speech (TTS) model family developed by Qwen (Alibaba), designed to generate highly natural, expressive, and controllable human-like speech from text. It stands out for its ultra-low latency streaming, multilingual support, and advanced voice control capabilities, making it suitable for both real-time applications and high-quality audio production. Built with a unified end-to-end architecture, Qwen3-TTS eliminates traditional pipeline limitations and delivers fast, high-fidelity speech with strong contextual understanding and emotional expression. This model costing on 5000 input characters.
Pricing: 10.35 joules / generation
Output will be displayed here
Fill in the inputs on the left and hit RUN to preview the results.
Examples
Explore different use cases and parameter configurations