tts-1.5-max — AI Audio Model

inworld

inworld

tts-1.5-max

Commercial useaudio

inworld/tts-1.5-max is a high-quality text-to-speech (TTS) model designed to generate natural, expressive, and human-like voice outputs from text. Built by Inworld AI, this model focuses on delivering realistic speech with emotional tone, clarity, and consistency—making it ideal for immersive experiences, interactive applications, and professional audio content. With support for nuanced voice modulation and low-latency performance, tts-1.5-max enables developers and creators to produce engaging voiceovers, character dialogues, and real-time conversational audio.

Model Type
coin

Pricing: 4.33 joules / generation

Input

voice

Voice to use for synthesis

temperature
Controls randomness when generating audio. Higher values produce more expressive results, lower values are more deterministic.
0.02.0
OutputAUDIO

Output will be displayed here

Fill in the inputs on the left and hit RUN to preview the results.

Examples

Explore different use cases and parameter configurations