Context--
Price (In / Out)Free / Free
Training CutoffPublic release timeline through 2026-08
CategoryAudio Model
About IndexTTS
IndexTTS is a controllable zero-shot text-to-speech model family from IndexTeam and Bilibili. The 2.5 release clones voices from a single reference clip, supports Chinese, English, Japanese, Spanish, and Arabic, and reports a 2.28x real-time-factor improvement over IndexTTS 2 in its technical report.
Capabilities
text to speechvoice cloningmultilingualemotion controlpronunciation controlspeed control
Input Modalities
textaudio
Output Modalities
audio
Technical Details
- API Identifier
- IndexTeam/IndexTTS-2.5
- Category
- Audio Model
Tags
text-to-speechvoice-cloningspeech-synthesismultilingualemotion-controlopen-weights
Pricing
Token pricing for IndexTTS API usage.
Input Tokens
Free
per million tokens
Output Tokens
Free
per million tokens
Pricing Calculator
Input cost$0.00
Output cost$0.00
Estimated monthly cost$0.00
Open model weights; hosting and GPU inference costs depend on the deployment environment.
IndexTeamopen source
Open-source speech models for controllable zero-shot TTS
1 ModelsFounded 2025China
View full profile