npx skills add ...
npx skills add inference-sh/skills --skill text-to-speech
Convert text to natural speech with Inworld TTS, ElevenLabs, DIA TTS, Kokoro, Chatterbox, and more via inference.sh CLI. Models: Inworld TTS-2 (100+ languages, emotion steering), Inworld TTS 1.5 (ultra-low latency), ElevenLabs (premium, 22+ voices, 32 languages), DIA TTS (conversational), Kokoro TTS, Chatterbox, Higgs Audio, VibeVoice (podcasts). Capabilities: text-to-speech, voice cloning, multi-speaker dialogue, podcast generation, expressive speech, emotion/delivery steering, character voices. Use for: voiceovers, audiobooks, podcasts, accessibility, video narration, IVR, voice assistants, gaming characters, avatar audio. Triggers: text to speech, tts, voice generation, ai voice, speech synthesis, voice over, generate speech, ai narrator, voice cloning, text to audio, elevenlabs, eleven labs, voice ai, ai voiceover, speech generator, natural voice, inworld, inworld tts, character voice, game voice, npc voice
npx skills add inference-sh/skills --skill text-to-speech
Install the belt CLI skill:
npx skills add belt-sh/cli
Convert text to natural speech via inference.sh CLI.

Requires inference.sh CLI (
belt). Install instructions
| Model | App ID | Best For |
|---|---|---|
| Inworld TTS-2 | inworld/text-to-speech-2 | 100+ languages, emotion steering with [brackets], delivery modes |
| Inworld TTS 1.5 Max | inworld/text-to-speech-1-5-max | Low latency (<200ms), 15 languages |
| Inworld TTS 1.5 Mini | inworld/text-to-speech-1-5-mini | Ultra-low latency (~120ms), 15 languages |
| ElevenLabs TTS | elevenlabs/tts | Premium quality, 22+ voices, 32 languages |
| DIA TTS | infsh/dia-tts | Conversational, expressive |
| Kokoro TTS | infsh/kokoro-tts | Fast, natural |
| Chatterbox | infsh/chatterbox | General purpose |
| Higgs Audio | infsh/higgs-audio | Emotional control |
| VibeVoice | infsh/vibevoice | Podcasts, long-form |
Inworld TTS-2 supports natural-language steering with [brackets] — control emotion, volume, speed, and non-verbals inline with text:
Delivery modes: STABLE (consistent), BALANCED (natural, default), CREATIVE (expressive).
Built-in voices (271+ across 15 languages): Sarah, Alex, Ashley, Dennis, Hana, Blake, Luna, Clive, and many more. Browse all voices in the Inworld TTS Playground or list programmatically via GET https://api.inworld.ai/voices/v1/voices.
For real-time applications where speed matters:
The easiest way to create a talking head video is P-Video-Avatar with built-in TTS — no separate audio step:
For models without built-in TTS (OmniHuman, PixVerse), generate speech first:
Browse all apps: belt app list