npx skills add ...
npx skills add inference-sh/skills --skill speech-to-text
Transcribe audio to text with ElevenLabs Scribe and Whisper models via inference.sh CLI. Models: ElevenLabs Scribe v2 (98%+ accuracy, diarization), Fast Whisper Large V3, Whisper V3 Large. Capabilities: transcription, translation, multi-language, timestamps, speaker diarization, audio event tagging. Use for: meeting transcription, subtitles, podcast transcripts, voice notes. Triggers: speech to text, transcription, whisper, audio to text, transcribe audio, voice to text, stt, automatic transcription, subtitles generation, transcribe meeting, audio transcription, whisper ai, elevenlabs stt, scribe, eleven labs transcribe
npx skills add inference-sh/skills --skill speech-to-text
Install the belt CLI skill:
npx skills add belt-sh/cli
Transcribe audio to text via inference.sh CLI.

Requires inference.sh CLI (
belt). Install instructions
| Model | App ID | Best For |
|---|---|---|
| ElevenLabs Scribe v2 | elevenlabs/stt | 98%+ accuracy, diarization, 90+ languages |
| Fast Whisper V3 | infsh/fast-whisper-large-v3 | Fast transcription |
| Whisper V3 Large | infsh/whisper-v3-large | Highest accuracy |
Whisper supports 99+ languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, Hindi, Russian, and many more.
Returns JSON with:
text: Full transcriptionsegments: Timestamped segments (if requested)language: Detected languageBrowse all audio apps: belt app list --category audio
belt app sample infsh/fast-whisper-large-v3 --save input.json
# {
# "audio": "https://podcast.mp3",
# "return_timestamps": "sentence"
# }
belt app run infsh/fast-whisper-large-v3 --input input.jsonbelt app run infsh/whisper-v3-large --input '{
"audio": "https://french-audio.mp3",
"task": "translate"
}'# Extract audio from video first
belt app run infsh/video-audio-extractor --input '{"video_file": "https://video.mp4"}' > audio.json
# Transcribe the extracted audio
belt app run infsh/fast-whisper-large-v3 --input '{"audio": "<audio-url>"}'# 1. Transcribe video audio
belt app run infsh/fast-whisper-large-v3 --input '{
"audio": "https://video.mp4",
"return_timestamps": "sentence"
}' > transcript.json
# 2. Pass the transcript segments as captions
belt app run infsh/caption-videos --input '{
"video_file": "https://video.mp4",
"segments": [{"start": 0.0, "end": 2.5, "text": "<segments-from-step-1>"}]
}'# ElevenLabs STT (98%+ accuracy, diarization)
npx skills add inference-sh/skills@elevenlabs-stt
# ElevenLabs TTS (reverse direction)
npx skills add inference-sh/skills@elevenlabs-tts
# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli
# Text-to-speech (reverse direction)
npx skills add inference-sh/skills@text-to-speech
# Video generation (add captions)
npx skills add inference-sh/skills@ai-video-generation
# AI avatars (lipsync with transcripts)
npx skills add inference-sh/skills@ai-avatar-video