npx skills add ...
npx skills add bytedance/deer-flow --skill podcast-generation
Use this skill when the user requests to generate, create, or produce podcasts from text content. Converts written content into a two-host conversational podcast audio format with natural dialogue.
npx skills add bytedance/deer-flow --skill podcast-generation
This skill generates high-quality podcast audio from text content. The workflow includes creating a structured JSON script (conversational dialogue) and executing audio generation through text-to-speech synthesis.
When a user requests podcast generation, identify:
/mnt/user-dataGenerate a structured JSON script file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}-script.json
The JSON structure:
Call the Python script:
Parameters:
--script-file: Absolute path to JSON script file (required)--output-file: Absolute path to output MP3 file (required)--transcript-file: Absolute path to output transcript markdown file (optional, but recommended)[!IMPORTANT]
- Execute the script in one complete call. Do NOT split the workflow into separate steps.
- The script handles all TTS API calls and audio generation internally.
- Do NOT read the Python file, just call it with the parameters.
- Always include
--transcript-fileto generate a readable transcript for the user.- The TTS provider and its concurrency are selected automatically from environment variables — you do not choose or tune them.
The script JSON file must follow this structure:
Fields:
title: Title of the podcast episode (optional, used as heading in transcript)locale: Language code - "en" for English or "zh" for Chineselines: Array of dialogue lines
speaker: Either "male" or "female"paragraph: The dialogue text for this speakerWhen creating the script JSON, follow these guidelines:
User request: "Generate a podcast about the history of artificial intelligence"
Step 1: Create script file /mnt/user-data/workspace/ai-history-script.json:
Step 2: Execute generation:
This will generate:
ai-history-podcast.mp3: The audio podcast fileai-history-transcript.md: A readable markdown transcript of the podcastRead the following template file only when matching the user request.
The generated podcast follows the "Hello Deer" format:
After generation:
/mnt/user-data/outputs/present_files toolThe following environment variables must be set:
VOLCENGINE_TTS_APPID and VOLCENGINE_TTS_ACCESS_TOKENMINIMAX_API_KEYVOLCENGINE_TTS_CLUSTER: Volcengine TTS cluster (optional, defaults to "volcano_tts")VOLCENGINE_TTS_VOICE_TYPE_MALE: Volcengine male voice type (optional, defaults to zh_male_yangguangqingnian_moon_bigtts)VOLCENGINE_TTS_VOICE_TYPE_FEMALE: Volcengine female voice type (optional, defaults to zh_female_sajiaonvyou_moon_bigtts)Voice type overrides are trimmed; unset or blank values use the listed defaults.
Auto-selected by environment variables:
VOLCENGINE_TTS_APPID + VOLCENGINE_TTS_ACCESS_TOKEN set → Volcengine TTS (default).MINIMAX_API_KEY set → MiniMax TTS (/v1/t2a_v2).PODCAST_GENERATION_PROVIDER=volcengine|minimax.MiniMax overrides: MINIMAX_API_HOST (default https://api.minimaxi.com),
MINIMAX_TTS_MODEL (default speech-2.6-hd), MINIMAX_TTS_VOICE_MALE
(default male-qn-qingse), MINIMAX_TTS_VOICE_FEMALE (default female-tianmei).
Concurrency is owned by each provider internally — MiniMax runs single-threaded to reduce rate-limit failures, Volcengine uses 4 workers. There is no caller-facing concurrency knob; transient rate limits are handled by automatic retry with backoff.