npx skills add ...
npx skills add higgsfield-ai/skills --skill higgsfield-video-explainer
Build a complete non-photoreal narrated explainer or story video from ordered 10-second blocks: one narrator, one universal style key, one Seed Audio take and one Gemini Omni clip per block, then server-side assembly with explainer_video. Use when: "make an explainer video", "explain this in a video", "turn this topic or document into a narrated video", "tell this story as an animated video", "make a faceless narrated video", or "show me explainer styles". Supports live CMS presets, custom style references, mascot/faceless modes, two aspects, and optional burned subtitles. NOT for: photoreal films, ads/UGC, talking heads, podcasts, motion typography reels, one-off clips without narration, or editing a finished video.
npx skills add higgsfield-ai/skills --skill higgsfield-video-explainer
Run the MCP video-explainer workflow through Higgsfield CLI. Lock one visual style key, write one narration line and one matching visual prompt per 10-second block, generate every voice take first, generate every clip second, then immediately assemble the ordered pairs with explainer_video.
Never use the monolithic video_explainer job in this skill.
| MCP workflow operation | CLI equivalent |
|---|---|
get_explainer_presets | higgsfield preset list video-explainer --json |
resolve_explainer_preset | higgsfield preset resolve video-explainer <preset_id> --json |
generate_image / nano_banana_pro | higgsfield generate create nano_banana_2 ... |
list_voices | higgsfield voices list --json |
generate_audio / seed_audio | higgsfield generate create seed_audio ... |
generate_video / gemini_omni | higgsfield generate create gemini_omni ... |
job_status | --wait --json or higgsfield generate wait <job_id> --json |
explainer_video | higgsfield generate create explainer_video ... |
nano_banana_2 is the public CLI id for the Nano Banana Pro style-key model used by the MCP workflow.
If higgsfield is unavailable, install it:
If higgsfield account status fails, ask the user to run higgsfield auth login, then wait.
Inspect the live contracts before the first submission:
Collect choices in two separate turns, in this order. Never merge them.
Always load the live CMS catalog:
Show the preset names with their thumbnail/video preview URLs. Say one short line asking the user to pick a preset, describe a custom style, or attach style-reference images, then end the turn. Do not ask production questions in the same turn. Choosing a style is mandatory; never choose silently unless the user explicitly says “you choose.”
Skip this turn only when the request already contains explainer preset id: <uuid>. Confirm that UUID exists in the live catalog and keep it for Phase 1.
Only after style selection, collect every unresolved setting:
N = duration_minutes × 6 fixed 10-second blocks.16:9 by default or 9:16 vertical.patrick, caveat, marker, or anton; never choose silently.Every choice belongs to the user unless they explicitly delegate it.
For local style donors, pass each path with a repeated --image. For a web image, download it locally first or use an existing uploaded media ID.
| Phase | Output | CLI |
|---|---|---|
| 0 Ask | style first; then duration, language, character, aspect, subtitles | preset list + user questions |
| R Research | verified facts and sources | available research tools |
| 1 Style key | one universal style image | preset resolve or nano_banana_2 |
| 2 Narration | N labeled narration lines | reasoning |
| 3 Block prompts | N labeled video prompts | reasoning |
| 4 Voice | user selects one voice; generate N takes | voices list + seed_audio |
| 5 Clips | generate N 10-second clips | gemini_omni |
| 6 Assemble | one final MP4 | explainer_video |
Read references/prompts.md before Phases 1–3.
For a real topic, use available web research tools and authoritative sources to verify enough facts for every block. Keep a short Sources list. Never script a factual explainer from memory alone.
For a personal story, skip web research and use only details supplied by the user. Invent nothing factual.
Write one reusable STYLE descriptor: medium, palette, line/fill behavior, texture/finish, then non-photorealistic, illustrated, not a photo, no live-action, no realism.
Resolve the hidden style image into the active workspace:
Keep the returned media_id as STYLE_KEY_ID. Skip image generation: this imported media is the style key. Build the STYLE descriptor from the returned preset name plus the mandatory non-photoreal rules. Do not recreate a preset from its name.
The preset reference controls framing. If it conflicts with the aspect requested in Phase 0, stop and let the user choose rather than silently fighting the reference.
Generate exactly one key image. Use the abstract swatch template from references/prompts.md, or its mascot variant when character mode is enabled. Repeat --image for every style donor:
Use 9:16 for vertical. Keep the completed image job UUID as STYLE_KEY_ID; later CLI generations can reuse a completed job UUID as an image reference.
Write exactly N labeled narration blocks in the selected language:
Write exactly N labeled English prompts using the template in references/prompts.md:
For mascot mode, Block 1 greets by gesture with mouth closed, the final block waves a sign-off, and middle blocks use consistent cameos only when useful. For faceless mode, use stylistic scenes only. Keep one clear action per block.
List the live voices, present the choices, and wait for the user to select one narrator:
Keep the selected voice's exact id and type (preset or element). Never invent or auto-pick a voice unless the user explicitly delegates it.
Generate one completed seed_audio job per narration block, always with the same voice:
Record every audio job UUID in block order. Regenerate only a failed or excessively long take. Shorten that block or adjust --speech_rate modestly when needed. Do not begin Phase 5 until all N audio jobs are complete.
Generate one completed 10-second gemini_omni clip per block. Attach the same style key to every call:
Use 9:16 when selected. Record every video job UUID in block order. Independent jobs may run concurrently inside this phase, but the audio-phase barrier is strict. Re-submit only failed blocks. Never silently replace gemini_omni; inspect the live video catalog if the model is unavailable.
Create blocks.json with at least two ordered block pairs. The CLI model contract requires typed references, so use the generic completed-job types:
Submit the server-side assembler immediately:
Use --width 720 --height 1280 for vertical. When subtitles are enabled, add the chosen font:
The assembler keeps each block at exactly 10 seconds: it centers short voice takes, pitch-safely speeds small overruns, never stretches video, concatenates blocks in order, and optionally burns timed captions. Total duration is exactly N × 10 seconds.
Do not use local ffmpeg, the legacy assembly scripts, or the monolithic video_explainer job.
N narration lines and prompts, one selected voice, and N completed audio jobs.N completed video jobs and exact one-to-one block pairing with no missing or duplicate IDs.preset list; never reuse or fabricate an ID.higgsfield generate wait <job_id> --json; never duplicate a running job.Return the final assembled video URL, exact duration, aspect, narration language, selected style, narrator, subtitle status, and a Sources list for researched topics. Keep intermediate job IDs and loose asset URLs internal unless requested.