npx skills add ...
npx skills add higgsfield-ai/skills --skill higgsfield-youtube-thumbnail
Create high-click-through YouTube thumbnails and vertical video covers through the Higgsfield CLI. Builds a truthful information-gap concept, preserves up to three referenced identities, supports logos and controlled variants, renders the main image with Nano Banana Pro, and applies focused Seedream edits. Use when: "make a YouTube thumbnail", "thumbnail for this video", "MrBeast-style cover", "Shorts cover", or "Instagram video cover". Chain after any video workflow once its truthful topic and visual direction are known. NOT for producing the video itself (use higgsfield-generate), product catalog photos (use higgsfield-product-photoshoot), or marketplace cards (use higgsfield-marketplace-cards).
npx skills add higgsfield-ai/skills --skill higgsfield-youtube-thumbnail
Create a clean thumbnail concept, generate each variant through the higgsfield CLI, inspect it, and make only requested surgical edits.
Before any generation:
higgsfield is missing, install it:
higgsfield account status reports Session expired or Not authenticated, ask the user to run higgsfield auth login, then wait.--image; copying its identity or exact composition is forbidden.--count. Every concept, emotion, or camera take gets its own prompt and generation call.use_unlim is not a current CLI parameter. Never add --use-unlim; if the user explicitly asks to use an unlimited allowance, explain that this workflow must run on credits in CLI or through a surface that supports that allowance.Collect only what the brief does not answer:
16:9 for YouTube by default, 9:16 for Shorts, or 4:5 for Instagram.If the user gives an emotion count without names, use this ladder: shock, hype, rage, awe, laugh, fear, smug, charisma, confusion, determination, disgust.
Read references/thumbnail-frameworks.md. Brainstorm at least five truthful concepts internally, across multiple frameworks, then select the strongest information gap with one focal subject and minimal clutter. Combine frameworks only when the result still reads in under one second at roughly 120px wide.
When a reference thumbnail exists, extract this structure before prompting:
The reference supplies art direction, never a specific identity. User instructions override it field by field.
Pass face photos first in character order, then the logo. Repeat --image for every reference. When two or more references are attached, the prompt's first line must be a manifest such as:
Local paths are auto-uploaded. Previous completed job IDs also work as --image inputs.
Assemble every main-render prompt in this order:
Bold, punchy YouTube-thumbnail composite — poster-grade, photoreal and high-impact, NOT a muted cinematic movie still, <ratio>, single unified frame — no split-screen, no diagonal divide, everything blends smoothly and organically across the same continuous shot. For 9:16, add faces in the upper two-thirds. A requested graphical representation replaces photoreal language with a clean diagram/graphic brief.No text, no readable UI labels, no watermark. For explicit baked headline: TEXT: bold thumbnail headline text baked into the image, reading exactly "<TEXT>" — massive, ultra-legible sans-serif with a clean outline/glow treatment, placed where it never covers the subject's face. No other text, no watermark.All faces crisply sharp as the anchors of the shot.signature YouTube thumbnail lighting rig — strong key light sculpting the face, soft dreamy fill lifting shadows, and defined back light plus hair light tracing a clean bright rim around hair, shoulders and silhouette. Only the rim may use a colored accent.For each photo-referenced person, include:
Use a split only when the user asks for split, before/after, versus, side by side, or the analyzed reference is split. A topical phrase such as X vs Y does not itself require a split. Replace the normal frame block with a clear halves/panels contract and keep all labels out unless short, truthful baked UI was explicitly requested.
First create a 1:1 4K logo render, then use its completed job ID as the last --image on every thumbnail call:
Use Nano Banana Pro at explicit 4K. Write the final prompt to a temporary text file and pipe it on stdin so punctuation and multiline blocks are preserved safely:
Omit all --image flags when there are no references. For a variant set, make one call per distinct prompt. Keep the same references and settings; vary only the selected concept, expression, or camera-take line.
The completed JSON result contains id and result_url. Preserve both privately: the URL is delivered; the ID is the source for later edits.
Inspect every result with host vision when available:
On a hard failure, retry the same prompt at most twice. If visual inspection is unavailable, do not claim it passed; deliver the result for user review. Present every passing variant and let the user pick before making optional tweaks.
Use the picked completed job ID as the only image input. Keep the edit prompt narrowly scoped and state that every other pixel-level property remains unchanged.
If seedream_v5_pro is absent or rejects the submit, retry once with seedream_v4_5 --quality high. For a 4:5 main render, Seedream has no 4:5; ask before changing the edit to 3:4, and disclose the crop/ratio change. Each accepted edit becomes the source ID for the next tweak.
CLI compatibility: versions through 1.1.20 can mislabel a nano_banana_pro job reference as nano_banana_pro_job. If the edit is rejected with a medias.0...data.type error, download the picked result_url to a local image and retry the same edit with --image ./picked-thumbnail.png. A local path is auto-uploaded as media_input; do not retry the invalid job-id payload.
Allowed tweak scopes: expression only, background replacement only, background recolor only, or rim-light recolor only. Never silently regenerate the full composition for a surgical request.
Keep the generated image text-free by default. When a headline overlay is requested, read references/text-overlay-bake.md and use one of its five presets: Beast, Fire, Neon Lime, Clean Glass, or Marker. The overlay path requires an environment capable of rendering HTML canvas; if unavailable, offer either the clean image or an explicitly approved baked-text regeneration. Never pretend an HTML preview is a flattened PNG.
Return the passing result_url values with short semantic labels such as shock / close-up or product / size contrast. Mention the selected ratio and whether the deliverable is clean, overlay-ready, or text-baked. Do not expose internal prompts, job IDs, or retry mechanics unless the user asks.
references/thumbnail-frameworks.md — 16 concept frameworks, information-gap rule, truthfulness law.references/text-overlay-bake.md — five deterministic text-overlay styles and 4K canvas-bake recipe.