npx skills add ...
npx skills add zc277584121/marketing-skills --skill browser-screenshot
Take focused, region-specific screenshots from web pages using a dedicated headed Chrome profile with persistent login state. Navigates to the right page based on user context (URL, search query, social media post), locates the target region via DOM selectors, and crops to a clean, focused screenshot without attaching to the user's daily browser.
npx skills add zc277584121/marketing-skills --skill browser-screenshot
Take focused screenshots of specific regions on web pages — a Reddit post, a tweet, an article section, a chart, etc. — not just a full-page dump.
Prerequisite:
agent-browsermust be installed. This skill uses its own headed Chrome with a dedicated persistent profile; the user's daily Chrome does not need remote debugging.
Use the same dedicated persistent profile convention as my-chrome-automation:
Run every browser command through ab. Do not use --auto-connect, generic CDP discovery, or the user's daily Chrome profile. The dedicated profile preserves cookies and login state across tasks while allowing the user's normal Chrome to remain open and untouched.
Choose a descriptive SESSION for the task when practical. Do not run multiple tasks against the same profile concurrently; reuse it sequentially to avoid profile-lock conflicts.
If the target site requires login, open it with ab, pause for the user to complete login or verification in the dedicated headed window, and then continue with the same profile and session.
A second Chrome launch with the same profile may exit before creating a DevTools endpoint. Never fall back to whichever Chrome happens to be open. Either reuse a known Agent Browser session for this profile, or explicitly attach only to a Chrome process whose command line contains the exact --user-data-dir=$PROFILE_DIR value.
On macOS, identify that dedicated process and its local CDP port like this:
Require non-empty, unambiguous PROFILE_PID and CDP_PORT values before attaching. If they cannot be verified, ask the user to close the dedicated automation window and retry. Do not close a browser that this task only attached to.
This skill handles the full pipeline:
NEVER output an uncropped full-viewport or full-page screenshot as a final result. Full screenshots contain too much noise (nav bars, sidebars, ads, unrelated content) and are unsuitable as article illustrations. Every screenshot MUST be cropped to a focused region.
The browser is for capturing, not for browsing. Before opening anything in Chrome, use text-based tools (WebSearch, WebFetch) to find candidate pages, read their content, and decide which ones are actually worth screenshotting.
This saves significant time — most candidate pages won't be worth screenshotting, and you can eliminate them without the overhead of browser navigation.
Skip the WebSearch/WebFetch phase and go directly to Chrome browsing when:
In these cases, Chrome browsing replaces WebSearch — navigate to the platform's search page, browse results, and evaluate pages visually before deciding what to screenshot.
The right page depends on the context of the article and how recent/notable the subject is:
| Subject Type | Best Page to Find | How to Find It |
|---|---|---|
| New model/feature launch (< 6 months) | Official blog post announcing it | WebSearch "<model name>" site:<vendor-domain> blog |
| Established product (> 6 months) | Product landing page or docs overview | WebSearch "<model name>" official page |
| Open-source model | HuggingFace model card or GitHub repo | Direct URL: huggingface.co/<org>/<model> |
| API service | API documentation page | WebSearch "<service name>" API docs |
Note: This table lists common subject types but is not exhaustive. Apply the same research-first strategy to any subject type — find the most authoritative and visually clean source page for the topic at hand.
Core principle: Less is more. Focus on content, not chrome.
A good screenshot source contains a focused, self-contained piece of information — a paragraph of text, a key quote, a data table, a diagram. It should NOT be a busy page full of buttons, navigation, sidebars, and interactive elements.
Rule of thumb: If the region you plan to capture contains more interactive UI elements (buttons, links, nav items) than readable text content, it's a bad crop. Find a more content-rich region, or pick a different page entirely.
Before opening in the browser, validate URLs with WebFetch (lightweight HEAD/GET) to avoid wasting time on 404s or redirects:
Think about what the article reader needs to see in this screenshot:
| Article Context | What to Capture | Target Region |
|---|---|---|
| Introducing a model in a lineup | Model name + key tagline/description | Blog hero section or HF model card header |
| Comparing capabilities | Feature highlights or spec table | Blog section showing specs/features |
| Discussing a specific feature | The feature description | Relevant section heading + 1-2 paragraphs |
| Showing a product/service | Brand identity + value prop | Landing page hero (title + subtitle + visual) |
The screenshot should make the reader think "ah, that's what this model/product is" — not "what am I looking at?"
Check if the page is already open in the dedicated session. Reuse its existing tabs when they have the correct login and page state.
| User Provides | Strategy |
|---|---|
| Direct URL | ab open <url> |
| Search query | ab open https://www.google.com/search?q=<encoded-query> → find and click the best result |
| Platform + topic | Construct platform search URL (see below) → locate target content |
| Vague description | Google search → evaluate results → navigate to best match |
| Platform | Search URL Pattern |
|---|---|
https://www.reddit.com/search/?q=<query> | |
| X / Twitter | https://x.com/search?q=<query> |
https://www.linkedin.com/search/results/content/?keywords=<query> | |
| Hacker News | https://hn.algolia.com/?q=<query> |
| GitHub | https://github.com/search?q=<query> |
| YouTube | https://www.youtube.com/results?search_query=<query> |
After navigation, wait for content to settle:
Note: Some sites (Reddit, X, LinkedIn) never reach
networkidle. Ifopenalready shows the page title in its output, skip the wait. Usewait 2000as a safe alternative.
This is the critical step. The goal is to find a CSS selector that precisely wraps the content to capture.
Take an annotated screenshot to understand the page layout:
Take a snapshot to see the page's accessibility tree:
Identify the target container element. Look for:
<article>, <main>, <section>[data-testid="..."], [data-id="..."]Verify with get box to confirm the element has a reasonable bounding box:
This returns { x, y, width, height }. Sanity-check:
If the selector is hard to find, use eval to explore the DOM:
Common container selectors for popular platforms:
| Platform | Target | Typical Selector |
|---|---|---|
| A post | shreddit-post, [data-testid="post-container"] | |
| X / Twitter | A tweet | article[data-testid="tweet"] |
| A feed post | .feed-shared-update-v2 | |
| Hacker News | A story + comments | #hnmain .fatitem |
| GitHub | A repo card | [data-hpc], .repository-content |
| YouTube | Video player area | #player-container-outer |
| Generic article | Main content | article, main, [role="main"], .post-content, .article-body |
These selectors may change over time. Always verify with
get boxbefore using.
If the selector matches multiple elements (e.g., multiple tweets on a timeline), narrow it down:
Then target a specific one using :nth-of-type(N) or a unique parent selector.
Best when the target element fits within the viewport.
Then crop using the bounding box (see Cropping).
Best when the target might be larger than the viewport or when precise cropping is needed.
Then crop (see Cropping).
Use ImageMagick (magick on IMv7, convert is deprecated) to crop the screenshot to the target region. Add padding for visual breathing room.
Critical: On macOS Retina displays, screenshots are captured at 2x resolution. A 1728x940 viewport produces a 3456x1880 image. You MUST account for this:
Detect the scale factor: Compare viewport size vs actual image dimensions:
Multiply get box coordinates by the scale factor before cropping:
Important:
get boxreturns floating-point values. Round them to integers before passing to ImageMagick.
Padding: Use 12–20px (viewport px). Increase to ~30px if the target has a distinct visual boundary (card, bordered box). Use 0 if the user wants a tight crop.
reddit-post-screenshot.png, tweet-screenshot.pngAfter cropping, read the output image to verify it captured the right content:
If the crop is wrong (missed content, too much whitespace, wrong element), adjust the selector or bounding box and retry.
When DOM-based location is uncertain — the selector might be wrong, multiple candidates exist, or the target is ambiguous — use JS-injected highlighting to visually confirm before cropping.
Inject a highlight border on the candidate element:
Take a screenshot and visually inspect:
Read the screenshot to check if the red border surrounds the correct content.
If correct, remove the highlight and proceed with cropping:
If wrong, try the next candidate or refine the selector, re-highlight, and re-check.
get box result looks suspicious (too large, too small, zero-sized)Before taking the final screenshot, clean up the page for a better result:
Use with caution: Hiding fixed elements might remove important context. Only run this when overlays visibly obstruct the target region.
Some cookie consent banners (e.g., Jina AI's Usercentrics) live in shadow DOM or iframes and cannot be dismissed via JS click() or remove(). Don't waste time with multiple JS attempts. Instead:
For consistent, high-quality screenshots, set the viewport before capturing:
Choose a viewport width that makes the target content render cleanly — not too cramped, not too stretched.
get box returns null or zero-sizedget count "<selector>" to verify.wait 2000 and retry.screenshot --full with get box (they use the same coordinate system).get box x values may be offset.get box and snapshot -i cannot see inside iframes.eval to access iframe content:
open succeeded but page content is wrongtab list to find the correct tab and tab goto <N> to switch.document.fonts.ready. Force-resolve it first:
PROFILE_PID="$(ps -axo pid=,command= | awk -v profile="--user-data-dir=$PROFILE_DIR" 'index($0, profile) && index($0, "--remote-debugging-port=") && $0 !~ /Helper/ {print $1; exit}')"
CDP_PORT="$(lsof -nP -a -p "$PROFILE_PID" -iTCP -sTCP:LISTEN | awk 'NR > 1 && $9 ~ /^127\.0\.0\.1:[0-9]+$/ {split($9, parts, ":"); print parts[2]; exit}')"
ab() {
agent-browser --namespace "$NAMESPACE" --session "$SESSION" --cdp "$CDP_PORT" "$@"
}WebFetch: <candidate-url>
→ Check status code, title, and content snippet
→ If 404 or redirect to unrelated page, try next candidateab tab listab wait --load networkidle