npx skills add ...
npx skills add zc277584121/marketing-skills --skill chrome-automation
Automate Chrome browser tasks using agent-browser CLI. Navigate pages, fill forms, click buttons, take screenshots, extract data, and replay Chrome DevTools Recorder exports.
npx skills add zc277584121/marketing-skills --skill chrome-automation
Automate browser tasks in the user's real Chrome session via the agent-browser CLI.
Prerequisite: agent-browser must be installed and Chrome must have remote debugging enabled. See
references/agent-browser-setup.mdif unsure.
This skill operates on a single Chrome process — the user's real browser. There is no session management, no separate profiles, no launching a fresh Playwright browser.
Before opening any new page, always list existing tabs first:
This returns all open tabs with their index numbers, titles, and URLs. Check if the page you need is already open:
Always use --auto-connect to connect to the user's running Chrome instance:
This auto-discovers Chrome with remote debugging enabled. If connection fails, guide the user through enabling remote debugging (see references/agent-browser-setup.md).
Chrome 144+ can expose remote debugging from chrome://inspect/#remote-debugging as a WebSocket-only endpoint. In that state the page shows Server running at: 127.0.0.1:9222, but the traditional discovery URLs return 404:
Older agent-browser versions such as 0.27.x may fail with No running Chrome instance found even though Chrome is ready. First try the latest CLI without changing the global install:
If this works, use npx -y agent-browser@latest <command> for the rest of the browser task. If it fails with an engine warning or install error, upgrade Node to 24+ or install the latest agent-browser globally.
The user may provide a recording exported from Chrome DevTools Recorder (JSON, Puppeteer JS, or @puppeteer/replay JS format). See Replaying Recordings below.
Use snapshot -i to see all interactive elements with refs (@e1, @e2, ...):
The output lists each interactive element with its role, text, and ref. Use these refs for subsequent actions.
| Action | Command |
|---|---|
| Navigate | agent-browser --auto-connect open <url> (optionally wait --load networkidle, but some sites like Reddit never reach networkidle — skip if open already shows the page title) |
| Click | snapshot -i → find ref → click @eN |
| Fill standard input | click @eN → fill @eN "text" |
| Fill rich text editor | click @eN → keyboard inserttext "text" |
| Press key | press <key> (Enter, Tab, Escape, etc.) |
| Scroll | scroll down <amount> or scroll up <amount> |
| Wait for element | wait @eN or wait "<css-selector>" |
| Screenshot | screenshot or screenshot --annotate |
| Get page text | get text body |
| Get current URL | get url |
| Run JavaScript | eval <js> |
fillkeyboard inserttextRefs (@e1, @e2, ...) are invalidated when the page changes. Always re-snapshot after:
After each significant action, verify the result:
JSON (recommended) — structured, can be read progressively:
@puppeteer/replay JS (import { createRunner })
Puppeteer JS (require('puppeteer'), page.goto, Locator.race)
navigate steps, reusing existing tabs when possible.snapshot -i) to see current interactive elementsaria/... selectors against the snapshottext/..., then CSS class hints, then screenshotsnapshot -i operates on the main frame only and cannot penetrate iframes. Sites like LinkedIn, Gmail, and embedded editors render content inside iframes.
snapshot -i returns unexpectedly short or empty resultsget text body content doesn't match what a screenshot showsUse eval to access iframe content:
Note: Only works for same-origin iframes.
Use keyboard for blind input: If the iframe element has focus, keyboard inserttext "..." sends text regardless of frame boundaries.
Use get text body to read full page content including iframes.
Use screenshot for visual verification when snapshot is unreliable.
If workarounds fail after 2 attempts on the same step, pause and explain:
find text "Dismiss" click or find text "Close" click)find text "..." click, or scroll to reveal with scroll down 300When pausing, explain clearly: what step you are on, what you expected, and what you see.
| Command | Description |
|---|---|
tab list | List all open tabs with index, title, and URL |
tab <index> | Switch to an existing tab by index |
tab new | Open a new empty tab |
tab close | Close the current tab |
open <url> | Navigate to URL |
snapshot -i | List interactive elements with refs |
click @eN | Click element by ref |
fill @eN "text" | Clear and fill standard input/textarea |
type @eN "text" | Type without clearing |
keyboard inserttext "text" | Insert text (best for contenteditable) |
press <key> | Press keyboard key |
scroll down/up <amount> | Scroll page in pixels |
wait @eN | Wait for element to appear |
wait --load networkidle | Wait for network to settle |
wait <ms> | Wait for a duration |
screenshot [path] | Take screenshot |
screenshot --annotate | Screenshot with numbered labels |
eval <js> | Execute JavaScript in page |
get text body | Get all text content |
get url | Get current URL |
set viewport <w> <h> | Set viewport size |
find text "..." click | Semantic find and click |
close | Close browser session |
snapshot -i cannot see inside iframes. See Iframe-Heavy Sites.find text strict mode: Fails when multiple elements match. Use snapshot -i to locate the specific ref instead.fill vs contenteditable: fill only works on <input> and <textarea>. For rich text editors, use keyboard inserttext.eval is main-frame only: To interact with iframe content, traverse via document.querySelector('iframe').contentDocument...When the user requests an action across multiple platforms (e.g., "publish this article to Dev.to, LinkedIn, and X"), do NOT attempt all platforms in a single conversation. Instead, launch sequential subagents/workers, one per platform. Selecting this skill is explicit permission to use subagents/workers for this multi-platform workflow.
Each platform operation consumes ~25-40K tokens (reference file + snapshots + interactions). Running 3-5 platforms in one context risks hitting the context limit and degrading late-platform accuracy. Each subagent/worker gets its own fresh context window.
Read /path/to/skills/chrome-automation/references/x.md)--auto-connect. Parallel subagents/workers would cause tab conflicts.When automating tasks on specific platforms, consult the relevant reference document for page structure details, common operations, and known quirks:
| Platform | Reference | Key Notes |
|---|---|---|
references/reddit.md | Custom faceplate-* components; networkidle never reached; unlabeled comment textbox; find text fails due to duplicate elements | |
| X (Twitter) | references/x.md | open often times out (use tab list to reuse existing tabs); click timestamp for post detail (not username); DraftJS contenteditable input (data-testid="tweetTextarea_0"); avoid networkidle |
references/linkedin.md | Ember.js SPA; Enter submits comments (use Shift+Enter for newlines); comment box and compose box share the same label; avoid networkidle; messaging overlay may block content | |
| Dev.to | references/devto.md | Fast server-rendered HTML (Forem/Rails); standard <textarea> for comments/posts (Markdown); 5 reaction types; Algolia-powered search; networkidle works normally |
| Hacker News | references/hackernews.md | Minimal plain HTML; all form fields are unlabeled; link "reply" navigates to separate page; networkidle works instantly; rate limiting on posts/comments |
For installation and Chrome setup instructions, see
references/agent-browser-setup.md.