npx skills add ...
npx skills add vercel/vercel-plugin --skill benchmark-e2e
npx skills add vercel/vercel-plugin --skill benchmark-e2e
End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill injection, launches dev servers, verifies everything works, analyzes conversation logs, and produces an improvement report for overnight self-improvement loops.
The same skill content is published under more than one repo. The install counts are split across them; any of these commands works.
Single-command pipeline that creates projects, exercises skill injection via claude --print, launches dev servers, verifies they work, analyzes conversation logs, and generates actionable improvement reports.
Options:
| Flag | Description | Default |
|---|---|---|
--quick | Run only first 3 projects | false |
--base <path> | Override base directory | ~/dev/vercel-plugin-testing |
--timeout <ms> | Per-project timeout (forwarded to runner) | 900000 (15 min) |
The orchestrator chains four stages sequentially, aborting on failure:
claude --print with VERCEL_PLUGIN_LOG_LEVEL=tracerun-manifest.json, extracts metricsreport.md and report.json with scorecards and recommendationsrun-manifest.jsonWritten by the runner at <base>/results/run-manifest.json. Links all downstream stages to the same run.
The analyzer and verifier read this manifest to correlate sessions precisely instead of guessing from directory listings.
events.jsonlThe orchestrator writes NDJSON events to <base>/results/events.jsonl tracking pipeline lifecycle:
report.jsonMachine-readable report at <base>/results/report.json for programmatic consumption:
Run the pipeline repeatedly with a cooldown between iterations:
Each run produces timestamped report.json and report.md files. Compare across runs to track improvement.
The pipeline enables a closed feedback loop:
bun run scripts/benchmark-e2e.ts exercises the plugin against realistic projectsreport.json lists which skills were expected but never injected, with exact slugssuggestedPatterns entries (copy-pasteable YAML) to add missing frontmatter patterns; use recommendations to fix hook logicreport.json across runs: verdict should trend from "fail" → "partial" → "pass"For overnight automation, combine with the loop above. Wake up to reports showing exactly what improved and what still needs work.
Prompts never name specific technologies — they describe the product and features, letting the plugin infer which skills to inject.
| # | Slug | Expected Skills |
|---|---|---|
| 01 | recipe-platform | auth, vercel-storage, nextjs |
| 02 | trivia-game | vercel-storage, nextjs |
| 03 | code-review-bot | ai-sdk, nextjs |
| 04 | conference-tickets | payments, email, auth |
| 05 | content-aggregator | cron-jobs, ai-sdk |
| 06 | finance-tracker | cron-jobs, email |
| 07 | multi-tenant-blog | routing-middleware, cms, auth |
| 08 | status-page | cron-jobs, vercel-storage, observability |
| 09 | dog-walking-saas | payments, auth, vercel-storage, env-vars |
// Each line is one JSON object:
{ "stage": "pipeline", "event": "start", "timestamp": "...", "data": { "baseDir": "...", "quick": false } }
{ "stage": "runner", "event": "start", "timestamp": "...", "data": { "script": "...", "args": [...] } }
{ "stage": "runner", "event": "complete", "timestamp": "...", "data": { "exitCode": 0, "durationMs": 120000 } }
// On failure:
{ "stage": "verify", "event": "error", "timestamp": "...", "data": { "exitCode": 1, "durationMs": 5000, "slug": "04-conference-tickets" } }
{ "stage": "pipeline", "event": "abort", "timestamp": "...", "data": { "failedStage": "verify", "exitCode": 1, "slug": "04-conference-tickets" } }interface ReportJson {
runId: string | null;
timestamp: string;
verdict: "pass" | "partial" | "fail";
gaps: Array<{
slug: string;
expected: string[];
actual: string[];
missing: string[];
}>;
recommendations: string[];
suggestedPatterns: Array<{
skill: string; // Skill that was expected but not injected
glob: string; // Suggested pathPattern glob
tool: string; // Tool name that should trigger injection
}>;
}while true; do
bun run scripts/benchmark-e2e.ts
sleep 3600
donerm -rf ~/dev/vercel-plugin-testing