npx skills add ...
npx skills add nvidia/skills --skill amc-run-video-calibration
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
npx skills add nvidia/skills --skill amc-run-video-calibration
Activate this skill when the user has pre-recorded MP4 files and wants to calibrate them via the AMC REST API. Typical prompts:
Drives calibration through the REST API on user-supplied pre-recorded MP4 files — no CLI scripts or Docker bind-mounts required, just a running microservice and your files.
Do not use this skill for live RTSP streams or rtsp://... URLs; route those requests to skills/amc-run-rtsp-calibration/SKILL.md.
Guide the agent through project creation, sorted MP4 upload, local asset resolution, UI fallback only when necessary, project verification, calibration, polling, evaluation, and optional VGGT refinement for a user-provided multi-camera dataset.
skills/amc-setup-calibration-stack/SKILL.md)https://<HOST_IP>:<MS_PORT> for remote AMC, or http://localhost:<MS_PORT> on loopback) and UI URLcam_00.mp4, cam_01.mp4, … time-synchronized, ~1920×1080requestsprojects/ directory, you know the host PROJECTS_DIR/calibrate unless CONFIRM_CALIBRATION=true is set. RUN_VGGT=true remains a separate opt-in for the optional VGGT step.Video files uploaded via this skill are transmitted to the AutoMagicCalib backend (REST endpoint). Only use this skill when the backend is deployed on a trusted platform / network.
VIDEO_DIR, BASE_URL, and PROJECT_NAME.CONFIG_FILE, ALIGNMENT_JSON, LAYOUT_PNG, GT_ZIP, FOCAL_LENGTHS, and DETECTOR_TYPE.CONFIRM_CALIBRATION, RUN_VGGT, PROJECTS_DIR, CALIBRATION_TIMEOUT_SECONDS, and VGGT_TIMEOUT_SECONDS.BASE_URL should use HTTPS for non-loopback hosts. Set ALLOW_INSECURE_HTTP=true only for trusted development setups that intentionally use remote plain HTTP.(Video-file naming and the microservice URL are specified under Prerequisites above — collect the inputs below.)
cam_*.mp4, uploaded sorted alphabetically.The script searches the videos dir, its first-level subdirectories, and its parent. If exactly one match is found, it is used; otherwise the script prints the searched locations and continues to explicit path or UI fallback:
| File | Candidate filenames | UI fallback |
|---|---|---|
| Calibration settings | settings.json, config.json, calibration_config.json | UI Step 3: Parameters |
| Alignment JSON | alignment_data.json | UI Step 4: Alignment |
| Layout PNG | layout.png | UI Step 4: Alignment |
Posting the settings file replaces UI Step 3 and may pin the detector (resnet/transformer), which is passed to /calibrate separately — see Step 4.
GT.zip with _World_Cameras_Camera_XX/ folders (enables evaluation metrics)1269.0, 1099.5, 1099.5resnet (default, fast) or transformer (slower, better under occlusion)See root README.md "Custom Dataset" section for input-video guidelines and ground-truth format.
| Script | Purpose | Key inputs |
|---|---|---|
| run_video_calibration.py | Executes create-project, upload, verify, calibrate, poll, evaluate, and optional VGGT refinement for a local MP4 dataset. | Required: BASE_URL, PROJECT_NAME, VIDEO_DIR. Optional: CONFIG_FILE, ALIGNMENT_JSON, LAYOUT_PNG, GT_ZIP, FOCAL_LENGTHS, DETECTOR_TYPE, CONFIRM_CALIBRATION, RUN_VGGT, PROJECTS_DIR, CALIBRATION_TIMEOUT_SECONDS, VGGT_TIMEOUT_SECONDS, ALLOW_INSECURE_HTTP. |
All endpoints below are implemented end-to-end in the Complete Python Script — the prose is the workflow plus the decisions the agent must make; the script is the authoritative runnable.
POST /v1/create_project (form field project_name) → save the returned project_id.
POST /v1/upload_video_files/<project_id> (multipart files). Upload sorted alphabetically — the server assigns camera indices by upload order. The bundled script rejects non-contiguous or non-zero-based camera sequences up front; the directory must contain cam_00.mp4, cam_01.mp4, ... with no gaps.
For each of calibration-settings, alignment, and layout, run this resolution:
Auto-use rule: if exactly one match is found, the script uses it automatically and prints the resolved path. No extra prompt occurs for that file.
VIDEO_DIR, one level of subdirectories under VIDEO_DIR, and VIDEO_DIR.parent for the candidate filenames (table above).verify_project as the source of truth for whether the UI-supplied alignment/layout data is complete.Upload each file resolved locally:
| File | Endpoint | Notes |
|---|---|---|
| Calibration settings | POST /v1/config/<project_id> (JSON, posted as-is) | Replaces UI Step 3 (rectification, bundle-adjustment, evaluation, detector, …). Non-2xx is surfaced — never silently fall back. Skip on the UI-fallback path. |
| Alignment | POST /v1/upload_alignment/<project_id> (alignment_data.json) | |
| Layout | POST /v1/upload_layout/<project_id> (layout.png) | |
| Ground truth (optional) | POST /v1/upload_gt_file/<project_id> (GT.zip) | Enables evaluation metrics |
| Focal lengths (optional) | POST /v1/upload_focal_length/<project_id> (repeated focal_length=) | Overrides GeoCalib estimates |
Upload all resolved local files first, in any order. After the local uploads are complete, continue to Step 5 only for unresolved files, then run Step 6 exactly once to verify the assembled project.
After a successful settings POST, parse the file for "detector" / "detector_type" — if it's "resnet" or "transformer", use that value for the /calibrate call in Step 7 (detector is a separate API parameter, not consumed by /config).
If any of settings / alignment / layout was not resolved in Step 3, direct the user to the appropriate UI step:
<project_id>, go to Step 3: Parameters, tune via the settings dialog (or accept defaults), click Save." Also: before the /calibrate call, ask the user which detector to use (resnet or transformer) using the host's question mechanism; if none is available, ask in chat and wait. UI Step 3 does not cover detector choice.<project_id>, go to Step 4: Alignment, upload layout, mark correspondence points, click Save."Wait for user confirmation. For non-interactive script runs, provide the needed files up front; the script exits with a clear message rather than waiting on input. Do not require local access to AMC's projects/ storage for the UI fallback; Step 6 is the canonical server-side verification step.
POST /v1/verify_project/<project_id> → must return {"project_state": "READY"} before calibrating.
Confirm the plan before calibrating. Whether the settings file and detector were auto-detected or asked, present a short summary and get explicit user confirmation before POST /calibrate using the host's question mechanism; if none is available, ask in chat and wait. The resolved values are the defaults, so confirming is one click, but the user can switch the detector or skip an auto-detected settings file. The standalone Python script prompts when stdin is interactive; in non-interactive runs it exits before /calibrate unless CONFIRM_CALIBRATION=true is set. Summarize:
resnet or transformer (the value to be sent).GET /v1/get_project_info/<project_id> every 10 s — project_info.project_state goes RUNNING → COMPLETED (or ERROR, pull the log). Typical time: 10–60 min depending on video length and detector. The bundled script defaults to a 90 minute cap through CALIBRATION_TIMEOUT_SECONDS=5400; raise that env var for longer runs instead of silently killing the process.
GET /v1/result/<project_id>/evaluation_statistics (only if GT was uploaded; includes Average L2 distance(m) and Average reprojection error 0(px)), and GET /v1/amc/calibrate/<project_id>/log for the calibration log. If GT was uploaded and evaluation_statistics returns non-200, surface that HTTP error instead of treating it as a missing-GT case.
get_project_infoproject_info.project_state is the AMC calibration lifecycle for the project: RUNNING → COMPLETED (or ERROR).
project_info.vggt_state is also per-project, a project-scoped VGGT refinement lifecycle rather than a direct global service or model-load status. A newly created project can report vggt_state: "INIT" even when the VGGT model is present and mounted. The expected VGGT lifecycle is INIT → READY after AMC calibration completes → RUNNING while VGGT refinement runs → COMPLETED (or ERROR).
Use vggt_state == "READY" only as the gate for optional VGGT refinement in Step 10. Interpret INIT on a new or uncalibrated project as normal project state. If AMC calibration is complete and the project remains in a non-ready VGGT state, confirm VGGT setup and model availability with the setup skill checks and service logs.
After AMC calibration completes, read vggt_state from GET /v1/get_project_info/<project_id>.
vggt_state == "READY", ask the user whether to run VGGT refinement using the host's question mechanism; if none is available, ask in chat and wait.POST /v1/vggt/calibrate/<project_id>, poll vggt_state via get_project_info, then GET /v1/vggt_results/<project_id>/evaluation_statistics.amc-setup-calibration-stack and rerun this optional step later.The standalone Python script prompts only when stdin is interactive. In non-interactive runs, set RUN_VGGT = True to opt in; otherwise the script prints that VGGT is ready and continues without blocking.
Use the bundled script from the amc-run-video-calibration skill package, not from the auto-magic-calib repo root. If the user points the agent at this skill folder directly instead of installing it, set AMC_VIDEO_SKILL_DIR to the directory containing this SKILL.md, or run the command from that directory. Set BASE_URL, PROJECT_NAME, and VIDEO_DIR; optional env vars are CONFIG_FILE, ALIGNMENT_JSON, LAYOUT_PNG, GT_ZIP, FOCAL_LENGTHS, DETECTOR_TYPE, RUN_VGGT, REPO_ROOT, PROJECTS_DIR, CONFIRM_CALIBRATION, CALIBRATION_TIMEOUT_SECONDS, VGGT_TIMEOUT_SECONDS, and ALLOW_INSECURE_HTTP. The script implements AMC readiness checks, UI fallback, explicit confirmation gating, bounded polling, and refined statistics retrieval.
Interactive run with auto-detection for local settings/alignment/layout:
Non-interactive run with all required local files supplied up front:
project_state == "COMPLETED" after polling.verify_project returned READY before calibration (including manual alignment/UI fallback paths).Average L2 distance(m) < 1.5Average reprojection error 0(px) < 5ERROR state.cam_*.mp4 datasets. Live RTSP streams and the bundled sample dataset are out of scope.CONFIRM_CALIBRATION=true.PROJECTS_DIR when AMC writes project outputs outside the default projects/ directory.| Issue | Fix |
|---|---|
verify_project state not READY | Confirm videos uploaded and alignment + layout are present (either via API or via UI manual alignment) |
| Manual alignment still not accepted after UI step | User likely did not click Save or the UI data is incomplete; rerun verify_project and repeat UI Step 4 |
Calibration stuck RUNNING > 90 min | GET /v1/amc/calibrate/<id>/log — usually insufficient tracklets (scene too static). See "Custom Dataset" guidelines in root README. |
Immediate ERROR state | Check video naming: must be cam_00.mp4, cam_01.mp4, … contiguous |
| Low L2 but high reprojection | Provide explicit focal_length override via Step 3 |
| VGGT stays non-ready after AMC completes | INIT is expected for a new project. After AMC calibration reaches COMPLETED, the project should transition to READY before optional VGGT refinement when VGGT is configured. If refinement is required and the state remains INIT or otherwise non-ready, confirm VGGT setup and model availability with setup skill Step 2 and MS logs. |
| Upload timeout | Large videos — bump timeout=300 to e.g. 600 in the script |
A downstream Multi-View 3D Tracking skill fetches the MV3DT-format calibration directly from the microservice (this skill does not download it; it returns the project_id). After this skill reports COMPLETED:
GET /v1/result/{project_id}/mv3dt_result?result_type=amc → mv3dt_output.zip (contains transforms.yml).COMPLETED (Step 10): ?result_type=vggt → vggt_mv3dt_output.zip.skills/amc-setup-calibration-stack/SKILL.md — start MS + UI first.skills/amc-run-sample-calibration/SKILL.md — verify the stack with the bundled sample before trying your own.skills/amc-run-rtsp-calibration/SKILL.md — same calibration tail, but sourcing footage from live RTSP streams through VIOS.Root README.md "Custom Dataset" and "Calibration Workflow (UI)" sections document input-video guidelines and the UI-driven alternative to this API flow.