npx skills add ...
npx skills add nvidia/skills --skill doca-bench
Run `doca_bench` (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or BlueField Arm. Use it to discover enabled benchmark libraries, capture a reproducible command/version/device/environment baseline, compare stable runs against a declared tolerance, or diagnose configuration, device-binding, workload-precondition, and measurement failures. Trigger for requests such as measuring BlueField compression speed, NIC RDMA throughput, crypto latency, or a pre-upgrade baseline. Do not use for application end-to-end timing, custom benchmark code, DOCA installation, or binary patches.
npx skills add nvidia/skills --skill doca-bench
doca_bench)Where to start: This is a tool skill for invoking doca_bench,
the cross-library micro-benchmark harness. Open
TASKS.md and start at
## configure for the three-axis decision
(target library × workload shape × measurement axis), then
## run for the smoke-before-bulk flow. Open
CAPABILITIES.md when the question is what
doca_bench can measure, which DOCA libraries it can drive, or
how to interpret throughput / latency / op-rate output without
fooling yourself on warm-up or steady-state. If DOCA is not
installed yet, route to
doca-setup first; if the install
version is < 2.7.0, doca_bench is not shipped on this host.
The CLASSES of doca_bench questions this skill is built to answer,
each with one worked example. The class is the load-bearing piece;
the worked example is one instance.
CAPABILITIES.md ## Capabilities and modes
TASKS.md ## run. The same shape answers
"send-side throughput of DOCA RDMA" — doca_bench is
cross-library, not single-library.doca_bench actually drive on this
install?" — worked example: "is doca_sha enumerable on a
granular-build install". Answered by the built-in query system
surfaced in
CAPABILITIES.md ## Capabilities and modes
TASKS.md ## configure step 2
(probe-before-bench). Empty enumeration = library not installed,
not bench failure.CAPABILITIES.md ## Error taxonomy
layer 5 + TASKS.md ## test (the eval-loop
overlay treats warm-up / steady-state / outliers as
re-iteration triggers, not one-shot facts).doca_bench shows
zero ops for AES-GCM but doca_caps says the device supports
it". Answered by the layered error taxonomy in
CAPABILITIES.md ## Error taxonomy
(config-syntax → device-binding → library-precondition →
workload-precondition → measurement-soundness → version →
cross-cutting) + TASKS.md ## debug.TASKS.md ## test (capture command line +
version + device + as-deployed environment alongside the
numbers; quoting numbers without the four-tuple is the
cross-version regression-hunt failure mode).doca_bench returns nothing for library X — what does that
mean?" — worked example: "empty output for DOCA SHA".
Answered by the empty-output interpretation rules in
TASKS.md ## debug +
CAPABILITIES.md ## Error taxonomy.
Re-route through
doca-caps for the coarse
per-device per-library capability ground truth, then back
into bench once the capability is confirmed present.This skill serves external operators, developers, and AI agents who need a reproducible, vendor-supported way to measure DOCA library performance on the user's actual install and device. Concretely:
doca_bench baseline against the new state.It is not for users debugging the doca_bench source code,
and not a substitute for the live public DOCA Bench guide on
docs.nvidia.com.
doca_bench is shipped as a tool (a single CLI binary plus a
companion app for the remote half of remote-memory / RDMA / Eth
scenarios), not a library you link against. The skill uses the
same kind: tool three-file shape as the rest of the bundle so
the agent's task-verb contract
(configure / build / modify / run / test / debug) is uniform
across libraries, services, and tools — even when individual
verbs collapse to a routing stub for a shipped binary.
Load this skill when the user is — or the agent needs to — invoke
doca_bench on a real host with DOCA ≥ 2.7.0 installed (or
inside the public NGC DOCA container with the equivalent version)
to measure performance of a DOCA library. Concretely:
tools/bench/doca_bench/configuration.hpp are not
interchangeable.TASKS.md ## debug).Do not load this skill for general DOCA orientation, library
API work, or installation. For those, use
doca-public-knowledge-map,
the matching libs/<library> skill, or
doca-setup. Do not load it for
application-level end-to-end benchmarking either — doca_bench
measures the DOCA library surface, not the user's application
above it.
This is a thin loader. Substantive material lives in two companion files:
CAPABILITIES.md — what doca_bench can measure (the
cross-library scope, the three-axis configuration model, the
documented operating modes, the warm-up / pipeline / multi-core
concepts that constrain measurement soundness), the version
overlay (doca-bench-specific facts on top of the canonical
doca-version rules), the layered error taxonomy
(config-syntax / device-binding / library-precondition /
workload-precondition / measurement-soundness / version /
cross-cutting), the observability surface (screen + CSV
output, real-time stats, query system), and the safety
posture (the public guide's "not for production" warning,
the host vs BlueField execution rule, the companion-app
attack surface).TASKS.md — step-by-step workflows for the in-scope task
verbs: configure (the three-axis decision + the
probe-before-bench step), build (route to install — the
binary is shipped, the companion app is shipped), modify
(refuse — do not patch the bench binary; modify the bench
invocation instead), run (the smoke-before-bulk flow),
test (the eval loop — warm-up, steady-state, outliers,
cross-version), debug (walk the error taxonomy layer by
layer), plus a Deferred task verbs block routing
out-of-scope questions and a Command appendix of
doca_bench-specific invocation classes.The skill assumes a host where DOCA ≥ 2.7.0 is already installed
(or the public NGC DOCA container is running at an equivalent
version) and the operator has whatever permissions the public
guide requires for doca_bench to bind devices and allocate
resources on their platform.
This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:
--help on the installed version are the
authoritative answer. Inventing a flag is the most common
hallucination failure for this skill.doca_bench CSV or stdout. The output formats are documented;
if a user wants to script against them, the right answer is
"read the live guide, write the parser against your installed
version".samples/ or reference/ subtree. This is a thin
loader for a documented CLI; substantive material lives on
the public page and in --help.SKILL.md first to confirm the user's question is
in scope (the user actually wants to invoke doca_bench for
measurement, not learn about a DOCA library in general).doca_bench measures, the three-axis model, the
version overlay, the error taxonomy, observability surface,
and safety posture, see CAPABILITIES.md.configure, build, modify, run, test,
debug — see TASKS.md.doca-public-knowledge-map
— routing to the public DOCA Bench page on docs.nvidia.com
and the rest of the public DOCA documentation set.doca-version — the canonical
version-detection chain, four-way match rule, NGC container
semantics, and headers-win-over-docs rule. The
## Version compatibility section in this skill is a thin
overlay on top of doca-version; the body lives there.doca-structured-tools-contract
— the bundle-wide contract for structured-output helper tools.
Bench-runner / bench-snapshot executables that satisfy the
detect-prefer-fallback-report loop are deferred to PR2; the
contract is consumed here in advance so the
## Command appendix in TASKS.md is infra-aware
from PR1.doca-setup — env preparation,
install verification, hugepages, NUMA awareness, and the I
have no install yet path with the public NGC DOCA container.doca-debug — the cross-cutting
debug ladder. Bench surfaces its own error taxonomy in
CAPABILITIES.md ## Error taxonomy;
when the cause turns out to be below DOCA (driver, firmware,
NUMA), the bench taxonomy hands off to doca-debug.doca-caps — the sibling DOCA tool
for the coarse per-device per-library capability snapshot.
Bench probes capability at finer grain via its own query
system; doca_caps is the cheaper first step to confirm the
device is even visible to DOCA.libs/<library> skill — e.g.
doca-comch,
doca-compress — for
the workload-side preconditions, capability-query rules, and
error-taxonomy overlays of the library under test. Bench
drives the library; the library skill explains what
"healthy" means for it.