npx skills add ...
npx skills add nvidia/nvflare --skill nvflare-fed-stats
Compute federated statistics over tabular data (count, sum, mean, stddev, var, histogram, quantile, noise-protected min/max) and image data (count, failure_count, pixel-intensity histogram) across NVFLARE sites via FedStatsRecipe — automatic and non-interactive from the dataset, feature names (header or supplied), and optionally a README or notes declaring which statistics to compute; do not use for model training conversion, hierarchical statistics, deployment, POC/production lifecycle, or failed-job diagnosis.
npx skills add nvidia/nvflare --skill nvflare-fed-stats
Data-first and automatic: point at tabular or image data and it runs end-to-end — no interaction, no user statistics code.
Use when the user asks to compute statistics, data summaries, histograms,
or quantiles across federated sites for tabular data (CSV, parquet, any
pandas-representable form) or image datasets (PNG/JPEG/BMP/TIFF folders;
DICOM/NIfTI with the matching loader), with or without an accompanying
README/notes or statistics script. Supported for tabular: count, sum,
mean, stddev, var, histogram, quantile, noise-protected min/max (variance
and stddev are distinct — never substitute one for the other); for
images: count, failure_count, pixel-intensity histograms. Both paths use
FedStatsRecipe generation, simulator validation, completeness checks.
Do not use for model training conversion (route to nvflare-convert-pytorch,
nvflare-convert-lightning, or nvflare-convert-huggingface), a failed or
stalled existing job (route to nvflare-diagnose-job), or generic
pandas/data-science help without federated intent.
If a request combines federated statistics and model-training conversion,
treat it as two independent jobs and workflows: do not merge or automatically
chain them, do not route the combination to nvflare-orient, and ask which
workflow to run first before generating or running either job. Recommend
nvflare-fed-stats first only when the user's purpose is to understand data
distribution; handle conversion later as a separate request.
Hierarchical statistics, production deployment, Kubernetes, POC lifecycle,
and privacy-policy design beyond the recipe's built-in knobs are out of
scope. Statistics outside the supported set — categorical counts,
correlations, custom aggregations — are reported as unsupported, never
silently dropped or approximated.
nvflare agent inspect data <path> --format json first; its dataset block is the evidence — do not hand-roll data
inspection. dataset.modality: image follows the image
path (references/image-statistics.md with
assets/image_stats_client.py); dataset.modality: tabular supplies site
layout, per-site row counts, and feature names with dtype classes when
header is present. On header: ambiguous (no names extracted),
names must come from the request, a README/metadata file, or a names
file — else fail closed with a precise missing-input report (ask once
only when an interactive channel exists); never invent or auto-number
names. A schema_agreement mismatch or columns_truncated schema
fails closed (the latter unless the user declares a feature subset);
counts_approximate: true means verify site sizes before bin-cap
decisions. On 2.8.x CLIs (no dataset block), apply the same rules from
references/statistics-mapping.md. Read any statistics script or
notebook as optional intent evidence (statistics, read options,
splits, histogram ranges) without importing or executing it.importlib.util.find_spec,
never a raising import. Quantiles additionally require fastdigest
(Rust toolchain to build): same preflight; on failure, fail that
statistic closed, report the product error, and complete the rest.
Load the shared dependency-install.md only when an install is needed.value_counts/nunique, correlations, custom
aggregations — numeric features only). count is always included
because the privacy cleansers need it. Continue with the supported
subset, stating what was excluded and why; load
references/statistics-mapping.md when requests exceed the standard set.client.py — image path: from assets/image_stats_client.py
per its reference; tabular: from assets/df_stats_client.py, a
DFStatisticsCore subclass whose load_data() reads the user's data —
a script's loading logic when one exists, else a plain pandas read
(supplied names for headerless data) — returning
{dataset_name: DataFrame} (default data) parameterized by site
identity. Do not port statistic math; DFStatisticsCore computes it
all. Pre-split per-site directories define site names and count; for
flat single-source data the site count must come from the request or a
declaration (missing fails closed), with deterministic seeded
partitions unless shared data is explicitly requested.nvflare recipe show fedstats --format json; for preflights/job.py use:
from nvflare.recipe import SimEnv; from nvflare.recipe.fedstats import FedStatsRecipe (never package root).
Load only SimEnv Execution from
../nvflare-shared/references/conversion-common.md before writing or validating the runner.
Use statistic_configs and one site list: FedStatsRecipe(..., sites=sites, ...); SimEnv(clients=sites, ...).
The recipe already assigns those clients; never use
SimEnv(num_clients=...) or both forms. Let SimEnv derive thread
count, or set num_threads=len(sites). Histograms default to 20 bins,
no range; set one only from a script, declaration, or user answer
(images: bit depth), else use protected min/max estimation. Reduce bins
when small sites demand it (20 bins needs 206+ rows per site); report
it. Keep and state StatsJob defaults: min_count=10, noise
0.1–0.3, and max_bins_percent=10.validation-evidence.md: compile
checks, recipe construction, one simulator run, then output
completeness — the output JSON exists, parses, and covers every
configured statistic per feature, site, and Global — using ephemeral
commands only. Generate NO validation scripts or helper files: beyond
client.py, job.py, and user-requested data preparation (seeded
partitions for flat data), the skill leaves nothing behind. Numeric
parity is harness-owned (references/stats-job-validation.md); stop
at the first failed rung and report the product error.count is non-null, so missingness shifts
denominators), and a compact per-site and global summary (aggregates
only — never raw rows or values) with the output JSON path and the
case-mix caveat: compare site rows before Global.count; stddev/var also require sum and mean
(second-round prerequisites — expand and state it). State the applied
default selection when the user expressed none.range (estimated from noise-protected
min/max, stated in the report).client.py, job.py, and user-requested data prep.fedstats recipe before constructing it; present selection and mapping
before generating code.client.py and job.py, keeping decisions within
this skill and its references. Report blockers: missing names,
non-numeric data, missing quantile dependency, undersized sites,
non-parameterizable loaders.dependency-install.md; audit and preview the
redacted plan, then confirm it unless unattended installation was explicitly
requested. Host permission remains an additional gate. After installation,
run requested validation without another execution prompt.Always read this SKILL.md. The standard tabular path is inline; load
details when their phase needs them: references/statistics-mapping.md
(mapping, config grammar), references/stats-job-validation.md
(validation, output locations, harness parity contract),
references/image-statistics.md plus assets/image_stats_client.py
(image path), assets/df_stats_client.py (tabular template), shared
references only for exceptions. Never preemptively; never depend on
NVFLARE repository examples being present.