npx skills add ...
npx skills add nvidia/skills --skill nemo-evaluator-plugin
Use when working on the Evaluator plugin CLI, jobs, SDK-backed specs, metric types, or plugin-owned Evaluator skills.
npx skills add nvidia/skills --skill nemo-evaluator-plugin
Use this skill for evaluation tasks against a running NeMo Platform server. The plugin-backed CLI interface is nemo evaluator; the legacy generated nemo evaluation API command group is not the target surface for new guidance.
nemo CLI: source .venv/bin/activateCheck plugin status from the CLI:
To view available metric names, run:
To view a specific metric schema, pass a metric name from the metric_types list above:
Inspect all the registered metric schema contracts:
Note: use
nemo evaluator evaluate explainas the source of truth for the current plugin input schema. It will return a large json schema response, so strongly prefernemo evaluator metric-typeswhen you only need metric names and corresponding schemas.
Evaluation spec is a payload that is provided to CLI as an input to execute evaluation.
At a high level, a spec describes:
metrics: bundled Evaluator SDK metric configurationsdataset: inline rows to evaluate or platform FilesetRef that contains the datasetparams: optional Evaluator SDK execution parameterstarget: optional model or agent target for online evaluationSee the LLM-judge spec example at assets/specs/llm_as_judge.json.
The checked-in spec examples use bundled SDK metrics. The fields under metrics[*].payload are generated by bundle_metric(metric, CloudpickleMetricBundlePackager()).
To see the pattern for configuring a pre-defined SDK metric, for example ExactMatchMetric, and converting it into bundled metric JSON, inspect build_metric_bundle_example() in generate_example_specs.py and run:
When using the nemo evaluator evaluate run command, results are saved into local temporary directories and the link is printed to stdout.
Prefer the --spec-file named argument over inline shell JSON because metric bundles include serialized payloads.
Examples of various specs are provided in the assets/specs directory.
exact-match metricSee the spec example at assets/specs/exact_match_metric.json.
LLM-Judge metricUses an LLM to score responses. See the spec example at assets/specs/llm_as_judge.json.
Use the nemo evaluator evaluate submit command to create a durable evaluation job. The response of this command returns a job handler object instead of the evaluation result.
The submit response includes the generated job's name field, for example nemo-evaluator-zlhn1ecd. Wait for the job to complete, then list and download the job results.
Evaluator Python SDK client is exposed as evaluator variable on NeMoPlatform instance:
See examples of using the plugin SDK interface in plugin_sdk_examples.py.
Make sure not to print any secrets to stdout since this can be collected as logs
For LLM-judge setup notes, see LLM Judge Notes.
For evaluator API key auth, see Evaluator API Auth.
For local and cluster troubleshooting, see Evaluation Troubleshooting.
nemo evaluator evaluate explainuv run --frozen python skills/nemo-evaluator-plugin/scripts/generate_example_specs.pynemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.jsonnemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_benchmark.jsonnemo evaluator evaluate run --spec-file skills/nemo-evaluator-plugin/assets/specs/llm_as_judge.jsonnemo evaluator evaluate submit \
--spec-file skills/nemo-evaluator-plugin/assets/specs/exact_match_metric.jsonnemo jobs get-status <job-name>
nemo jobs get <job-name>
nemo jobs results list <job-name>
nemo jobs results download aggregate-scores --job <job-name> --output-file aggregate-scores.json
nemo jobs results download row-scores --job <job-name> --output-file row-scores.jsonlfrom nemo_platform import NeMoPlatform
platform_client = NeMoPlatform(base_url="http://localhost:8080")
status = platform_client.evaluator.plugin_status()