npx skills add ...
npx skills add nvidia/k8s-launch-kit --skill k8s-launch-kit-config
Use this skill when the user needs help understanding, creating, or editing a k8s-launch-kit (l8k) configuration file (l8k-config.yaml or cluster-config.yaml). Activate for: config file questions, parameter tuning, subnet configuration, NV-IPAM setup, DOCA driver settings, maintenance concurrency, NIC configuration operator settings, changing MTU, VFs, resource names, or understanding what any config field does.
npx skills add nvidia/k8s-launch-kit --skill k8s-launch-kit-config
PREREQUISITE: Read
../k8s-launch-kit-shared/SKILL.mdfor install paths, global flags, and output modes.
Understand, create, or edit l8k configuration files.
| File | Source | Purpose |
|---|---|---|
cluster-config.yaml | Generated by l8k discover | Hardware inventory plus the resolved deployment profile |
l8k-config.yaml | User-created or copied from cluster-config | Full config with both hardware + deployment settings |
Fresh discovery and file-backed generation resolve and persist settings with this precedence:
Discovery with --user-config follows a stricter refresh contract: only
clusterConfig is replaced. Every other section remains as loaded unless an
explicit CLI flag overrides its corresponding field. Generation still fills
missing profile fields when it consumes that config.
| Section | What It Controls |
|---|---|
networkOperator | Operator namespace, version, image repository, Helm repository, and skipHelmChart ownership switch |
docaDriver | OFED/DOCA driver image, version, blacklist settings |
maintenance | Maintenance Operator, SR-IOV drain, and legacy OFED upgrade concurrency |
nvIpam | NV-IPAM per-node block size, IP pool ranges, and subnet generation |
sriov | VF count, resource prefix, MTU, link type |
hostdev | Host device resource name |
rdmaShared | RDMA shared device resource name |
ipoib | IPoIB master interface, resource name |
macvlan | MacVLAN master interface, mode |
nicConfigurationOperator | NIC firmware template settings |
spectrumX | OVS bridge config, multiplane mode, RDMA settings |
profile | Profile selection criteria (fabric, deployment, multirail) |
clusterConfig[] | Per-group hardware: NICs, nodes, capabilities, selectors |
Each clusterConfig[] entry has these key fields:
identifier — group name (used for NicNodePolicy naming). For groups with both machineType and gpuType resolved, this is the lowercased machine/GPU identity with complete NVIDIA segments removed and common machine segments shortened (ThinkSystem → ts, PowerEdge → pe), bounded to 30 bytes with balanced component prefixes and a 6-character deterministic hash when needed; otherwise a fallback group-N. The Launch Kit machine node label uses the same value.machineType — server model (e.g. PowerEdge-XE9680); populated from nvidia.com/gpu.machine label or DMI fallback.gpuType — GPU SKU (e.g. NVIDIA-H200); populated from nvidia.com/gpu.product label or nvidia-smi fallback. Note: this field used to be called productType — the rename happened to disambiguate it from the server model. Old productType: keys in hand-authored configs must be renamed to gpuType:.netplanManaged — true when any worker has an NVIDIA PF whose current MAC
is selected by a host netplan match.macaddress stanza with a non-empty
set-name. The flag identifies a potential conflict with NCO udev naming;
host-specific MACs are not persisted. If generation would emit a
NicInterfaceNameTemplate for this group, clean up the affected set-name
stanzas and re-run discovery instead of editing this flag by hand.capabilities.nodes.{sriov,rdma,ib} — what the underlying hardware supports.pfs[] — physical function list with PCI address, device ID, RDMA device,
network interface, traffic class, rail, NUMA, GPU affinity, and model (the
VPD model/description string read from NicDevice.Status.modelName).
model is genuinely dual-port (2-port/Dual-port) keeps a rail per port. Run l8k discover --collapse-nic-rails=false to emit one rail per PF (legacy/dev behaviour).nodeSelector — Kubernetes node selector for this group. Source groups key on the machine label written by l8k discover: nvidia.kubernetes-launch-kit.machine: <identifier>. Auto-merged groups (different machineTypes sharing a GPU type) key on nvidia.kubernetes-launch-kit.gpu: <gpuType> instead — the GPU label retains its discovered value, including NVIDIA, so the merged selector binds correctly across source machineTypes.workerNodes — explicit hostnames (populated by discovery).For the full field-by-field reference with types, defaults, and descriptions,
read references/config-reference.md.
The presets/ directory contains pre-recorded topologies for known hardware combinations. A preset is a topology.yaml file with the following shape:
Lookup is exact-match on (machineType, gpuType). No any-GPU fallback — a preset that doesn't declare gpuType: is rejected at load time. Multi-variant presets for the same machine (different GPU SKUs) live in separate directories with composite names like PowerEdge-XE9680-H200 / PowerEdge-XE9680-B200. The directory name is shown by l8k preset list and is what l8k generate --for <name> accepts.
Validation deviations. When the matched preset's PF count, PCI addresses, or device IDs don't exactly match discovered hardware, the preset is NOT applied — discovery keeps the live-discovered topology (traffic/rail/NUMA), because overlaying a preset onto a different PCI layout would corrupt the classification (a coincidentally-overlapping PCI address would inherit the preset's unrelated traffic/rail/GPU fields). The discrepancies are recorded under clusterConfig[*].presetDeviation and every subsequent config load re-emits a warning listing each deviation. The preset's authoritative topology is overlaid only on an exact match (zero deviations); presetApplied: true appears only in that case. Part-number and PSID differences are expected (firmware/SKU variants) and never block application.
l8k discover) to generate a baseline, then edit it.l8k generate without
repeating profile flags; use flags only for overrides.l8k schema to discover the Network Operator release keys supported by
the installed l8k version.nvIpam subnets are auto-generated if not specified — one per rail using non-routable ranges.nvIpam.perNodeBlockSize defaults to 10; generation warns when it is lower than sriov.numVfs because the per-node allocation may be too small for all VFs.docaDriver.unloadThirdPartyRDMAModules: true auto-populates UNLOAD_THIRD_PARTY_RDMA_MODULES from discovered OFED-dependent modules.docaDriver.env is an advanced escape hatch. Values override generated MOFED environment entries by name, duplicate custom names use the last value, and bad driver options can disrupt node networking.--overwrite-existing.references/config-reference.md — Complete annotated YAML, including maintenance value restrictions