npx skills add ...
npx skills add google/skills --skill gke-node-notready
Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations. Use when nodes show NotReady, when the kubelet stops posting node status, or when workloads are evicted or stuck Pending due to node health. Don't use for pod-level application failures (use gke-workload-troubleshooting), autoscaler scale-up/scale-down decisions (use gke-cluster-autoscaler), or non-GKE compute.
npx skills add google/skills --skill gke-node-notready
Use this skill to systematically diagnose why one or more GKE nodes report a
NotReady (or Ready: Unknown) status and to propose safe remediations. A
NotReady status means the node's kubelet is not reporting to the control plane
correctly, so Kubernetes stops scheduling new Pods on the node, which can reduce
application capacity and cause downtime.
This skill operates non-interactively and enforces a read-only diagnostics
boundary: gather evidence first, then propose a fix (a kubectl/gcloud
command or a GitOps manifest change) for a human to apply. Never mutate the
cluster, drain, delete, or recreate nodes automatically.
[!IMPORTANT] First rule out an expected
NotReady: a node that is newly provisioning, upgrading, being repaired, cordoned, or scaling down will transiently reportNotReady. Only treat it as a fault if it persists beyond the expected window.
project_id, cluster_name,
cluster_location, and node_name non-interactively from the user prompt,
active SETTINGS.md, or environment defaults (kubectl config current-context,
gcloud config get-value project).gcloud container clusters get-credentials {cluster_name} --location {cluster_location} --project {project_id}.
If the cluster is unreachable or commands fail (sandbox/dry-run/offline),
present the exact diagnostic commands for a human to run and continue the
analysis from the reported symptoms.{issue_time} (explicit, relative, or now) and
center a 1-hour window around it (start = issue_time - 30m,
end = issue_time + 30m) for all log/metric queries.Equivalent via Cloud Logging (preferred when kubectl access is limited or for
historical events). Open it as a Logs Explorer deep link — URL-encode the
query and append the project and Step 0 time window:
https://console.cloud.google.com/logs/query;query={URL_ENCODED_QUERY};timeRange={start}%2F{end}?project={project_id}
(encode / as %2F, or use ;duration=PT1H for a rolling hour):
Interpret the Conditions table:
Ready: False / Ready: Unknown with reason KubeletNotReady /
NodeStatusUnknown ("Kubelet stopped posting node status") → kubelet or
runtime problem; continue to Step 2.MemoryPressure: True, DiskPressure: True, PIDPressure: True → resource
exhaustion; go to Step 4b.NetworkUnavailable: True → networking/CNI problem; go to Step 4d.Open these kubelet logs as a Logs Explorer deep link using the same
logs/query;query={URL_ENCODED_QUERY};timeRange=...?project=... pattern as Step 1.
Also review the node's serial-console logs (log_id("serialconsole.googleapis.com/serial_port_1_output")
or the resource.type="gce_instance" serial logs) for kernel TaskHung,
OOM-killer, or disk I/O errors that correlate with the kubelet failures.
| Kubelet / event signature | Likely root cause | Go to |
|---|---|---|
runtime is down, Container runtime not ready, errors on /run/containerd/containerd.sock (connection refused / DeadlineExceeded) | Container runtime (containerd) down or unresponsive | Step 4a |
Got sys oom event from cadvisor / kernel OOM-killer in serial logs | System (node-level) OOM killed critical processes | Step 4b |
PLEG is not healthy | PLEG stalled, usually node overload (CPU/disk) | Step 4c |
TaskHung for containerd/kubelet, high disk latency | Disk throttling / I/O starvation | Step 4b |
failed to ensure lease, leases.coordination.k8s.io ... namespace kube-node-lease ... terminating | kube-node-lease termination → NotReady flapping | Step 4f |
| Kubelet cannot reach API server, TLS/dial timeouts | Kubelet ↔ control-plane connectivity | Step 4d |
NetworkPluginNotReady, cni plugin not initialized, NetworkUnavailable | CNI plugin failure | Step 4d |
| Node-critical DaemonSet Pods (CNI, kube-proxy, metadata) blocked from admission | Admission webhook interference | Step 4e |
Only generic NodeNotReady, no other signature | Cause unclear — widen to Step 4d, then escalate | Escalation |
containerd) downConfirm the kubelet cannot talk to containerd (socket errors above). Check for
containerd restarts/crashes in serial logs. Remediation (propose, don't run):
recreate/repair the node (kubectl drain then let the node pool recreate it, or
gcloud container clusters upgrade/node auto-repair); if it recurs across nodes,
suspect a node image or custom DaemonSet interfering with containerd.
Cloud Monitoring metrics to inspect (read-only): kubernetes.io/node/memory/used_bytes,
kubernetes.io/node/cpu/core_usage_time, kubernetes.io/node/ephemeral_storage/used_bytes.
requests/limits,
reduce over-commit, or use larger machine types. Distinguish system OOM
(node-wide, kills kubelet/runtime) from cgroup OOM (single container).PLEG is not healthy almost always means the node is overloaded (CPU saturation,
disk latency, or too many Pods/containers per node) so the runtime can't relist
in time. Correlate with 4b metrics. Remediation: reduce node density, add
CPU/disk headroom, or spread workloads.
NetworkPluginNotReady): the CNI DaemonSet
(netd/calico/dataplane) is not running on the node → inspect those Pods'
logs/events.A misconfigured/failing validating or mutating webhook with a broad scope can block node-critical system Pods from being admitted, keeping the node NotReady.
Look for webhooks that intercept kube-system / node-critical objects with
failurePolicy: Fail. Remediation (propose): scope the webhook out of
kube-system/node-critical namespaces or set an appropriate namespaceSelector.
kube-node-lease termination flappingIf the node flaps NotReady with leases.coordination.k8s.io ... namespace kube-node-lease ... is being terminated, the kube-node-lease namespace was
deleted/terminating. Remediation (propose): do not delete the
kube-node-lease namespace; if terminating, identify the finalizer/actor holding
it and restore the namespace.
Escalate when either:
_Default bucket defaults to 30 days, so
incidents older than that are permanently deleted); orIn those cases, do all three:
_Default retention window and are permanently deleted")._Required bucket, default
400-day retention).This skill is derived from public Google Cloud documentation:
kube-node-lease / CNI / admission-webhook signatures and their remediations.resource.type="k8s_node", log_id("kubelet")) and log-bucket
retention (_Default 30 days, _Required 400 days).logs/query;query=... deep-link
format used above).