npx skills add ...
npx skills add google/skills --skill gke-manifest-generation
Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters. Use when creating or modifying GKE deployment manifests, configuring container security contexts, setting CPU/memory resource limits, defining readiness/liveness/startup probes, mounting secrets and volumes, configuring GKE Gateway API routes, targeting Spot VMs, or deploying AI model inference workloads (vLLM, TGI, Gemma). Don't use for live cluster operations, pod troubleshooting (use gke-workload-troubleshooting), or cluster infrastructure provisioning (use gke-cluster-creation).
npx skills add google/skills --skill gke-manifest-generation
This skill provides guidelines, tooling integration, and templates to translate natural language descriptions or application code changes into secure, compliant, and cost-effective Kubernetes YAML manifests optimized for both GKE Autopilot and GKE Standard clusters.
When generating or updating YAML manifests, you must strictly adhere to the following rules:
namespace: {namespace} explicitly
in the metadata of every resource (Deployments, Services, ConfigMaps,
Secrets, PVCs, Roles, bindings). Map it to the namespace configured in your
active SETTINGS.md. Never omit the namespace.default
ServiceAccount. Always create and reference a dedicated ServiceAccount
(e.g., devteam-agent-sa) for each microservice.Resources Requests & Limits: Always specify CPU and Memory requests and limits for all containers.
Density Defaults: For stateless apps or sidecars on GKE Standard,
default to conservative requests (e.g., requests.cpu: "100m" or "200m",
requests.memory: "256Mi" or "512Mi") with burstable limits. Use a
reasonable overcommit ratio for limits (e.g., 2x to 4x requests, like
limits.cpu: "400m" to "800m", and limits.memory: "512Mi" to "1Gi").
Avoid excessive overcommit limits (like limits.cpu: "4" for a 100m
request) to prevent severe CPU throttling and latency degradation under
heavy scheduling load, particularly in environments without guaranteed node
shares.
Spot VMs for Staging/Dev: For non-production workloads (e.g., namespaces
containing -test, -dev, or -staging), or if the user requests cost
optimization, automatically target GKE Spot VMs. This requires injecting
both the nodeSelector targeting Spot VMs AND the corresponding toleration
to tolerate the Spot VM taint:
(On GKE Standard, this assumes a Spot node pool is configured).
securityContext at the Pod level
(and container level if overriding) to run as a non-root user (e.g.,
runAsNonRoot: true, runAsUser: 10000, runAsGroup: 10000, fsGroup: 10000). This is strictly enforced on GKE Autopilot and is a critical
security baseline for GKE Standard.allowPrivilegeEscalation: false and
seccompProfile: {type: RuntimeDefault}.readOnlyRootFilesystem: true to prevent
modifications to the container image filesystem.
readOnlyRootFilesystem is enabled,
mount a local emptyDir volume to /tmp or /var/run/ to allow
applications (like Java/Nginx) to write temp files without crashing.volumes spec with defaultMode: 0400) instead of
mapping them as environment variables, unless the application framework
exclusively supports env-var based configuration. This prevents secrets
leaking into application logs.Liveness & Readiness Probes: Every Deployment container must define both
livenessProbe and readinessProbe.
httpGet probes.tcpSocket probes.exec probes (e.g.,
exec.command: ["redis-cli", "ping"]).Startup Probes for Slow-Starting Apps: For applications with slow boot
times (e.g., Java spring boot, complex Python scripts, LLM model servers),
you must also define a startupProbe. When a startupProbe is defined,
the liveness and readiness probes are disabled until it succeeds, preventing
Kubernetes from prematurely killing the pod during startup:
Sensible Defaults: Set initialDelaySeconds: 5 to 15 depending on
startup time (e.g., Java requires a longer delay than Go/Nginx).
type: ClusterIP. Never use type: LoadBalancer or NodePort unless the workload
is explicitly intended to be publicly accessible from the internet.name: http-web or name: grpc-api) to enable
automatic protocol discovery, tracing, and Web App routing.Gateway and HTTPRoute resources) over legacy Ingress
objects to enable advanced L7 routing and security features (e.g., Cloud
Armor).ConfigMap or Secret to
an application directory containing other files (like Nginx public
directories), always use subPath to overlay only the specific file.
Caveat: Note that containers using subPath volume mounts do not receive
automatic configuration updates if the underlying ConfigMap or Secret is
modified; pods must be restarted manually to pick up changes.standard-rwo
(default balanced PD) or premium-rwo (SSD PD).standard (default PD) or premium
(SSD PD) if standard-rwo/premium-rwo are not configured.premium-rwo or premium)
only when the prompt explicitly requests high IOPS, low latency, or
database storage.podAntiAffinity
or topologySpreadConstraints with topologyKey: "kubernetes.io/hostname"
to distribute pods across GKE nodes and availability zones.PodDisruptionBudget to guarantee minimum replica availability during
voluntary GKE node upgrades and maintenance cycles.name). You must keep the name key stable when modifying
properties of an existing list item. Renaming the name key will cause SSA
to create a brand new entry and leave the old entry intact (orphaned) rather
than modifying it.For model serving workloads, prioritize using optimized tooling like GKE Inference Quickstart if available. If generating manually:
nvidia.com/gpu in both requests and limits.nodeSelector or node affinity targeting the desired GKE
accelerator tag (e.g., cloud.google.com/gke-accelerator: nvidia-l4)./dev/shm) for inter-process
communications. Always declare and mount an emptyDir volume with
medium: Memory to /dev/shm.csi.storage.gke.io) as readOnly: true for efficient
cold-starts.When generating manifests, you should leverage the following tooling to reduce hallucinations and optimize configurations:
Inference Workloads (GKE Inference Quickstart CLI):
Make sure you have the Google Cloud SDK installed.
For all AI/LLM inference workloads (e.g. model serving), you must
prioritize using the gcloud CLI GKE Inference Quickstart command to
generate the optimized manifests instead of writing them manually:
Constraint: You must include all resources returned by this command (Deployments, Services, PodMonitoring, etc.) without filtering.
Grounding in Official Documentation (Developer Knowledge API):
answer_query: Use this to ask direct questions (e.g., "How to
configure GCS Fuse CSI driver in GKE"). This is the preferred tool
for general queries.search_documents: Use this to search for relevant GKE guides
or examples when you don't have a specific question.get_document: Use this to fetch full document contents when
you have a specific document ID.For detailed, production-ready manifest templates, consult the following reference guides:
/dev/shm
shared memory boost, and startup probes.Gateway and HTTPRoute resources).