npx skills add ...
npx skills add google/skills --skill agent-platform-deploy
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available models, check if a specific model is deployable (`gcloud ai model-garden models list-deployment-config`), query deployment cost, troubleshoot deployment errors (like quota limits), or undeploy/clean up endpoints. Also use when copying and deploying a 1P Tuned Model. Don't use for pure listing/discovery questions of the form "is X deployed?", "list my endpoints", or "which regions have models running?" — for those use `agent-platform-endpoint-management`. Don't use for public Vertex AI deployments (use the `vertex-deploy` skill) or for running model evaluations (use the `agent-platform-eval-flywheel` skill).
npx skills add google/skills --skill agent-platform-deploy
This skill provides instructions for deploying Open Models from Agent Platform Model Garden to endpoints, and subsequently undeploying them to clean up resources.
If you need to copy a 1P (First-Party) Tuned Model from a source project to a destination region or project and deploy it to a newly created endpoint, refer to the 1P Tuned Model Copy & Deployment Guide.
Before executing any commands on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:
list, describe, list-deployment-config)
deploy, undeploy-model)
undeploy-model, you MUST first verify that the endpoint and deployed
model exist; if describe or list returns a 404 or empty result, you
MUST halt and inform the user rather than attempting undeployment.delete)
Before deploying, ensure you have the correct project and region set. The
commands below use placeholder variables PROJECT_ID and LOCATION_ID.
Ensure you are authenticated:
You can list models available in Model Garden and check if they can be self-deployed.
To see what machine types and accelerators are supported for a specific model,
pass a MODEL_ID you obtained from the models list output above. Substitute
<PUBLISHER>/<FAMILY>@<VERSION-ID> below with the exact string from the catalog
output — the placeholder is deliberately not a real model ID:
[!NOTE] Some models, especially Hugging Face models, might require a Hugging Face Access Token for deployment.
[!TIP] Model Recommendation Instructions: Whenever you are about to name a specific model version in a response, do NOT recommend from memory. This applies in all of the following situations — not just direct deploy requests:
- The user asks to deploy a model without naming one.
- You are volunteering a next-step suggestion after a
list,describe, orundeployoperation (e.g. "Would you like me to deploy<model>to this endpoint?").- The user asks a general "what should I use?" / "what's a good model for X?" question.
- You are filling in a
MODEL_IDvalue in an example command you are showing the user (as opposed to a placeholder like<PUBLISHER>/<FAMILY>@<VERSION-ID>).New model versions ship frequently and older ones may be deprecated, so training-corpus knowledge of which models exist is unreliable. Follow this procedure:
- Clarify the use case if it isn't already clear from context (task type, quality vs. latency vs. cost priorities, hardware/quota constraints, license constraints). Skip if the user has already given enough signal.
- Query the live catalog with
gcloud ai model-garden models list. Narrow with--filterwhen appropriate (e.g.--filter="name~gemma",--filter="name~llama",--filter="name~qwen",--filter="name~deepseek"). Never name a specific model version to the user until you have seen it in the catalog output for this project.- Pick the latest generally-available version in the family that fits the use case. When multiple size variants exist, pick the one that matches the user's hardware/cost tolerance. Prefer a newer major version over an older one unless it is marked preview/experimental and the user explicitly asked for a stable option.
- Verify the exact model ID is deployable with
gcloud ai model-garden models list-deployment-config --model="<publisher>/<family>@<version>"before naming it in your response.- Cite the model ID verbatim in your recommendation, exactly as it appears in the catalog. Do not paraphrase to a family label ("Gemma", "Llama").
The
MODEL_IDvalues in the §3 examples below are intentionally non-substantive placeholders (<PUBLISHER>/<FAMILY>@<VERSION-ID>). Do NOT replace them with a remembered model name for a user-facing recommendation — always re-run steps 2-4 first, then cite the exact string from the catalog.
[!NOTE] Skip this section if the user is asking to deploy an open-weights model from Model Garden (Gemma, Llama, DeepSeek, Qwen, or any user-supplied weights) — i.e. anything served via
gcloud ai model-garden models deployonto a dedicated endpoint. These models have no per-region availability restriction; the Model Garden catalog is global. The real failure modes for an unusual region are (a) the requested accelerator/machine type isn't offered in that region, or (b) the project has no quota — both surface as a clean error at deploy time before any resources are provisioned (§3's cost-confirm gate catches them). Go straight to §3.Apply this section only if the user is asking to serve a first-party managed Gemini model (
google/gemini-*) or a fine-tuned Gemini LoRA adapter — both of which route through a publisher endpoint whose regional availability actually varies.
Before responding to any deploy request that names a specific region for a
first-party managed model (google/gemini-*) or a fine-tuned Gemini LoRA
adapter, you MUST verify the model is actually available in that region by
making a live API call. Do not rely on Google Search, training-corpus knowledge,
or publisher documentation for availability claims — regional availability
changes frequently and grounded text can be stale or wrong.
Probe only the exact model and region the user asked about. Do not probe other models as a "control" — you cannot infer anything about model A's availability from model B's status, because a different model may itself be unavailable in the reference region for unrelated reasons.
For first-party publisher models (google/*), probe with a real
:generateContent call using a minimal valid payload:
For fine-tuned Gemini LoRA models (deploying a user-tuned adapter on top of a
base Gemini model), probe the base model in the target region using the same
:generateContent call above with ${MODEL_ID} set to the base (e.g.
gemini-2.5-flash if the adapter was tuned on gemini-2.5-flash). The LoRA
adapter cannot serve in a region where its base model isn't available.
Interpret the probe result and act:
gcloud ai model-garden models list --filter="name~$MODEL_NAME" without --region). Do not silently switch
regions. Do not proceed to write deploy code or SDK initialization for the
unsupported region. Do not run additional "control" probes to double-check
the 404 — the target-region probe is authoritative.[!WARNING] Deploying models, especially large ones, consumes significant compute resources and incurs costs.
You MUST compute an hourly $ estimate for the requested
--machine-typebefore proposing a deploy. Try, in order, and fall through on any failure (tool unavailable, tool returnsstatus != "success", script exits non-zero, script rejects the machine type):a. If the
estimate_costtool is available AND returnsstatus == "success", use its result -- it returns live SKU-resolved pricing (machine + accelerator + total) fromCostEstimationServicerather than a hardcoded snapshot. On any other status (includingerror), fall through to (b).b. Otherwise, run
scripts/calculate_cost.py. The accelerator type and count are fixed per machine type in Model Garden and derived automatically. Example:If the script exits non-zero (unknown
--machine-type— a routine state for machines in the Model Garden catalog but not yet in the price snapshot, e.g. A4/B200 today), fall through to (c). Do NOT invent a number.c. Fall back to Agent Platform prediction pricing if the tool is unavailable AND the script does not know the requested machine type. Read the accelerator + hourly rate directly off that page and cite the URL in the estimate you present to the user.
You MUST present this cost estimation to the user and warn them that this is the list price, which may differ from their actual bill due to potential discounts, reservations, or non-
us-central1regions.You MUST ALWAYS request explicit confirmation from the user agreeing to the estimated cost before executing any
deploycommand.
To deploy a model, use the deploy command. It is highly recommended to use the
--asynchronous flag for long-running deployments, and then poll the status if
necessary.
Here is a typical bash script to deploy a model. You can run this block directly.
To deploy a model using custom weights, you can use the exact same deploy
command. Instead of providing the model garden model ID, provide the Google
Cloud Storage (GCS) URI to your custom weights folder in the --model flag.
When you deploy a model asynchronously using the --asynchronous flag, the
deploy command will return an operation ID. You can use this ID to check the
ongoing status of the deployment.
[!NOTE] As an agent, you can also offer to check the status of a deployment for the user if they provide an operation ID or if they just initiated the deployment with you.
Alternatively, you can list your endpoints to see if it shows up and check the Cloud Console under the "Online prediction" tab.
Note: Large models (roughly 20B+ parameters) may take 15-20 minutes to fully deploy and start serving.
If the model is successfully deployed, verify by making a prediction call to
test. Because Model Garden models are often deployed to Dedicated Endpoints, you
shouldn't use gcloud ai endpoints predict. Instead, you must fetch the
endpoint's dedicated DNS name and send a curl request.
[!TIP] Ask the user to try using their own prompt to see the results. Otherwise use the default.
Use the following script:
To stop incurring charges, you must undeploy the model from the endpoint. This is a multi-step process if you don't already have the exact endpoint and deployed model IDs.
Here is a bash script demonstrating how to find the IDs and undeploy the model.
[!WARNING] Failing to undeploy a model will result in continuous charges for the allocated compute resources, even if you are not sending prediction requests. Always clean up after testing.
If your deployment fails (or stays in an error state) due to QUOTA_EXCEEDED or
RESOURCE_EXHAUSTED errors, the specific hardware requested (e.g., NVIDIA_L4
or g2-standard-24) is either not available in your chosen region or exceeds
your project's quota limits.
Solution: Look closely at the error message returned. It will often
recommend an alternative region or machine type that currently has availability.
Ask the user for confirmation to retry the deployment using the suggested
--region or --machine-type parameters.
[!WARNING] If the alternative suggestions involve changing the machine type or accelerator, you MUST recalculate the estimated cost by re-running
scripts/calculate_cost.pywith the new params (see §3), warn the user about list prices versus actual billing, and get their explicit confirmation for the new cost before retrying the deployment.
gcloud ai model-garden models list-deployment-config \
--model="<PUBLISHER>/<FAMILY>@<VERSION-ID>"curl -sS -o /dev/null -w "%{http_code}\n" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://${LOCATION_ID}-aiplatform.googleapis.com/v1/projects/${PROJECT_ID}/locations/${LOCATION_ID}/publishers/google/${MODEL_ID}:generateContent" \
-d "{\"contents\":{\"role\":\"user\",\"parts\":{\"text\":\"${PROBE_TEXT:-hi}\"}}}"