npx skills add ...
npx skills add temporalio/skill-temporal-cloud --skill temporal-cloud
Fix Temporal Cloud connection, auth, and config problems. Use when users hit login failures, can't connect to Cloud, get x509/TLS errors, have namespace or endpoint mismatches, paste broken SDK connection snippets, are confused about which endpoint to use, see "no pollers" or RESOURCE_EXHAUSTED, struggle with PrivateLink/PSC, or need help setting up a new namespace. Also use for HA namespace failover and DNS issues. Not for worker performance tuning or scaling.
npx skills add temporalio/skill-temporal-cloud --skill temporal-cloud
Help users diagnose and resolve Temporal Cloud connectivity, authentication, and configuration issues using tcld and temporal CLI.
Cloud issues are frustrating because they sit at the intersection of configuration, networking, authentication, and Temporal-specific code. Most problems fall into predictable patterns. This skill provides systematic diagnosis to quickly identify root causes and prescribe fixes.
References:
references/cloud-troubleshooting-reference.md for full CLI command reference and error codesreferences/common-scenarios.md for step-by-step setup walkthroughsOut of scope: Worker performance tuning, scaling, metrics interpretation, SDK-specific config, deployment patterns. Those topics are covered by separate worker-focused skills.
| Category | Key Symptoms | First Check |
|---|---|---|
| tcld Login | login failed, token refresh failed, wrong account | tcld account get |
| Connection/Auth | can't connect, access denied, handshake failures | Endpoint format + DNS + port connectivity |
| Ambiguous Runtime Errors | context deadline exceeded, workflow is busy | Identify the operation and layer first |
| mTLS/Certs | x509 errors, unknown authority, expired | openssl x509 -enddate |
| Namespace | namespace not found, SNI mismatch | Namespace name format |
| HA / Failover | Failover not working, wrong region, DNS stale | DNS CNAME resolution |
| Worker | Tasks not picked up, stale connections | temporal task-queue describe |
| Private Connectivity | PrivateLink/PSC errors | VPC endpoint status |
| Rate Limiting | RESOURCE_EXHAUSTED | APS limits |
Ask the user:
For SDK/client snippet reviews:
HostPort / address are you using?temporal CLI, or tcld?For tcld issues:
tcld account get?For connection issues:
HostPort?For ambiguous runtime errors:
For certificate issues:
For worker issues:
temporal task-queue describe show?Use the appropriate decision tree based on category (see below).
Give specific commands to resolve the issue, with verification steps.
Always include a confidence score for the proposed diagnosis or fix:
Confidence: 9-10/10 when the symptom, operation, and confirming signals line up cleanlyConfidence: 6-8/10 when the evidence is good but one plausible alternative remainsConfidence: 1-5/10 when the issue is still ambiguous and the "fix" is really the next discriminating checkIf the problem is ambiguous, say so explicitly and keep the recommendation scoped to the next check rather than presenting a speculative root cause as settled.
Docs: Environment configuration - SDK connection options
Endpoint check before network debugging:
| Use case | Recommended endpoint | Notes |
|---|---|---|
| Workers & clients (all auth) | <namespace>.<account>.tmprl.cloud:7233 | Namespace Endpoint - works for both mTLS and API key auth. Recommended for all namespaces. |
| Multi-region HA (advanced) | <region>.<cloud_provider>.api.temporal.io:7233 | Regional Endpoint - only needed for advanced HA routing. See namespace access docs. |
| tcld / Cloud Ops API | saas-api.tmprl.cloud | Control plane |
Exception: Namespaces using Flexible Auth (pre-release) cannot use Namespace Endpoints yet.
Do not assume these are pure connectivity failures. Classify them by operation first.
| Error text | Common interpretations | First discriminator |
|---|---|---|
context deadline exceeded | wrong endpoint, network timeout, oversized payload, blocked execution path, client-side timeout | Where in the flow does it occur? |
workflow is busy / RESOURCE_EXHAUSTED: Workflow is busy | operation-level contention, workload pressure, confusing user-facing error semantics | Which operation returned it? |
no pollers | no connected workers, workers present but misconfigured, stale/misleading metrics | Does temporal task-queue describe show pollers? |
Use this decision sequence:
If the operation and surrounding signals still do not make the error interpretable, label it as ambiguous and gather more context before prescribing a fix.
When responding, attach a confidence score from 1-10 to the proposed diagnosis or next step. Ambiguous cases should carry a low-confidence score and a narrow next check rather than a broad claimed fix.
When the user pastes SDK config, validate the config itself before suggesting lower-level networking checks.
Review in this order:
HostPort: should be Namespace Endpoint (<ns>.<acct>.tmprl.cloud:7233) for most cases<namespace>.<account-id>)tls.Config{} is normal for API key auth; client cert/key required for mTLSTEMPORAL_ADDRESS, TEMPORAL_NAMESPACE, TEMPORAL_API_KEY, TEMPORAL_TLS_CLIENT_CERT_PATH, TEMPORAL_TLS_CLIENT_KEY_PATHCommon snippet diagnoses:
*.api.temporal.io) when Namespace Endpoint would work → simplify to <ns>.<acct>.tmprl.cloud:7233HostPort + Cloud namespace/auth → missing explicit Cloud endpointIf the snippet is wrong, fix that first. Do not lead with DNS/TLS debugging until the endpoint and namespace are plausible.
This skill diagnoses Cloud connectivity issues for workers. Worker performance tuning, scaling, and deployment patterns are out of scope.
Docs: Environment configuration - SDK connection setup
Scope clarification:
| Issue Type | In Scope? |
|---|---|
| tcld login, certs, namespace, private connectivity | Yes |
| Worker scaling, metrics, tuning, deployment | No |
| "Workers not picking up tasks" | Yes - diagnose, hand off if not a Cloud issue |
Docs: HA namespace connectivity
HA (multi-region) namespaces use a hierarchical DNS structure:
<ns>.<acct>.tmprl.cloud (CNAME to active region)<region>.region.tmprl.cloudWorker placement for HA:
See references/common-scenarios.md for step-by-step walkthroughs:
| Pitfall | Why It Happens | Fix |
|---|---|---|
| Wrong namespace format | Using short name instead of name.account-id | Use full namespace from tcld namespace list |
| Self-signed certs | Trying to use self-signed without CA | Generate CA first, sign certs with it, upload CA |
| Regional endpoint when unnecessary | Using *.api.temporal.io when Namespace Endpoint works | Switch to <namespace>.<account>.tmprl.cloud:7233 |
| Old endpoint docs | Following stale examples from before Namespace Endpoints were universal | Use Namespace Endpoint: <ns>.<acct>.tmprl.cloud:7233 |
| Expired certs | Not monitoring expiry | Set up alerts, rotate before expiry |
| tcld wrong account | Logged into different org | Use tcld account get, then verify the namespace in tcld namespace list |
| Stale tcld login | Cached auth state is no longer valid | tcld logout && tcld login |
| DNS caching | K8s pods caching old DNS | Restart pods after endpoint changes |
| Missing port | Firewall blocks 7233 | Ensure egress allowed on port 7233 |
<ns>.<acct>.tmprl.cloud:7233 works for both mTLS and API key auth