npx skills add ...
npx skills add beclab/olares --skill olares-doctor
Runtime diagnosis for Olares apps and the system via olares-cli — find the root cause when an app won't install or start, crashes, cannot pull an image, is `running` but unreachable, or is slow; includes doctor images and thirdleveldomain. Use for diagnosing catalog and dev app failures, not for authoring or editing charts.
npx skills add beclab/olares --skill olares-doctor
Shared front door: load
../olares-shared/SKILL.mdfor suite routing, active-profile selection, platform entry points, and the auth proceed/stop gate. Load its auth reference only when login, profile switching, token storage, or auth recovery is actually needed.
This skill is a thin diagnostic router over Market, Cluster, and Dashboard. Load the shared application-state model when interpreting lifecycle states, TTLs, serialized downloads, or running. Use olares-cli doctor <verb> --help for syntax.
running; an app won't start.ImagePullBackOff / ErrImagePull / wrong arch), or you want to find unused local images.running but its entrance is unreachable / errors / times out.node-pressure).Both catalog apps (installed via market) and your own dev apps (deployed via chart) route runtime failures here. Once the root cause is found, the fix for a dev app you authored is usually a chart edit — hand back to ../olares-chart/SKILL.md.
Mental model:
doctoranswers "why is this broken and what do I do next?" Diagnosis is read-only by default; the only mutation is the explicitly approvedthirdleveldomain --force-deduperepair. The four-skill develop->deploy->debug combo ischart+market+olares-shared+doctor.
| Symptom | Reference |
|---|---|
Install/upgrade stuck; never reaches running; sits in pending / downloading / installing / initializing; a fresh install ended in stopped | references/olares-doctor-app-stuck.md |
App crashes / restarts (CrashLoopBackOff, non-zero exit, CreateContainerConfigError, permission errors) | references/olares-doctor-app-crash.md |
Image won't pull (ImagePullBackOff / ErrImagePull / InvalidImageName / arch mismatch); finding unused local images | references/olares-doctor-image.md |
App is running but the entrance is unreachable / 5xx / times out / blank | references/olares-doctor-running-unhealthy.md |
System or app slow; resource pressure; GPU/compute binding rejected (node-pressure) | references/olares-doctor-resources.md |
A model that is configured but does not answer is diagnosed one layer up first: olares-router separates the gateway, its access control and the model application's own download/engine state from the pod-level failures here, and routes back when the cause is below the application.
First, rule out the normal queue. Before declaring an install stuck, check whether another app is
downloading— app-service runs one download at a time, so apendingrow is often just queuing (see the appstate reference and the app-stuck reference).
| Command | Purpose | Read when triggered |
|---|---|---|
images | Full local image inventory annotated with workload references; unused candidates | image diagnosis |
thirdleveldomain | Audit duplicate/reserved third-level domains; optional repair | domain audit and repair |
olares-market.olares-cluster.olares-dashboard.Correlate evidence by time and object ownership. A Market timeout is not failure; running proves only entrance TCP reachability; a fresh install can settle at stopped after scheduling failure without a *Failed lifecycle state.
thirdleveldomain --force-dedupe mutates Application resources. Show the proposed changes and obtain explicit approval first.