olares-doctor
Runtime diagnosis for Olares apps and the system via olares-cli — find the root cause when an app won't install or start, crashes, cannot pull an image, is `running` but unreachable, or is slow; includes doctor images and thirdleveldomain. Use for diagnosing catalog and dev app failures, not for authoring or editing charts.
How do I install this agent skill?
npx skills add https://github.com/beclab/olares --skill olares-doctorIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The olares-doctor skill provides diagnostic tools for the Olares platform using the olares-cli tool. It allows for inspection of application logs, cluster states, and resource usage, which can expose sensitive information. It also includes the ability to modify Kubernetes domain configurations and is susceptible to indirect prompt injection from untrusted log data.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
doctor (runtime diagnosis)
Shared front door: load
../olares-shared/SKILL.mdfor suite routing, active-profile selection, platform entry points, and the auth proceed/stop gate. Load its auth reference only when login, profile switching, token storage, or auth recovery is actually needed.
This skill is a thin diagnostic router over Market, Cluster, and Dashboard. Load the shared application-state model when interpreting lifecycle states, TTLs, serialized downloads, or running. Use olares-cli doctor <verb> --help for syntax.
When to use
- An install/upgrade is stuck or never reaches
running; an app won't start. - An app crashes / restarts repeatedly (CrashLoopBackOff, exit codes, config errors).
- An image won't pull (
ImagePullBackOff/ErrImagePull/ wrong arch), or you want to find unused local images. - An app is
runningbut its entrance is unreachable / errors / times out. - The system or an app is slow, or a GPU/resource binding is rejected (
node-pressure).
Both catalog apps (installed via market) and your own dev apps (deployed via chart) route runtime failures here. Once the root cause is found, the fix for a dev app you authored is usually a chart edit — hand back to ../olares-chart/SKILL.md.
Mental model:
doctoranswers "why is this broken and what do I do next?" Diagnosis is read-only by default; the only mutation is the explicitly approvedthirdleveldomain --force-deduperepair. The four-skill develop->deploy->debug combo ischart+market+olares-shared+doctor.
Symptom routing
| Symptom | Reference |
|---|---|
Install/upgrade stuck; never reaches running; sits in pending / downloading / installing / initializing; a fresh install ended in stopped | references/olares-doctor-app-stuck.md |
App crashes / restarts (CrashLoopBackOff, non-zero exit, CreateContainerConfigError, permission errors) | references/olares-doctor-app-crash.md |
Image won't pull (ImagePullBackOff / ErrImagePull / InvalidImageName / arch mismatch); finding unused local images | references/olares-doctor-image.md |
App is running but the entrance is unreachable / 5xx / times out / blank | references/olares-doctor-running-unhealthy.md |
System or app slow; resource pressure; GPU/compute binding rejected (node-pressure) | references/olares-doctor-resources.md |
A model that is configured but does not answer is diagnosed one layer up first: olares-router separates the gateway, its access control and the model application's own download/engine state from the pod-level failures here, and routes back when the cause is below the application.
First, rule out the normal queue. Before declaring an install stuck, check whether another app is
downloading— app-service runs one download at a time, so apendingrow is often just queuing (see the appstate reference and the app-stuck reference).
Verb index
| Command | Purpose | Read when triggered |
|---|---|---|
images | Full local image inventory annotated with workload references; unused candidates | image diagnosis |
thirdleveldomain | Audit duplicate/reserved third-level domains; optional repair | domain audit and repair |
How doctor gathers evidence (orchestration, not ownership)
- Lifecycle state/source comes from
olares-market. - Pods, events, logs, and workloads come from
olares-cluster. - Pressure and utilization come from
olares-dashboard. - Namespace discovery follows the shared platform model.
Correlate evidence by time and object ownership. A Market timeout is not failure; running proves only entrance TCP reachability; a fresh install can settle at stopped after scheduling failure without a *Failed lifecycle state.
Safety and escalation
- Diagnosis is read-only by default. Do not restart, delete, scale, cancel, prune, or edit a chart as an automatic diagnostic step.
thirdleveldomain --force-dedupemutates Application resources. Show the proposed changes and obtain explicit approval first.- An unused-image report is evidence, not authorization to remove images.
- System namespace evidence commonly requires admin visibility. On 403/404, use the app's own evidence and report the missing visibility; do not switch identities without approval.
- Stop when the app/user/namespace is ambiguous, required logs need a higher role, or the fix crosses into chart editing, lifecycle mutation, or host administration.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/beclab/olares/olares-doctor">View olares-doctor on skillZs</a>