Skill-Details
k8s-debug
Useful Kubernetes operations specialty, but narrower than general DevOps.
Vor Nutzung prüfen
Die automatische Prüfung bewertet Relevanz, nicht Sicherheit oder Empfehlung. Lies vor der Nutzung die Quellanweisungen.
SKILL.md
Dieser Auszug wurde bei der Prüfung gespeichert. Die externe Quelle enthält die vollständige und aktuelle Version.
--- name: k8s-debug description: Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl. --- # Kubernetes Debugging Skill ## Overview Systematic toolkit for debugging Kubernetes clusters, workloads, networking, and storage with a deterministic, safety-first workflow. ## Trigger Phrases Use this skill when requests resemble: - "My pod is in `CrashLoopBackOff`; help me find the root cause." - "Service DNS works in one pod but not another." - "Deployment rollout is stuck." - "Pods are `Pending` and not scheduling." - "Cluster health looks degraded after a change." - "PVC is pending and pods cannot mount storage." ## Prerequisites Run from the skill directory (`devops-skills-plugin/skills/k8s-debug`) so relative script paths work as written. ### Required - `kubectl` installed and configured. - An active cluster context. - Read access to namespaces, pods, events, services, and nodes. Quick preflight: ```bash kubectl config current-context kubectl auth can-i get pods -A kubectl auth can-i get events -A kubectl get ns ``` ### Optional but Recommended - `jq` for more precise filtering in `./scripts/cluster_health.sh`. - Metrics API (`metrics-server`) for `kubectl top`. - In-container debug tools (`nslookup`, `getent`, `curl`, `wget`, `ip`) for deep network tests. Fallback behavior: - If optional tools are missing, scripts continue and print warnings with reduced output. - If `kubectl top` is unavailable, continue with `kubectl describe` and events. ## When to Use This Skill Use this skill for: - Pod failures (CrashLoopBackOff, ImagePullBackOff, Pending, OOMKilled) - Service connectivity or DNS resolution issues - Network policy or ingress problems - Volume and storage mount failures - Deployment rollout issues - Cluster health or performance degradation - Resource exhaustion (CPU/memory) - Configuration problems (ConfigMaps, Secrets, RBAC) ## Safety Rules for Disruptive Commands Default mode is read-only diagnosis first. Only execute disruptive commands after confirming blast radius and rollback. Commands requiring explicit confirmation: - `kubectl delete pod ... --force --grace-period=0` - `kubectl drain ...` - `kubectl rollout restart ...` - `kubectl rollout undo ...` - `kubectl debug ... --copy-to=...` Before disruptive actions: ```bash # Snapshot current state for rollback and incident notes kubectl get deploy,rs,pod,svc -n <namespace> -o wide kubectl get pod <pod-name> -n <namespace> -o yaml > before-<pod-name>.yaml kubectl get events -n <namespace> --sort-by='.lastTimestamp' > before-events.txt ``` ## Reference Navigation Map Load only the section needed for the observed symptom. | Symptom / Need | Open | Start section | | --- | --- | --- | | You need an end-to-end diagnosis path | `./references/troubleshooting_workflow.md` | `General Debugging Workflow` | | Pod state is `Pending`, `CrashLoopBackOff`, or `ImagePullBackOff` | `./references/troubleshooting_workflow.md` | `Pod Lifecycle Troubleshooting` | | Service reachability or DNS failure | `./references/troubleshooting_workflow.md` | `Network Troubleshooting Workflow` | | Node pressure or performance regression | `./references/troubleshooting_workflow.md` | `Resource and Performance Workflow` | | PVC / PV / storage class issues | `./references/troubleshooting_workflow.md` | `Storage Troubleshooting Workflow` | | Quick symptom-to-fix lookup | `./references/common_issues.md` | matching issue heading | | Post-mortem fix options for known issues | `./references/common_issues.md` | `Solutions` sections | ## Scripts Overview | Script | Purpose | Required args | Optional args | Output | Fallback behavior | | --- | --- | --- | --- | --- | --- | | `./scripts/cluster_health.sh` | Cluster-wide health snapshot (nodes, workloads, events, common failure states) | None | `--strict`, `K8S_REQUEST_TIMEOUT` env var | Sectioned report to stdout | Continues on check failures, tracks them in summary and exit cVollständige Quelle auf GitHub lesen (öffnet externe Seite)