Chương 61: GKE Troubleshooting Runbook — Các Vấn Đề Phổ Biến & Giải Pháp
Tại sao điều này quan trọng
GKE là platform phức tạp với nhiều lớp tương tác: container runtime, networking, storage, scheduling, control plane. Khi vấn đề xảy ra trong production, khả năng diagnose nhanh chóng và chính xác là yếu tố then chốt để giảm MTTR (Mean Time To Recovery).
Các vấn đề phổ biến (Pod pending, Node NotReady, network connectivity failures) không phải là random — chúng tuân theo các pattern có thể dự đoán được. Pre-written runbooks giúp team tránh việc phải "khám phá lại bánh xe" khi incident xảy ra.
Chương này không tập trung vào lý thuyết mà vào quy trình thực thi: cho từng loại vấn đề, kiểm tra gì, tools nào để dùng, cây quyết định nào để follow.
Điều kiện tiên quyết
- Chương 39–43: Cloud Monitoring, Cloud Logging, GKE Observability, Production Debugging Methodology
- Chương 53–54: Production GKE Debugging Framework, Incident Response
- Kiến thức nền: Kubernetes object model, pod lifecycle, node concepts, networking, storage
- Công cụ: kubectl, gcloud CLI, Cloud Console
Mức độ sâu & cách sử dụng
Mức độ: 4/5 — Runbook này dành cho engineers, SREs, platform team members cần debug issues ngay trong production. Mỗi section là standalone quy trình có thể thực thi ngay, không cần đọc các sections khác trước.
Cách sử dụng:
- Khi incident xảy ra, xác định loại vấn đề (Pod pending? Node NotReady? Networking?)
- Nhảy tới section tương ứng
- Follow các bước debugging theo thứ tự, collect data
- Dùng cây quyết định để define root cause
- Apply fix hoặc escalate nếu cần
Các chủ đề con
Runbook này bao gồm 8 vấn đề phổ biến nhất trong production GKE:
Pod Creation Failures: Troubleshooting Checklist
- Khi Pods không thể được tạo hoặc rejected ngay
- Resource constraints, storage, CNI, webhooks, quotas
Scheduling Failures: Resolving Pending Pods
- Khi Pods ở status Pending (không được schedule)
- Insufficient resources, autoscaler blocking, capacity limits
Networking Issues: Connectivity Test Procedures
- Khi traffic giữa các Pods, services, hoặc external không work
- Connectivity tests, policies, packet capture, dataplane issues
Storage Issues: Volume Attachment Failures
- Khi PersistentVolumes không mount, attach timeouts, fsGroup issues
- PV provisioning, filesystem mismatches, disk attach limits
Control Plane Issues: API Server & etcd Health
- Khi API server latency cao, etcd bị overload
- Metrics collection, bottleneck analysis, performance tuning
Node Issues: NotReady Diagnosis
- Khi nodes report NotReady status
- Kubelet health, resource pressure, kernel issues
Workload Identity Failures: Token Exchange Debugging
- Khi Pods không thể authenticate với GCP APIs
- Token exchange, metadata server, quota issues
Autoscaling Issues: HPA/CA Troubleshooting
- Khi Pods không scale (HPA) hoặc nodes không được provision (CA)
- Metrics, autoscaler logs, scale-up blockage
Structured Debugging Methodology
Mỗi runbook section follow quy trình này:
1. Symptoms Recognition
Dấu hiệu nào cho biết vấn đề này?
2. Information Gathering
Collect ngay: logs, events, metrics nào?
3. Diagnostic Decision Tree
Cây quyết định tuần tự: kiểm tra A → kết quả là B? → đi tới C hoặc D?
4. Common Root Causes & Fixes
Với mỗi root cause, fix nào (immediate workaround vs permanent)?
5. Prevention & Monitoring
Để tránh incident này lần tiếp theo, cần cái gì?
6. Escalation Criteria
Khi nào phải escalate cho platform team / Google Support?
Tools & Commands Reference
Diagnostic Tools:
kubectl logs,kubectl describe— pod/node state inspectionkubectl events— Kubernetes events stream- Cloud Logging filters — historical analysis
- Cloud Monitoring dashboards — metrics correlation
gcloud compute instances describe— VM-level diagnostics- Connectivity Tests — GCP network diagnostics
tcpdump/ packet capture — network-level debugging
GCP Console Shortcuts:
- Cluster Dashboard → Workloads tab → check pod status
- Cluster Dashboard → Nodes tab → check node health
- Cloud Logging → Resource = gke_container → filter by pod/node
- Cloud Monitoring → Kubernetes → Pods → custom dashboards
Real-World Incident Context
Các runbook này ditulis dựa trên real incidents mà platform teams gặp phải ở scale production:
- Resource quota misconfiguration blocking critical system pods
- CNI plugin failures khi scaling > 500 nodes
- PersistentVolume mount timeouts với millions of files
- API server latency spikes từ watch cache exhaustion
- Workload Identity token exchange hitting rate limit
- Cluster autoscaler stuck do node pool at maximum size
Mỗi troubleshooting path đã được validated với real incident scenarios.
Quick Reference: When to Use Which Runbook
| Vấn đề | Triệu chứng | Section |
|---|---|---|
| Pod không tạo, rejected immediately | Error: <reason> | Pod Creation Failures |
| Pod pending, waiting | Pending status, Unschedulable event | Scheduling Failures |
| Pod running nhưng không reach được | Connection timeout, refused | Networking Issues |
| Pod CrashLoop, mount error | FailedMount, FileNotFound | Storage Issues |
| API slow, kubectl commands hang | High latency, timeouts | Control Plane Issues |
| Nodes marked NotReady | NotReady condition | Node NotReady Diagnosis |
| Pod fails with auth error | 403 Forbidden to GCP APIs | Workload Identity Failures |
| Pod replicas không scale up/down | HPA not scaling, CA not provisioning | Autoscaling Issues |
Notes on Automation & Escalation
Automated Detection: GKE Health Status dashboard trong Cloud Console tự động detect nhiều vấn đề này. Luôn check console health dashboard trước khi vào runbook chi tiết.
When to Escalate to Google Cloud Support:
- Node issues không resolve với reboot/auto-repair
- Control plane issues persist after scaling down workloads
- Storage issues related to infrastructure (disk quota, zone capacity)
- Unexplained network packet loss
- GCP service-level outages