Skip to content

Chương 61: GKE Troubleshooting Runbook — Các Vấn Đề Phổ Biến & Giải Pháp

Tại sao điều này quan trọng

GKE là platform phức tạp với nhiều lớp tương tác: container runtime, networking, storage, scheduling, control plane. Khi vấn đề xảy ra trong production, khả năng diagnose nhanh chóng và chính xác là yếu tố then chốt để giảm MTTR (Mean Time To Recovery).

Các vấn đề phổ biến (Pod pending, Node NotReady, network connectivity failures) không phải là random — chúng tuân theo các pattern có thể dự đoán được. Pre-written runbooks giúp team tránh việc phải "khám phá lại bánh xe" khi incident xảy ra.

Chương này không tập trung vào lý thuyết mà vào quy trình thực thi: cho từng loại vấn đề, kiểm tra gì, tools nào để dùng, cây quyết định nào để follow.

Điều kiện tiên quyết

  • Chương 39–43: Cloud Monitoring, Cloud Logging, GKE Observability, Production Debugging Methodology
  • Chương 53–54: Production GKE Debugging Framework, Incident Response
  • Kiến thức nền: Kubernetes object model, pod lifecycle, node concepts, networking, storage
  • Công cụ: kubectl, gcloud CLI, Cloud Console

Mức độ sâu & cách sử dụng

Mức độ: 4/5 — Runbook này dành cho engineers, SREs, platform team members cần debug issues ngay trong production. Mỗi section là standalone quy trình có thể thực thi ngay, không cần đọc các sections khác trước.

Cách sử dụng:

  1. Khi incident xảy ra, xác định loại vấn đề (Pod pending? Node NotReady? Networking?)
  2. Nhảy tới section tương ứng
  3. Follow các bước debugging theo thứ tự, collect data
  4. Dùng cây quyết định để define root cause
  5. Apply fix hoặc escalate nếu cần

Các chủ đề con

Runbook này bao gồm 8 vấn đề phổ biến nhất trong production GKE:

  1. Pod Creation Failures: Troubleshooting Checklist

    • Khi Pods không thể được tạo hoặc rejected ngay
    • Resource constraints, storage, CNI, webhooks, quotas
  2. Scheduling Failures: Resolving Pending Pods

    • Khi Pods ở status Pending (không được schedule)
    • Insufficient resources, autoscaler blocking, capacity limits
  3. Networking Issues: Connectivity Test Procedures

    • Khi traffic giữa các Pods, services, hoặc external không work
    • Connectivity tests, policies, packet capture, dataplane issues
  4. Storage Issues: Volume Attachment Failures

    • Khi PersistentVolumes không mount, attach timeouts, fsGroup issues
    • PV provisioning, filesystem mismatches, disk attach limits
  5. Control Plane Issues: API Server & etcd Health

    • Khi API server latency cao, etcd bị overload
    • Metrics collection, bottleneck analysis, performance tuning
  6. Node Issues: NotReady Diagnosis

    • Khi nodes report NotReady status
    • Kubelet health, resource pressure, kernel issues
  7. Workload Identity Failures: Token Exchange Debugging

    • Khi Pods không thể authenticate với GCP APIs
    • Token exchange, metadata server, quota issues
  8. Autoscaling Issues: HPA/CA Troubleshooting

    • Khi Pods không scale (HPA) hoặc nodes không được provision (CA)
    • Metrics, autoscaler logs, scale-up blockage

Structured Debugging Methodology

Mỗi runbook section follow quy trình này:

1. Symptoms Recognition

Dấu hiệu nào cho biết vấn đề này?

2. Information Gathering

Collect ngay: logs, events, metrics nào?

3. Diagnostic Decision Tree

Cây quyết định tuần tự: kiểm tra A → kết quả là B? → đi tới C hoặc D?

4. Common Root Causes & Fixes

Với mỗi root cause, fix nào (immediate workaround vs permanent)?

5. Prevention & Monitoring

Để tránh incident này lần tiếp theo, cần cái gì?

6. Escalation Criteria

Khi nào phải escalate cho platform team / Google Support?


Tools & Commands Reference

Diagnostic Tools:

  • kubectl logs, kubectl describe — pod/node state inspection
  • kubectl events — Kubernetes events stream
  • Cloud Logging filters — historical analysis
  • Cloud Monitoring dashboards — metrics correlation
  • gcloud compute instances describe — VM-level diagnostics
  • Connectivity Tests — GCP network diagnostics
  • tcpdump / packet capture — network-level debugging

GCP Console Shortcuts:

  • Cluster Dashboard → Workloads tab → check pod status
  • Cluster Dashboard → Nodes tab → check node health
  • Cloud Logging → Resource = gke_container → filter by pod/node
  • Cloud Monitoring → Kubernetes → Pods → custom dashboards

Real-World Incident Context

Các runbook này ditulis dựa trên real incidents mà platform teams gặp phải ở scale production:

  • Resource quota misconfiguration blocking critical system pods
  • CNI plugin failures khi scaling > 500 nodes
  • PersistentVolume mount timeouts với millions of files
  • API server latency spikes từ watch cache exhaustion
  • Workload Identity token exchange hitting rate limit
  • Cluster autoscaler stuck do node pool at maximum size

Mỗi troubleshooting path đã được validated với real incident scenarios.


Quick Reference: When to Use Which Runbook

Vấn đềTriệu chứngSection
Pod không tạo, rejected immediatelyError: <reason>Pod Creation Failures
Pod pending, waitingPending status, Unschedulable eventScheduling Failures
Pod running nhưng không reach đượcConnection timeout, refusedNetworking Issues
Pod CrashLoop, mount errorFailedMount, FileNotFoundStorage Issues
API slow, kubectl commands hangHigh latency, timeoutsControl Plane Issues
Nodes marked NotReadyNotReady conditionNode NotReady Diagnosis
Pod fails with auth error403 Forbidden to GCP APIsWorkload Identity Failures
Pod replicas không scale up/downHPA not scaling, CA not provisioningAutoscaling Issues

Notes on Automation & Escalation

Automated Detection: GKE Health Status dashboard trong Cloud Console tự động detect nhiều vấn đề này. Luôn check console health dashboard trước khi vào runbook chi tiết.

When to Escalate to Google Cloud Support:

  • Node issues không resolve với reboot/auto-repair
  • Control plane issues persist after scaling down workloads
  • Storage issues related to infrastructure (disk quota, zone capacity)
  • Unexplained network packet loss
  • GCP service-level outages

References