Skip to content

Phương Pháp Debugging GKE Ở Production

Tại sao lại quan trọng?

Hầu hết GKE troubleshooting không phải là random kubectl exec. Nó là phương pháp nhập hệ thống để tìm hiểu cách Pod, Service, Node và control plane vận hành thực tế dưới lớp abstraction. Khi một ứng dụng production bị down:

  • Bạn cần hiểu ngay Pod status thực sự là gì (không phải tên), tại sao nó stuck ở state đó
  • Bạn cần biết cách correlate events từ nhiều layer (scheduler → kubelet → container runtime → kernel logs)
  • Bạn cần có mental model về failure cascade: nếu điều A không hoạt động, điều B, C, D sẽ fail theo cách nào

"Debugging" ở đây không phải là "tôi đọc logs và tìm xem lỗi là gì". Nó là reverse-engineering cách hệ thống vận hành từ observable behavior.

Scope & Prerequisites

Chapter này giả định bạn đã hiểu:

Structured Approach to Debugging

GKE failures có pattern. Hầu hết vấn đề rơi vào một số category:

1. Pod Lifecycle Failures

Pod không start (Pending), crash liên tục (CrashLoopBackOff), OOM, init container fail. Đây là level ứng dụng/kubelet.

2. Service Connectivity Issues

Pod chạy tốt nhưng không thể nói chuyện với nhau. DNS fail, NetworkPolicy block, iptables messed up. Đây là level networking.

3. Node-Level Problems

Node bị NotReady, disk pressure, CPU throttling, containerd crash. Kubelet không thể schedule pod.

4. Control Plane Failures

API latency cao, etcd slow, webhook timeout, scheduler fail to place pods.

5. Cross-Cutting Issues

Request gọi từ client → load balancer → node → pod, nhưng fail ở đâu không biết. Cần distributed tracing.

Mỗi category có debugging methodology riêng, tooling riêng, và cách tư duy riêng.

Debugging Mental Model

Khi troubleshoot, hỏi những câu này theo thứ tự:

  1. Pod Status là gì? (Pending, CrashLoopBackOff, Running, etc.) — Status không phải chỉ là string, nó có backing event history
  2. Pod description ra gì? (kubectl describe pod) — Events phải kể câu chuyện về lý do pod stuck
  3. Điểm fail ở layer nào? (Scheduling → admission → kubelet start → container runtime → kernel)
  4. Logs ở đâu? (Pod logs, kubelet logs, containerd logs, kernel logs)
  5. Có metrics support không? (CPU, memory, disk, network)
  6. Failure cascade như thế nào? (Nếu A fail → B cũng fail → C cũng fail → D fail theo thứ tự)

Debugging Toolchain Overview

ToolLayerUse Case
kubectl get/describe/logsPod metadataPod status, events, basic logs
kubectl execPod-insideDiagnose từ trong container (DNS, curl, netstat)
SSH → toolbox → tcpdumpNode networkPacket capture, MTU, conntrack
gcloud logging readHistorical logsKubelet, containerd, scheduler logs across time
Cloud MonitoringMetricsCPU, memory, disk, network time series
Cloud TraceDistributedRequest path through system
Gemini Cloud AssistAIAuto-diagnose common patterns

Chapters Included

Phần tiếp theo breakdown debugging methodology thành từng domain:

  • Pod Lifecycle Debugging — Pending pods, CrashLoopBackOff, OOMKilled, init container failures. Cách kubelet scheduler nó ra sao, cách container runtime start nó, cách ứng dụng crash.

  • Service Connectivity Debugging — DNS resolution, NetworkPolicy enforcement, service IP routing, port-forward bypassing. Cách iptables/eBPF route traffic, cách DNS client tìm resolver.

  • Node Debugging — NotReady nodes, disk pressure, CPU throttling, network issues. Cách kubelet health check node state, cách kernel report resource exhaustion.

  • Control Plane Debugging — API latency, etcd performance, webhook timeouts, scheduler failures. Cách etcd store state, cách API server route requests, cách webhook block admission.

  • Cross-Cutting Debugging — Request tracing từ client through load balancer through pod. Cách correlate access logs, app logs, traces thành unified story.

Real-World Debugging Scenario

Giả sử ứng dụng của bạn suddenly stop serving requests. Quy trình:

1. kubectl get pods → thấy pod status là Running
2. Tức là pod đang chạy, nhưng tại sao request fail?
3. kubectl exec → curl localhost:8080 → timeout
4. Tức là pod có lẽ hung hoặc stuck ở blocking operation
5. kubectl logs → không có logs mới trong 5 phút cuối
6. Tức là ứng dụng stuck, kubelet chưa restart nó
7. kubectl describe pod → liveness probe fail không?
   - Nếu liveness probe chạy, pod should restart. Nếu không restart → liveness probe không chạy
   - Nếu liveness probe không chạy → kubelet có problem
8. gcloud logging read resource.type=k8s_node ... → kubelet logs
   - Nếu kubelet log có "exec /bin/sh" fail → container runtime problem
   - Nếu kubelet log không ada apa-apa → kubelet stuck
9. SSH node → sudo systemctl status kubelet
   - Nếu kubelet service down → restart nó
   - Nếu kubelet service running nhưng memory 99% → node OOM
10. top → thấy process nào hogging memory

Ini adalah structured debugging: bukan random guess, tapi follow the chain of what could go wrong, check each layer methodically.

Key Principles

  1. Understand Pod Status String — Status bukan random; itu result dari kubelet state machine. Pending = scheduler pending. CrashLoopBackOff = container exited + restartPolicy = Always. Dll.

  2. Follow Event Historykubectl describe pod menampilkan events. Itu event log real. Kalau ada "Failed to pull image", itu bukan log message yang bisa hilang. Itu persistent event object di etcd.

  3. Logs Multi-Layered — Pod logs adalah stdout/stderr dari container. Kubelet logs adalah kubelet process logs. Containerd logs adalah container runtime logs. Ketiganya berbeda dan sometimes semua diperlukan untuk understand root cause.

  4. Metrics adalah Time Series — Kalau anda bilang "CPU high", check metrics time series, bukan snapshot. Kalau CPU spike ke 100% for 1 second then drop, itu berbeda dari sustained 100%.

  5. Correlate Across Time — Debugging bukan about one log line. Itu about building timeline: what happened at 14:23:45, then at 14:23:46, etc.

  6. Understand Failure Cascades — Kalau Node OOM, kernel OOM killer kill pod. Kubelet detect pod died, restart per restartPolicy. Kalau restart happen 5x dalam 5 menit, kubelet set pod status CrashLoopBackOff dan stop trying. Each step ada reason, understand semua.

Debugging Mindset

Ketika anda debugging production:

  • Don't guess — Ada hypothesis (contoh: "disk full"), verify dengan data (check df), follow evidence
  • Think in layers — Failure bisa di layer 1 (pod tidak start), layer 2 (pod start tapi container crash), layer 3 (container running tapi application hang)
  • Use primitive tools first — Sebelum deploy fancy APM, gunakan kubectl get, kubectl logs, kubectl describe dulu
  • Check prerequisites systematically — Sebelum bilang "networking down", pastikan Pod actually running dan container process not crashed
  • Build mental model of how each component works — Kalau anda tahu bagaimana scheduler work (filtering → scoring → binding), anda tahu why pod stuck Pending
  • Understand observable behavior — Pod status, events, logs, metrics adalah observables. Anda debugging adalah reverse-engineer internal state dari observables

Next Steps

Baca setiap chapter dalam order:

  1. Pod Lifecycle (memahami bagaimana pod start dan fail)
  2. Service Connectivity (memahami bagaimana traffic route)
  3. Node Debugging (memahami bagaimana node health)
  4. Control Plane (memahami bagaimana sistem decide placement)
  5. Cross-Cutting (memahami bagaimana request flow end-to-end)

Setiap chapter membangun mental model. Setelah selesai, anda bisa debug GKE production issues secara methodical, bukan trial-and-error.