Skip to content

SỔ TAY KỸ THUẬT GCP CẤP ĐỘ SẢN XUẤT

Hệ Thống Toàn Diện cho Platform Engineers & Staff/Principal Cloud Architects


PHẦN I: NỀN TẢNG KIẾN TRÚC & TỔNG QUAN GCP


Chương 1: GCP Resource Hierarchy & Tổ Chức Tài Nguyên

Tại sao quan trọng: Mọi quyết định IAM, billing, network boundary, org policy phụ thuộc vào hiểu được resource hierarchy. Sai ở tầng này → blast radius tối đa ảnh hưởng toàn tổ chức.

Chapter 1 Full Index & Learning Paths

Các chủ đề con:

  1. Resource Hierarchy Fundamentals - Organization → Folder → Project → Resource: phân cấp, inheritance, override mechanics

  2. Resource Manager API - Programmatic resource management, propagation delay, consistency model, eventual consistency handling

  3. Project Naming & Automation - Project ID constraints, immutability, soft-delete windows, naming automation patterns

  4. IAM Policy Propagation - Three-layer propagation, eventual consistency, caching behavior, testing strategies

  5. Quota Management - Quota types (allocation/rate/concurrent), project vs organization-level, exhaustion scenarios

  6. Labels, Tags & Organization - Labels vs Tags vs Network Tags: usage cho billing, firewall, IAM conditions, cost allocation

  7. Cloud Asset Inventory - Query resource state across hierarchy, drift detection, compliance auditing

  8. Resource Protection - Locking, deletion protection, soft-delete recovery, backup strategies

  9. Shared VPC Model - Host project vs service projects, centralized network management, cross-project connectivity

  10. Service Account Scoping - Cross-project access patterns, keys vs tokens, workload identity, impersonation chains

  11. Billing Hierarchy - Cost attribution, billing account structure, chargeback models, budget alerts

  12. Organization Policies - Constraint framework, managed/custom constraints, conditional policies, CEL expressions


Chương 2: GCP Physical Network Architecture — Jupiter Fabric & Andromeda

Tại sao quan trọng: GCP networking khác hoàn toàn so với on-prem và AWS. Jupiter spine-leaf fabric, Andromeda SDN, global routing — hiểu cơ chế này giải thích latency, failover behavior, bandwidth allocation.

Chapter 2 Full Index & Learning Paths

Các chủ đề con:

  1. Andromeda: GCP Software-Defined Networking Stack - Control plane vs data plane, 5-step packet processing pipeline, VPC logical overlay, production patterns, anti-patterns

  2. Jupiter Fabric: Spine-Leaf Topology - Physical datacenter topology, per-server bandwidth, ECMP routing, oversubscription implications, zone placement

  3. Google Points of Presence (PoP) - Edge node hierarchy, traffic entry points, PoP failover mechanisms, DDoS scrubbing, anycast routing

  4. GCP Global Backbone: Premium vs Standard Tier - User-centric vs region-centric routing, private fiber backbone, SLA differences, cost tradeoffs

  5. Latency SLA & Fiber Path Engineering - Fiber infrastructure, latency components, multi-path redundancy, inter-region latencies, propagation delays

  6. Anycast Routing with Global Load Balancer - BGP anycast mechanism, automatic geo-routing, single IP multiple locations, failover transparency

  7. Cold Potato vs Hot Potato Routing Strategies - Egress point optimization, cold potato (backbone) vs hot potato (internet), strategic routing decisions

  8. Network Service Tiers: Practical Datapath Implications - Premium vs Standard tier queuing, SLA achievement mechanics, health checking differences

  9. Bandwidth Allocation & Egress Pricing Architecture - Per-zone capacity, bandwidth quotas, egress pricing model, burst allowance mechanics

  10. Regional vs Global Services: Data Sovereignty - Data residency requirements, GDPR/CCPA/HIPAA compliance, regional constraints, multi-region architectures

  11. Traffic Engineering & Multi-path Load Balancing - ECMP routing, capacity planning, failure scenarios, multi-path resilience, cascade failure prevention


Chương 3: GCP VPC Model — Kiến Trúc Mạng Ảo Toàn Cầu

Tại sao quan trọng: VPC là nền tảng của mọi thứ trong GCP. Hiểu cấu trúc global-regional, subnet design, routing primitives là điều kiện bắt buộc.

Chapter 3 Full Index & Learning Paths

Các chủ đề con:

  1. VPC là Global Resource, Subnet là Regional - VPC scope vs subnet scope, implications cho multi-region, tại sao GCP khác AWS/Azure

  2. Auto-mode vs Custom-mode VPC - Auto-mode limitations (10.128.0.0/9), custom-mode flexibility, tại sao production luôn custom, migration strategies

  3. Subnet Design & CIDR Planning - Primary vs secondary ranges, GKE Pod CIDR allocation, IP address management at scale, overlap constraints

  4. Alias IP Ranges & GKE Pods - VPC-native pod routing (không NAT), anti-spoofing checks, container networking patterns, firewall interactions

  5. Static Routes & Next Hops - Subnet routes, custom static routes, next hop types (VMs, ILBs, VPN), route conflict resolution

  6. Dynamic Routes & Cloud Router - BGP sessions, route learning/advertisement, regional vs global mode, on-premises connectivity

  7. System-generated Routes - Default route, subnet routes, special paths (GFE, IAP, Serverless), reserved ranges

  8. Firewall Rules Fundamentals - Stateful inspection, priority 0-65535, ingress/egress asymmetry, connection tracking limits

  9. Network Tags vs Service Accounts - Tags vs SAs for firewall targeting, decision matrix, multi-tier patterns, IAM integration

  10. Hierarchical Firewall Policies - Organization → folder → project evaluation order, allow/deny semantics, exceptions, multi-org scenarios

  11. Cloud NGFW & L7 Inspection - FQDN filtering, TLS interception, IDS/IPS, threat intelligence, latency overhead, throughput ceilings

  12. VPC Peering Deep Dive - No-transitivity principle, mesh topology, hub-and-spoke routing, DNS resolution, cross-project patterns

  13. Shared VPC & Centralized Management - Host/service projects, subnet sharing, IAM role separation, multi-tenancy isolation, cost attribution

  14. Private Google Access - 199.36.153.x/30 routing, Google APIs access without internet, private vs restricted endpoints

  15. VPC Flow Logs Analysis - Sampling mechanics, metadata fields, BigQuery export, cost analysis, troubleshooting patterns

  16. Network Intelligence Center - Topology visualization, connectivity tests, performance insights, firewall analysis

  17. VPC Service Controls - Service perimeters, access levels, ingress/egress rules, data exfiltration prevention, compliance


Chương 4: Cloud DNS Architecture & Production Patterns

Tại sao quan trọng: DNS là attack surface ẩn. Misconfiguration dẫn đến outages và data exfiltration. Cloud DNS for GKE là mandatory cho Autopilot.

Chapter 4 Full Index & Learning Paths

Các chủ đề con:

  1. Managed Zones: Public vs Private - Public/private zone fundamentals, zone characteristics, naming conventions

  2. Split-Horizon DNS: Internal vs External Resolution - Same domain multiple answers, internal/external topology, failover patterns

  3. DNS Peering: Hybrid On-Premises Resolution - Multi-project, multi-VPC, on-premises integration, hub-spoke architecture

  4. DNS Forwarding: Cấu hình Upstream Resolvers - Forwarding zones, external DNS, resolver chains, failure handling

  5. Private DNS Zones: VPC Binding & Zone Discovery - VPC attachment, zone discovery, multi-VPC patterns, GKE integration

  6. Cloud DNS for GKE: Alternatives & Performance at Scale - GKE DNS stack, kube-dns vs CoreDNS, external service discovery, multi-cluster

  7. Response Policy Zones (RPZ): Internal Overrides & Security - RPZ mechanisms, security use cases, malware blocking, internal redirects

  8. NodeLocal DNSCache: Latency Reduction & Caching Mechanics - Local caching, performance impact, deployment, troubleshooting

  9. DNS Resolution Path: Pod → NodeLocal → Cloud DNS → Upstream - Complete flow, layer-by-layer troubleshooting, debugging tools

  10. DNS Query Logging: Detection, Audit & Compliance - Query logging, exfiltration detection, BigQuery analysis, alerting

  11. TTL Tuning: High-Churn Environments & Consistency - TTL mechanics, environment-specific tuning, eventual consistency

  12. DNSSEC: Validation & Key Management - DNSSEC architecture, validation, key signing, operational considerations

  13. Multi-Cluster DNS: Cloud Service Directory Patterns - ServiceImport/Export, Service Directory, cross-cluster routing, failover


PHẦN II: GOOGLE KUBERNETES ENGINE — KIẾN TRÚC TOÀN DIỆN


Chương 5: GKE Control Plane Internals — Stateful Systems at Scale

Tại sao quan trọng: Control plane là "bộ não" của cluster. Hiểu cơ chế reconciliation, etcd behavior, control plane limitations là điều kiện tiên quyết debug scheduling failures, API server latency, upgrade issues.

Chapter 5 Full Index & Learning Paths

Các chủ đề con:

  1. GKE Managed Control Plane Model — Standard vs Autopilot - Google quản lý gì, customer quản lý gì, implications cho operations

  2. Kiến Trúc Control Plane Components — API Server, Scheduler, Controller-Manager - Mỗi component role, dependencies, failure modes, interoperability

  3. etcd vs Spanner Backend — GKE State Storage & Consistency Model - Storage backends, consistency guarantees, latency implications, backup strategies

  4. etcd Architecture Deep Dive — Quorum, Replication, Watch Mechanism, Compaction - Raft consensus, replication log, watch caching, compaction schedule, performance limits

  5. Watch Caching & API Server Local Cache — Stale Reads, Reconnection Behavior - Cache mechanics, stale reads, watch connection handling, cache invalidation

  6. Kubernetes Informer Pattern — List-Watch Protocol, Local Cache, Resync Intervals - List-watch protocol, informer cache, resync mechanics, shared factory pattern

  7. Controller Reconciliation Loops — Level-Triggered vs Edge-Triggered Design - Reconciliation patterns, level vs edge-triggered, failure modes, idempotency

  8. API Priority and Fairness (APF) — Flow Schemas, Priority Levels, Rate Limiting - Request prioritization, flow classification, token bucket algorithm, debugging rejections

  9. Admission Control Pipeline — MutatingAdmissionWebhook, ValidatingAdmissionWebhook - Request processing pipeline, webhook execution order, failure modes, cluster stability

  10. Mutating Admission Policies — CEL-Based Policies, Webhook Alternatives - CEL expressions, policy enforcement, webhook alternatives, performance tradeoffs

  11. API Server Request Lifecycle — Authentication → Authorization → Admission → Storage - Full request path, latency breakdown, bottleneck analysis

  12. Control Plane Scalability — Request Rate Limits, Watch Connection Limits, Burst Handling - Scale limits, capacity planning, failure at scale, workarounds

  13. Control Plane Connectivity — DNS-Based vs IP-Based Endpoint, Authorized Networks - Endpoint types, authorized networks, network security implications

  14. Private Cluster Control Plane — Private Endpoint, Cloud NAT, Node Access - Private endpoint setup, node connectivity, security benefits

  15. Credential Rotation & Zero-Downtime Updates — SSL Certificates, CA Rotation, IP Rotation - Certificate lifecycle, CA rotation, zero-downtime strategies

  16. Control Plane SLA, Release Channels, & Versioning Policy - Availability guarantees, release cadence, version support windows, version skew policy


Chương 6: GKE Node Lifecycle & Pool Management

Tại sao quan trọng: Node management là nơi xảy ra phần lớn operational incidents. Node not ready, OOM kills, disk pressure — hiểu lifecycle giúp thiết kế clusters chịu lỗi tốt hơn.

Điều kiện tiên quyết: Chương 5, Container-Optimized OS cơ bản

Mức độ sâu: 5/5

Chapter 6 Full Index & Learning Paths

Các chủ đề con:

  1. COS, Node Bootstrap, Node Conditions và Auto-Repair - COS hardening/immutable filesystem, kubelet registration, startup taints, node conditions, eviction behavior, auto-repair trigger và cơ chế thay node

  2. Node Pool Upgrades, Draining, Maintenance Windows và Cluster Disruption Budget - Surge vs blue-green, maxSurge/maxUnavailable, cordon vs drain, PDB, maintenance windows/exclusions, giới hạn tần suất gián đoạn

  3. Spot, ARM, Confidential Nodes, Reservations và Node Labeling Strategy - Spot preemption, grace shutdown, ARM T2A compatibility, confidential computing, reservation affinity, chiến lược labels cho scheduling

  4. Max Pods, Flex Pod CIDR, Boot Disk, Local SSD và Capacity Design - max pods per node, alias IP sizing, discontiguous Pod CIDR, boot disk performance, local SSD patterns

  5. kubelet & containerd Configuration cho Production GKE - kubelet tuning, eviction thresholds, cgroup v2 migration, registry mirrors, custom TLS CA, image pulling behavior


Chương 7: GKE Networking Internals — VPC-Native, CNI, Dataplane V2 Deep Dive

Tại sao quan trọng: GKE networking là nơi phức tạp nhất. Hiểu packet path từ pod đến pod, qua service, ra internet là điều kiện tiên quyết debug latency, packet drops, network policy violations.

Điều kiện tiên quyết: Chương 3, Linux networking (namespaces, iptables, veth pairs, bridge)

Mức độ sâu: 5/5

Chapter 7 Full Index & Learning Paths

Các chủ đề con:

  1. VPC-Native Architecture — Alias IP, Pod CIDR Sizing & Migration - VPC-native vs routes-based, alias IP ranges trên NIC node, routes-based deprecation, Pod CIDR sizing & max-pods-per-node, secondary subnet sizing, discontiguous Pod CIDR, IP migration

  2. CNI Evolution & Dataplane V2 — kubenet, Calico, eBPF/Cilium - kubenet legacy, Calico iptables ceiling, GKE Dataplane V2 (anetd DaemonSet, eBPF programs, no kube-proxy), eBPF vs iptables 260K endpoint limit, Cilium identity model

  3. Detailed Packet Path Analysis — 5 Đường Đi Của Gói Tin - Same-node & cross-node pod-to-pod, pod-to-Service (ClusterIP DNAT), pod-to-external (masquerade/Cloud NAT), external-to-pod (LoadBalancer/NEG container-native)

  4. kube-proxy & Service Dataplane — iptables vs eBPF - iptables mode chains/DNAT/session affinity, chain explosion O(Services × Endpoints), lock contention, rule resyncing & control-plane latency, Dataplane V2 eBPF replacement

  5. NetworkPolicy Enforcement — Calico iptables vs Dataplane V2 eBPF - Mô hình default-deny, Calico ipset theo IP, Dataplane V2 theo Cilium identity, FQDN egress, NetworkPolicy logging, anti-patterns isolation

  6. Troubleshooting Toolkit — tcpdump, nsenter, Hubble, Connectivity Tests - tcpdump trong pod network namespace, nsenter cấp node, Hubble, GCP Connectivity Tests, ip route/arp/iptables, conntrack limits, eBPF tracing với bpftrace


Chương 8: GKE Scheduler — Algorithms, Affinity, Resource Model

Tại sao quan trọng: Scheduling failures là nguyên nhân hàng đầu Pod stuck in Pending. Hiểu cơ chế scoring/filtering giúp thiết kế node pools, resource requests đúng ngay từ đầu.

Điều kiện tiên quyết: Chương 6, 7; Kubernetes resource model (requests/limits)

Mức độ sâu: 5/5

Chapter 8 Full Index & Learning Paths

Các chủ đề con:

  1. Scheduler Architecture & Workflow — Scheduling Framework, Cycle & Queue - Scheduling cycle vs binding cycle, toàn bộ extension point (PreFilter→Filter→PostFilter→Score→Reserve→Permit→Bind), optimistic locking, ba hàng đợi activeQ/backoffQ/unschedulablePods, QueueingHints, scheduler metrics

  2. Filter & Score Plugins — Lọc Node & Chấm Điểm - Filter plugins (NodeResourcesFit, NodeAffinity, TaintToleration, PodTopologySpread, VolumeBinding), Score plugins, LeastAllocated vs MostAllocated (spread vs bin-packing), GKE optimize-utilization, percentageOfNodesToScore

  3. Node Affinity & Inter-Pod Affinity/Anti-Affinity - nodeAffinity (required/preferred, operators, weight), inter-pod affinity/anti-affinity (topologyKey, namespaceSelector), chi phí O(pods×namespaces), anti-pattern required anti-affinity hostname, tương tác autoscaler

  4. Pod Topology Spread Constraints — Phân Bố Theo Failure Domain - Công thức skew, maxSkew, minDomains, whenUnsatisfiable (DoNotSchedule vs ScheduleAnyway), nodeAffinityPolicy/nodeTaintsPolicy, matchLabelKeys, so sánh với podAntiAffinity

  5. Taints & Tolerations — Ràng Buộc "Đẩy" Node - Ba effect NoSchedule/PreferNoSchedule/NoExecute, tolerationSeconds, operator Equal/Exists, taint-based eviction theo node condition, default toleration 300s, taints mặc định GKE (GPU/Spot/cordon)

  6. Resource Model, QoS & Node-Pressure Eviction - requests vs limits, CPU CFS quota throttling, memory OOM kill, QoS (Guaranteed/Burstable/BestEffort), node-pressure eviction (soft/hard threshold), oom_score_adj, overcommit, vì sao eviction không tôn trọng PDB

  7. Pod Priority & Preemption - PriorityClass (value, globalDefault, preemptionPolicy Never), thuật toán chọn victim, nominatedNodeName, PDB best-effort, cross-node preemption, cascading eviction, starvation, ResourceQuota giới hạn priority

  8. Extended Resources & GPU Scheduling - requests=limits cho extended resources, nvidia.com/gpu, taint GPU + ExtendedResourceToleration, device plugin/driver, GPU sharing (time-sharing/MIG), stranded GPU, TPU & Dynamic Workload Scheduler

  9. GKE Autopilot Scheduling, Custom ComputeClasses & Scheduler Extenders - Autopilot ép tỷ lệ CPU:memory & từ chối/điều chỉnh request, compute classes, custom ComputeClasses (priorities/fallback, activeMigration, consolidation, nodePoolAutoCreation), scheduler extenders vs plugins, Kueue/Volcano


Chương 9: GKE Autoscaling — HPA, VPA, Cluster Autoscaler, KEDA

Tại sao quan trọng: Autoscaling là trái tim cost optimization và reliability. Hiểu sai autoscaling → chậm scale-up (outage), expensive over-provisioning, hoặc flapping destabilizing cluster.

Điều kiện tiên quyết: Chương 8, Cloud Monitoring metrics

Mức độ sâu: 5/5

Chapter 9 Full Index & Learning Paths

Các chủ đề con:

  1. HorizontalPodAutoscaler — Control Loop & Thuật Toán - Control loop chu kỳ 15s, công thức desiredReplicas, tolerance 0.1, dampening Pod chưa Ready/thiếu metric, stabilization window, log atomic vs final recommendation (hpa-controller), debug qua conditions

  2. HPA — Behavior Policies, Metrics Sources & Debugging - behavior autoscaling/v2 (scaleUp/scaleDown, selectPolicy, stabilizationWindowSeconds), Resource/Custom/External metrics, Performance HPA Profile (1000/5000 objects), xung đột HPA+VPA, tương tác rolling update, AbleToScale/ScalingActive/ScalingLimited

  3. VerticalPodAutoscaler — Kiến Trúc, Recommender & Update Modes - Recommender/Updater/Admission Controller, histogram phân rã half-life 24h, OOM bump, update modes (Off/Initial/Recreate/Auto/InPlaceOrRecreate), In-Place Pod Resize, controlledValues, giới hạn VPA

  4. Multidimensional Pod Autoscaling — HPA và VPA Cùng Lúc - Vì sao HPA+VPA xung đột, MultidimPodAutoscaler (CPU ngang + memory dọc), spec & constraints, migration, so sánh với HPA custom metric + VPA Off, failure modes

  5. Cluster Autoscaler — Cơ Chế Scale-Up & Scale-Down - Pod Pending trigger, fake scheduling simulation, expander (least-waste/priority...), location_policy BALANCED/ANY, ngưỡng scale-down 0.5 & các delay, điều chặn scale-down, drain sequence, autoscaling profile

  6. Node Auto-Provisioning — Tự Động Tạo Node Pool - NAP tự tạo/xóa pool, resourceLimits, chọn machine type, khuôn mặc định (Shielded/SA/auto-upgrade), GPU/TPU/Spot, tích hợp ComputeClass, ngưỡng 200 pool, NAP trên Autopilot

  7. CA Troubleshooting, Capacity Buffers & Provisioning Requests - Visibility events (scaleUp/scaleDown/nodePoolCreated), noScaleUp/noScaleDown reasons, Cloud Logging queries, capacity buffer với pause Pod, Provisioning Requests & Dynamic Workload Scheduler, Kueue

  8. KEDA — Kubernetes Event-Driven Autoscaling - Kiến trúc KEDA (operator/metrics-apiserver/webhooks) tạo HPA, ScaledObject vs ScaledJob, scale-to-zero (activation/scaling), defaults (pollingInterval/cooldownPeriod), Pub/Sub & Prometheus scaler, Cloud Tasks/BigQuery


Chương 10: GKE Admission Control & Policy Enforcement — Securing the API

Tại sao quan trọng: Admission control là cửa ngõ security. Misconfigured webhooks → down toàn cluster. Hiểu admission pipeline bắt buộc cho platform engineers.

Điều kiện tiên quyết: Chương 5, Kubernetes API fundamentals

Mức độ sâu: 5/5

Chapter 10 Full Index & Learning Paths

Các chủ đề con:

  1. Admission Pipeline & Built-in Plugins - Vị trí admission trong vòng đời request, hai pha bất biến Mutating → Validating, danh sách plugin bật mặc định, bốn plugin then chốt LimitRanger/ResourceQuota/PodSecurity/NodeRestriction, vì sao trên GKE không sửa được --enable-admission-plugins

  2. Mutating & Validating Webhooks — Cơ Chế Gọi & Dry-Run - WebhookConfiguration (rules/clientConfig), vòng AdmissionReview request/response, JSON Patch, reinvocationPolicy IfNeeded & idempotency, matchPolicy/objectSelector/namespaceSelector, sideEffects & dry-run, audit vs enforce

  3. Webhook Failure Modes, Performance & Stability - failurePolicy Fail vs Ignore, timeoutSeconds (10s/30s) & p99 latency, đường ghi nóng, anti-pattern bắt kube-system/tự-validate/thiếu HA, chiến lược ổn định control plane, break-glass

  4. Webhook Certificate Management — CA Bundle & cert-manager - Webhook là HTTPS server, SAN <service>.<ns>.svc, caBundle & verify, cert-manager + CA Injector tự bơm caBundle, rotation không downtime, self-signed CA, các lỗi x509

  5. PodSecurity Admission (PSA) — Modes & Profiles - Ba mode enforce/audit/warn qua label namespace, version pinning, ba profile privileged/baseline/restricted với từng control chi tiết, enforce áp Pod vs audit/warn áp workload, exemptions, thay thế PSP

  6. Gatekeeper / Policy Controller (OPA) — ConstraintTemplate & Constraint - Kiến trúc webhook + audit controller, ConstraintTemplate (Rego) → Constraint, enforcementAction deny/dryrun/warn, audit loop & status violations, referential constraints, Policy Controller trên GKE (Config Sync/fleet/bundles), Gatekeeper vs PSA

  7. ResourceQuota & LimitRange — Quản Trị Tài Nguyên Namespace - ResourceQuota compute/storage/object-count, scoped quota theo PriorityClass, quy tắc bắt buộc khai requests/limits, LimitRange default/min/max/maxLimitRequestRatio, thứ tự LimitRanger (mutating) → ResourceQuota (validating)

  8. ValidatingAdmissionPolicy (CEL) — Policy In-Process Không Cần Webhook - VAP (GA 1.30) + Binding + paramRef, biến CEL object/oldObject/request/params/namespaceObject, matchConditions/variables, validationActions Deny/Warn/Audit, vì sao CEL loại bỏ failure mode webhook, MutatingAdmissionPolicy, ma trận chọn engine

  9. Organization Policies for GKE & Admission Debugging - Org Policy chặn ở GCP API layer (cluster config) vs Kubernetes admission (Pod config), custom constraints CEL trên container.googleapis.com/Cluster & NodePool, debugging qua audit log Policy Denied/dry-run/log webhook/metric apiserver_admission


Chương 11: GKE Storage — PV/PVC, StorageClasses, CSI Drivers

Tại sao quan trọng: Storage là nơi stateful workloads sống. Hiểu PV/PVC lifecycle, storage classes, volume binding ngăn data loss và performance bottlenecks.

Điều kiện tiên quyết: Chương 6, Kubernetes storage concepts

Mức độ sâu: 5/5

Chapter 11 Full Index & Learning Paths

Các chủ đề con:

  1. Volume Types & Storage Taxonomy — Bản Đồ Toàn Cảnh - Kubernetes volume types (emptyDir, configMap, secret, projected, downwardAPI, hostPath, PVC), phân loại Block/File/Object, access modes RWO/ROX/RWX/RWOP và ngữ nghĩa node-vs-pod, khung quyết định chọn storage

  2. PV/PVC Lifecycle & Dynamic Provisioning - Vòng đời provisioning→binding→mounting→releasing→reclaiming, reclaimPolicy Delete vs Retain, dynamic provisioning end-to-end qua StorageClass/CSI, volume binding modes Immediate vs WaitForFirstConsumer, StorageClass mặc định GKE

  3. Persistent Disk CSI — Block Storage Nền Tảng - PD types và quan hệ IOPS-dung lượng, attach/detach và per-node limit, giới hạn RWX của block device, Regional PD replication đồng bộ, snapshots/cloning/expansion, Stateful HA Operator force-attach

  4. Hyperdisk — Block Storage Thế Hệ Mới - Tách IOPS/throughput khỏi dung lượng, năm loại (balanced/extreme/throughput/ml/balanced-ha), per-VM performance limit, Hyperdisk ML multi-attach ROX, Storage Pools thin provisioning, VolumeAttributesClass

  5. Local SSD & Ephemeral Storage — Tốc Độ Đổi Lấy Độ Bền - Local SSD NVMe physical, emptyDir và ephemeral storage, quy luật mất dữ liệu khi node recreate, provisioning ephemeral-storage-local-ssd, use case đúng và anti-pattern chết người

  6. Filestore CSI — Shared NFS cho ReadWriteMany - Khi nào thật sự cần RWX, service tiers (BASIC_HDD/SSD, Zonal, Enterprise/Regional), Multishares gộp nhiều PVC nhỏ, volume snapshots, NFS tradeoffs về latency/consistency/locking

  7. Cloud Storage FUSE — Object Storage Với File Semantics - Cơ chế FUSE giả lập filesystem trên GCS, sidecar gke-gcsfuse-sidecar, Workload Identity, file cache/metadata cache/parallel downloads, ngữ nghĩa khác POSIX, use case AI/ML read-heavy

  8. Parallelstore & Managed Lustre — Filesystem Song Song cho AI/ML - Parallelstore nền DAOS với erasure coding 2+1 và mô hình temporary storage, Managed Lustre cho HPC, CSI driver, tích hợp GCS, khung chọn Parallelstore/Lustre/Filestore/GCS FUSE

  9. StatefulSets, Volume Expansion & Backup for GKE - StatefulSet volumeClaimTemplates và Pod identity bền vững, PVC giữ khi scale-down, volume expansion online vs cold, Backup for GKE backup config+volume, khác biệt PD snapshot, snapshot lifecycle và chiến lược DR


Chương 12: GKE Security — Hardening, RBAC, Pod Security

Tại sao quan trọng: GKE security có nhiều lớp. Một cấu hình sai có thể phơi bày toàn bộ cluster. Production hardening là bắt buộc, không phải tùy chọn.

Điều kiện tiên quyết: Chương 5, 10, IAM fundamentals

Mức độ sâu: 5/5

Chapter 12 Full Index & Learning Paths

Các chủ đề con:

  1. Security Model & Shared Responsibility — Ai Bảo Vệ Cái Gì - Mô hình trách nhiệm chung GKE, ranh giới dịch chuyển giữa Standard và Autopilot, bảy lớp phòng thủ (org/project → control plane → identity → node → pod → network → supply chain), threat model và các pattern hardening control plane (private cluster, authorized networks)

  2. Authentication & Identity — Bốn Cổng Của Một Request - Luồng request bốn cổng, mô hình hai cổng IAM ↔ RBAC, các phương thức xác thực (Google identity/OIDC, gke-gcloud-auth-plugin, X.509 legacy), token ServiceAccount legacy vs bound (TokenRequest, audience-bound, hết hạn), automountServiceAccountToken: false

  3. RBAC Deep Dive — Role, Binding & Least Privilege - Role vs ClusterRole, quy tắc scope của binding, aggregated ClusterRole, ánh xạ IAM predefined role ↔ RBAC, default role (view/edit/admin/cluster-admin), anti-pattern (cluster-admin cho SA, wildcard, system:authenticated), kiểm tra bằng kubectl auth can-i

  4. Workload Identity Federation for GKE — Hết Long-Lived Key - Hiểm họa của service account key dạng JSON, workload identity pool PROJECT_ID.svc.id.goog, định dạng principal, ba bước trao đổi token qua GKE metadata server, direct binding vs annotation legacy, federation với external IdP

  5. Node Security — Shielded, Confidential, gVisor, COS - Shielded Nodes (Secure Boot, vTPM, Integrity Monitoring), gVisor (runtimeClassName: gvisor, userspace kernel), Confidential Nodes (AMD SEV, mã hóa bộ nhớ), Container-Optimized OS (rootfs read-only, seccomp), node service account tối thiểu, metadata concealment

  6. Pod & Workload Security — Pod Security Standards & securityContext - Ba mức Pod Security Standards (Privileged/Baseline/Restricted), Pod Security Admission (enforce/audit/warn, namespace label), securityContext từng trường (runAsNonRoot, readOnlyRootFilesystem, allowPrivilegeEscalation, drop capabilities), seccomp RuntimeDefault, AppArmor

  7. Network Policy Security — Default-Deny & Đông-Tây - Pattern default-deny, bẫy chặn DNS, Dataplane V2 (Cilium/eBPF), FQDNNetworkPolicy cho egress theo tên miền, Network Policy logging phục vụ điều tra, kiểm soát lateral movement

  8. Admission Control Security — Enforcement Tại Cổng API - Admission như cơ chế enforcement bảo mật, trade-off failurePolicy Fail/Ignore, rủi ro của mutating webhook, ValidatingAdmissionPolicy/CEL in-tree, OPA/Gatekeeper vs Kyverno, Policy Controller managed và constraint framework

  9. Binary Authorization — Chỉ Deploy Image Đáng Tin - Mô hình attestation (digest → attestor → attestation ký số → policy), Artifact Analysis note, execution path qua admission + Binary Authorization API, policy modes (allowlist/require-attestation/dryRun), break-glass có audit, Continuous Validation, Cloud Build/SLSA provenance

  10. Audit Logging, Security Posture & Hardening Checklist - Bốn loại Cloud Audit Logs (Admin Activity, Data Access, System Event, Policy Denied) và bẫy chi phí, Kubernetes audit log và query forensics, GKE Security Posture (config scanning + workload vulnerability scanning), tích hợp Security Command Center, checklist hardening đầy đủ bảy lớp


Chương 13: GKE Workload Identity & Service Accounts — Modern Authentication

Tại sao điều này quan trọng: Workload Identity là cơ chế hiện đại giúp các Pod xác thực với Google APIs mà không cần sử dụng các khóa dịch vụ (service account keys) tồn tại lâu dài. Nếu cấu hình không chính xác, Pod sẽ không thể truy cập hoặc gọi các Google APIs. Vì vậy, việc hiểu rõ luồng trao đổi token (token exchange flow) là yếu tố then chốt để triển khai, vận hành và khắc phục sự cố hiệu quả.

Điều kiện tiên quyết: Chương 12, IAM service accounts, OIDC basics

Mức độ sâu: 5/5

Chapter 13 Full Index & Learning Paths

Các chủ đề con:

  1. Workload Identity Architecture — Cluster Như Một OIDC Provider - Mỗi cluster là một OIDC issuer độc lập, Workload Identity Pool PROJECT_ID.svc.id.goog làm cây cầu để IAM hiểu danh tính Kubernetes, bốn dạng định danh principal/principalSet (theo tên KSA, theo UID, cấp namespace, cấp cluster), identity sameness giữa các cluster cùng project, Fleet Workload Identity

  2. ServiceAccount Token & Projection Mechanics — Danh Tính Được Ký - TokenRequest API và bound token thay legacy secret-based token, projected volume với audience/expirationSeconds/path, cấu trúc JWT (iss issuer cluster, aud sts.googleapis.com, exp, claim kubernetes.io), OIDC issuer endpoint và JWKS để STS verify offline, vòng đời tự refresh

  3. Metadata Server & Token Exchange Path — Trái Tim Của Cơ Chế - gke-metadata-server DaemonSet một Pod/node chặn request 169.254.169.254, trust boundary cấp node và rủi ro hostNetwork bypass, token exchange năm bước qua Security Token Service, caching/refresh lifetime 1 giờ, scale bottleneck (500 conn/node, 3000 SA/cluster, quota 6000 req/phút), network policy egress

  4. IAM Binding Models — Cấp Quyền Cho Danh Tính Workload - Mô hình trực tiếp bind role thẳng cho principal KSA vs mô hình impersonation qua annotation iam.gke.io/gcp-service-account và roles/iam.workloadIdentityUser, principalSet cấp namespace/cluster, cross-project với credential-quota-project, Autopilot luôn bật, return-principal-id-as-email

  5. Workload Identity Federation cho External IdP — Liên Bang Danh Tính Đa Đám Mây - Workload Identity Pool + Provider cho external IdP, token exchange RFC 8693 qua sts.googleapis.com, IdP hỗ trợ (AWS, Entra ID, GitHub Actions, GitLab, Kubernetes, Okta, AD FS, OIDC/SAML), attribute mapping CEL google.subject/attribute.NAME, attribute condition chống confused deputy, direct vs impersonation

  6. Truy Cập Dịch Vụ & Application Default Credentials Patterns - ADC behavior và thứ tự dò credential, vì sao client library tự hoạt động không sửa code, Secret Manager qua Workload Identity, pattern Cloud Storage/Pub-Sub/BigQuery KSA-per-workload, credential helper Artifact Registry, anti-pattern GOOGLE_APPLICATION_CREDENTIALS, khác biệt ADC local-vs-cluster

  7. Debugging Workload Identity — Khi Token Exchange Thất Bại - Quy trình bốn tầng (token gốc, metadata server, STS, IAM binding), debug từ trong Pod bằng curl metadata, verify GKE_METADATA mọi node pool, verify IAM binding và principal string, token validity check, cây quyết định lỗi (unable to detect environment, 403, 404, treo, lỗi rải rác scale)


Chương 14: GKE Observability — Metrics, Logs, Traces

Tại sao quan trọng: GKE sinh ra lượng telemetry khổng lồ và phân tầng. Biết metric nào nằm ở tầng nào, và correlate telemetry để đi từ triệu chứng tới nguyên nhân, là kỹ năng production cốt lõi.

Điều kiện tiên quyết: Chương 5–13, Cloud Monitoring/Logging basics

Mức độ sâu: 5/5

Chapter 14 Full Index & Learning Paths

Các chủ đề con:

  1. Observability Stack — Telemetry Phân Tầng & Mental Model - Ba tầng telemetry (control plane/system/workload), ba loại signal (metric/log/trace) với mô hình chi phí riêng, tích hợp GKE với Cloud Monitoring/Logging/Trace và Managed Prometheus, resource label nhất quán làm nền cho correlation

  2. Control Plane Metrics — Quan Sát Bộ Não Cluster - API server (request rate/error/latency percentile, etcd op latency, inflight, admission webhook), scheduler (pending_pods, scheduling attempt duration, preemption), controller-manager (workqueue depth, reconciliation, node eviction), cách bật --monitoring, mô hình chi phí

  3. System & Workload Metrics — kube-state-metrics, cAdvisor, DCGM GPU - System metrics node, kube-state-metrics (kube_* trạng thái object), cAdvisor (container_*, CPU CFS throttling, memory working set), DCGM GPU metrics (utilization, framebuffer, power, profiling, XID), cardinality

  4. Application Metrics, Startup Latency & Cost Allocation - Golden signals (rate/error/duration/saturation), auto-instrumentation vs custom metric, phân rã startup latency (image pull/init/readiness), GKE cost allocation theo namespace/label (requested vs consumed), FinOps loop

  5. GKE Logs — System, Workload, Audit & Log Control - Logging agent fluent-bit, gói log (SYSTEM/WORKLOAD/API_SERVER/...), system component logs, workload stdout/stderr và structured logging, bốn loại audit log (Admin Activity/Data Access/System Event/Policy Denied), Log Router/sink, exclusion/sampling/retention

  6. Managed Service for Prometheus — PodMonitoring, Rules, PromQL - Managed collection (gmp-operator, collector DaemonSet scrape colocated node, rule-evaluator, alertmanager) và push model, PodMonitoring/ClusterPodMonitoring CRDs, Rules/ClusterRules/AlertmanagerConfig, PromQL trong Cloud Monitoring, high cardinality và metricRelabeling

  7. Managed OpenTelemetry & Custom Metrics cho HPA - Managed OpenTelemetry cho GKE (in-cluster OTLP collector, Instrumentation CRD, signal routing), Google-Built OpenTelemetry Collector, custom metric cho HPA (Custom Metrics Stackdriver Adapter vs Prometheus Adapter, không chạy đồng thời), ServiceMonitor/PodMonitor, liên kết KEDA

  8. Self-Managed Observability — Elastic Stack trên GKE - Khi nào tự vận hành (data sovereignty, multi-cloud, log analytics nâng cao, anti-lock-in), Elastic Cloud on Kubernetes (ECK), performance tuning (Hyperdisk, JVM heap 50%/≤31GB, shard sizing, ILM hot-warm-cold), khung quyết định managed vs self-managed, pattern hybrid

  9. Troubleshooting & Dashboard — Tích Hợp Metrics, Logs, Traces - GKE dashboard trong Cloud Console, workflow correlate dashboard → metric → log → trace qua resource label chung, runbook (Pod Pending, OOMKill, latency spike, API server overload, node NotReady), alerting SLO/burn-rate tránh alert fatigue


Chương 15: GKE Upgrade Mechanics & Disruption Management

Tại sao quan trọng: Sai upgrade strategy → production outage. Hiểu upgrade mechanics, release channels, node draining là nền tảng để thực hiện zero-downtime upgrades.

Điều kiện tiên quyết: Chương 5, 6, 7, 13

Mức độ sâu: 5/5

Chapter 15 Full Index & Learning Paths

Các chủ đề con:

  1. Release Channels, Versioning & Version Skew Policy - GKE release channels (Rapid/Regular/Stable/Extended) cadence, auto-upgrade triggers, capping behavior, Kubernetes version skew policy control plane ↔ kubelet, n-2 support model, patch version advance notice

  2. Cơ Chế Upgrade Cluster GKE: Control Plane, Node Pool & Autopilot - Control plane upgrade zonal vs regional, node pool sequencing, auto-upgrade vs manual upgrade, Autopilot managed upgrade mechanics, rollout sequencing trong fleet, upgrade notifications Pub/Sub

  3. Node Upgrade Strategies: Surge vs Blue-Green - Surge upgrade (maxSurge/maxUnavailable mechanics, pod scheduling, quota implications), blue-green upgrade (5 phases, parallel pool creation, pod migration, rollback), autoscaled blue-green, chiến lược chọn theo workload type, concurrent node pool upgrades

  4. Maintenance Windows, Exclusions & Cluster Disruption Budget - Maintenance windows (UTC timezone, RRULE recurrence, 48h/32d requirement), ba loại maintenance exclusion (no-upgrades/no-minor/no-minor-node), precedence rules, cluster disruption budget cho fleet, rollout sequencing patterns

  5. Workload Disruption Readiness: PDB, Annotations & Upgrade Notifications - PodDisruptionBudget semantics (minAvailable/maxUnavailable, 1-giờ hard limit, PDB + topology spread), pod-deletion-cost dynamic annotation, safe-to-evict, terminationGracePeriodSeconds + preStop hooks, upgrade notification automation, workload checklist

  6. Troubleshooting Stuck Upgrades & Testing Upgrade Strategy - Diagnose stuck upgrade (PDB blocking, quota exhaustion, node affinity, webhook failures), manual intervention (force drain, rollback blue-green), staging cluster validation, kubectl drain testing, API deprecation checks, post-upgrade validation checklist


Chương 16: GKE Autopilot Mode — Managed Infrastructure

Tại sao quan trọng: Autopilot thay đổi cách tư duy về infrastructure. Hiểu Autopilot mechanics, resource enforcement, compute classes giúp tránh resource waste và Pods bị rejected.

Điều kiện tiên quyết: Chương 5, 8, 9

Mức độ sâu: 4/5

Chapter 16 Full Index & Learning Paths

Các chủ đề con:

  1. Autopilot vs Standard — Managed Node Model, Billing, Feature Gaps - Ranh giới trách nhiệm, billing per-Pod vs per-node, feature comparison đầy đủ

  2. Resource Enforcement — Min/Max Requests, CPU:Memory Ratio - Luồng xử lý khi submit Pod, automatic adjustment, minimum/maximum theo compute class, tỷ lệ CPU:memory enforcement

  3. Compute Classes — Balanced, Scale-Out, Performance, Accelerator - Mapping VM families, resource limits theo class, khi nào dùng mỗi class, Custom ComputeClasses

  4. Security Hardening — Pod Security, Privileged Workloads, Org Policy - Pod Security Standards mặc định, Linux capabilities bị drop, allowlist cho privileged workloads, org policy constraints

  5. Spot Pods & Extended Duration Pods - Preemption behavior (25s grace period), design patterns cho batch jobs, Extended Duration bảo vệ khỏi node upgrades (7 ngày)

  6. Cluster Upgrades — Zero-Downtime, Surge, Maintenance Windows - Control plane zero-downtime, surge upgrade strategy, maintenance windows/exclusions, tương tác với PDB và Extended Duration Pods

  7. Networking — IP Allocation, VPC-Native, hostPort - Fixed 32 Pods per node, Pod CIDR sizing, Cloud DNS requirement, hostPort limitations, Dataplane V2/Cilium

  8. Observability — Metrics, Logs, Monitoring - System metrics available trong Autopilot, Managed Prometheus, structured logging, debugging mà không có SSH access

  9. Migration từ Standard sang Autopilot - Pre-flight check, incompatibility checklist đầy đủ, blue-green và MCS migration strategies, Running Autopilot Pods trong Standard clusters


Chương 17: GKE Multi-Tenancy & Workload Isolation

Tại sao quan trọng: Multi-tenant GKE menghemat cost tapi require careful isolation. Understand boundaries dari namespace isolation, resource quotas, network policies untuk design correctly.

Điều kiện tiên quyết: Chương 8, 10, 12

Mức độ sâu: 4/5

Chapter 17 Full Index & Learning Paths

Các chủ đề con:

  1. Multi-Tenancy Models — Soft vs Hard, Mô Hình Tin Cậy - Soft (namespace logic) vs hard (kernel/hardware/cluster), phổ isolation, blast radius, framework quyết định namespace vs cluster riêng

  2. Namespace Isolation & Hierarchical Namespace Controller - Namespaced vs cluster-scoped resources, shared vs isolated, HNC subnamespace, policy/RBAC propagation theo cây, HierarchicalResourceQuota

  3. RBAC cho Multi-Tenancy — Role vs ClusterRole, Impersonation - ClusterRole + RoleBinding pattern, con đường privilege escalation (create pods, escalate/bind, impersonate), Google Groups for RBAC, tách IAM/RBAC

  4. NetworkPolicy Isolation — Default-Deny, Dataplane V2 - Default-deny per-namespace, ingress/egress, namespaceSelector, chặn metadata server, Dataplane V2 (Cilium/eBPF), endPort/Cilium Identity limits

  5. ResourceQuota & LimitRange — Fair Allocation - Quota tổng namespace vs LimitRange per-object, object count/LoadBalancer limits, default requests, QoS class, noisy neighbor

  6. Pod Security Standards per Namespace - Privileged/baseline/restricted, enforce/audit/warn, label per-namespace, rollout an toàn, giới hạn so với Policy Controller, Autopilot override

  7. GKE Sandbox (gVisor) — Kernel Interception - Sentry/Gofer, syscall re-implementation, overhead syscall-heavy, unsupported features, RuntimeClass gvisor, node pool requirements

  8. Node-Level Isolation — Dedicated Pools, Sole-Tenant Nodes - Taint/toleration (trusted only, có thể bị bypass), sole-tenant hardware isolation, HIPAA/PCI, license-bound, NAP/autoscaler limits

  9. Multi-Tenant Operations — Logging & Cost Attribution - Per-namespace log sinks, Log Views/IAM isolation, GKE Cost Allocation (requested vs consumed), billing export BigQuery, showback/chargeback


Chương 18: GKE Fleet Management & Multi-Cluster Architecture

Tại sao quan trọng: Production GKE deployments biasanya multi-cluster. Fleet management mengurangi toil untuk platform teams operating 10s-1000s of clusters.

Điều kiện tiên quyết: Chương 5–17

Mức độ sâu: 4/5

Chapter 18 Full Index & Learning Paths

Các chủ đề con:

  1. Fleet Concepts & Hub Membership — Tổ Chức Multi-Cluster - Fleet định nghĩa kỹ thuật, hub project, membership resource, Connect Agent, namespace sameness, identity sameness, team scopes

  2. Fleet Workload Identity — Định Danh Thống Nhất Xuyên Cluster - Per-cluster vs fleet WIF pool format, OIDC token structure, principal identifier, IAM binding granularity, migration từ per-cluster

  3. Config Sync Architecture — GitOps Engine Bên Trong - RootSync vs RepoSync, reconciler pipeline (importer/parser/applier), server-side apply, drift remediation, ResourceGroup inventory

  4. Config Sync Sources & Hierarchical Repository — Tổ Chức Cấu Hình Ở Scale - Git/OCI/Helm sources, polling interval, authentication, Helm rendering, hierarchical vs unstructured repo, RootSync+RepoSync delegation model

  5. Policy Controller — OPA Constraint Engine Trên Fleet - ConstraintTemplate Rego policy, Constraint CRD, audit vs enforce mode, admission webhook pipeline, fleet-level policy distribution

  6. Multi-Cluster Services (MCS) — Service Discovery Xuyên Cluster - ServiceExport/ServiceImport CRDs, Cloud DNS integration, endpoint propagation, locality-aware LB, eventual consistency model

  7. Multi-Cluster Ingress & Multi-Cluster Gateway — Global Load Balancing - MCI config cluster pattern, MCG Gateway API, global GCLB, NEG creation, traffic routing, health checks xuyên clusters

  8. Fleet RBAC & Fleet Observability — Governance Xuyên Cluster - Fleet IAM → Kubernetes RBAC mapping, team scopes, OIDC-based authentication, cross-cluster metrics aggregation, fleet-level dashboards

  9. Config Controller — Quản Lý GCP Resources Qua Kubernetes CRDs - Config Connector (KCC) CRD-to-GCP-resource reconciliation, deletion protection, dependency management, Config Controller vs standalone KCC

  10. Anthos Service Mesh Multi-Cluster — Mesh Xuyên Cluster - Trust federation SPIFFE identity, istiod cross-cluster discovery, mTLS xuyên cluster, traffic management, observability

  11. Network Connectivity Center — Hub-and-Spoke Network Topology - Hub orchestration, spoke types (VPC/hybrid/router appliance), data transfer giữa spokes, site-to-site connectivity, multi-region cluster patterns


PHẦN III: NETWORKING & TRAFFIC MANAGEMENT


Chương 19: VPC Routing & Connectivity — Deep Dive

Tại sao quan trọng: Routing & connectivity control plane là nơi sinh ra phần lớn sự cố mạng nghiêm trọng — không phải do "cấu hình sai một dòng" mà do hiểu sai mô hình. Chương này đi sâu hơn Chương 3: thay vì "VPC là gì và cấu hình thế nào", nó mổ xẻ control plane thực sự vận hành thế nào dưới lớp abstraction (route resolution, Cloud Router/BGP internals, propagation, transitivity, firewall engine, VPC-SC). Hiểu để debug và thiết kế, không phải để cấu hình.

Điều kiện tiên quyết: Chương 2 (Andromeda/Jupiter), Chương 3 (VPC Model), BGP fundamentals

Mức độ sâu: 5/5

Chapter 19 Full Index & Learning Paths

Các chủ đề con:

  1. Routing Control Plane Model — Route table sống ở đâu - VPC route table như control-plane object program vào Andromeda, tách control/data plane, không có router trung tâm, scope global vs regional, eventual consistency

  2. Route Resolution & Priority — Thuật toán chọn route đầy đủ - Special paths → policy-based → subnet → custom; longest-prefix → priority → preference category → route-type → ECMP (5-tuple/3-tuple hash), next-hop validity

  3. Cloud Router Architecture — Control plane, không phải datapath - Control-plane-only (không forward packet), regional, HA nội tại, program route vào VPC, vì sao Cloud Router fail không lập tức rớt traffic, giới hạn prefix/peer/session

  4. BGP trên GCP — ASN, session, timers, BFD, multipath - ASN 16/32-bit, link-local 169.254/16, BGP FSM (debug "stuck session"), keepalive/hold timers, BFD failover sub-giây vs hold-timer 60-180s, MD5, graceful restart, MP-BGP

  5. Route Advertisement & Propagation — Điều khiển route đi đâu - Advertisement mode default/custom, dynamic routing mode regional vs global, MED→priority + inter-region cost, route policies import/export, asymmetric reachability, summarization

  6. Hybrid Connectivity Routing — HA VPN & Interconnect - SLA 99.9% vs 99.99% đến từ topology vật lý, edge availability domain, dư thừa giả "2 tunnel 1 router", failover = topology + BFD + route standby + resolution

  7. Transitive Routing & NCC — Phi-transitivity của peering và lời giải - Vì sao spoke-to-spoke không thông qua peering (vắng route, không phải firewall), thiết kế có chủ đích, Network Connectivity Center hub/spoke re-export route, any-to-any

  8. Firewall Policy Evaluation Engine — Thứ tự nhiều tầng & goto_next - Pipeline hierarchical → regional system → global → regional → VPC legacy → implied; allow/deny terminal vs goto_next; priority 0-2.147.483.547; secure tags vs SA; stateful tracking

  9. VPC Service Controls — Perimeter ở tầng API, không phải firewall - VPC-SC như control plane thứ ba (trực giao IAM/firewall), enforce ở API layer, restricted VIP 199.36.153.4/30, access levels, ingress/egress, dry-run, chống exfiltration

Lưu ý phạm vi: Chương này được tổ chức như một deep dive về routing & connectivity control plane, bổ sung (không lặp lại) Chương 3 (VPC Model). Các chủ đề nền tảng — subnet design/CIDR, alias IP, static/system routes cơ bản, firewall fundamentals, VPC Peering cơ bản, Shared VPC, Private Google Access, VPC Flow Logs, Network Intelligence Center — đã được trình bày chi tiết ở Chương 3. Chương 19 tập trung vào cơ chế bên trong: thuật toán route resolution, Cloud Router/BGP internals, propagation control, transitive routing, firewall evaluation engine, và VPC-SC perimeter.


Chương 20: Cloud Load Balancing — Architecture & Mechanics

Tại sao quan trọng: Load balancer là điểm vào của 100% traffic — từ internet, giữa các microservice, tới mọi GKE Service. Cloud Load Balancing không phải một sản phẩm mà là một họ hơn một chục loại LB xây trên ba họ datapath hoàn toàn khác nhau (GFE, Envoy, Maglev/Andromeda). Sai cấu hình ở tầng này → phân phối lệch, health check fail, mất client IP, SSL handshake fail. Chương này đi từ cơ chế datapath bên trong (proxy vs passthrough, terminate ở đâu, state sống ở đâu) thay vì bảng tra "loại nào dùng việc gì".

Điều kiện tiên quyết: Chương 19 (VPC Routing/Andromeda), Chương 7 (GKE Networking), HTTP/HTTPS, TLS handshake, TCP fundamentals (5-tuple, connection state)

Mức độ sâu: 5/5

Chapter 20 Full Index & Learning Paths

Các chủ đề con:

  1. LB Taxonomy & Architecture Model — Ba họ datapath, một mô hình resource - Bản đồ đầy đủ các loại LB; mô hình resource thống nhất (forwarding rule → target proxy → URL map → backend service → backend); ba họ datapath GFE/Envoy/Maglev; proxy vs passthrough là trục quan trọng nhất (source IP, nơi terminate, state, DSR)

  2. Global External Application LB — Anycast, GFE two-tier, Maglev, backend selection - Anycast VIP quảng bá ở biên, Maglev đứng trước GFE, kiến trúc two-tier GFE (first-layer terminate TLS + parse URL map, second-layer chọn backend), balancing mode RATE/UTILIZATION/CONNECTION, capacity scaler, waterfall region selection, session affinity L7, tích hợp CDN/Armor

  3. Envoy-based Regional & Internal ALB — Proxy-only subnet và datapath trong VPC - Envoy managed proxy fleet, proxy-only subnet (vì sao bắt buộc, sizing, source IP tới backend), regional/cross-region scope, locality LB policy & outlier detection, khác biệt cơ chế so với GFE

  4. Passthrough Network LB — Maglev, DSR, consistent hashing & connection tracking - Không proxy, Direct Server Return giữ IP đích, Maglev (external) & Andromeda (internal), consistent hashing chọn backend, connection tracking table, mode PER_CONNECTION vs PER_SESSION, session affinity 5/3/2-tuple, connection persistence, ILB như next-hop

  5. Network Endpoint Groups (NEGs) — Container-native LB và mô hình endpoint - Các loại NEG (zonal GCE_VM_IP_PORT/GCE_VM_IP, serverless, internet, PSC, hybrid), container-native LB bỏ hop kube-proxy, NEG controller reconciliation, readiness gate đồng bộ vòng đời Pod với LB

  6. Health Check Architecture — Distributed prober & backend health model - Distributed prober model (tần suất thực cao hơn cấu hình), dải IP 35.191.0.0/16 & 130.211.0.0/22 (và dải riêng cho passthrough), interval/timeout/threshold semantics, proxy vs passthrough health check, failure modes (firewall, path, flapping)

  7. GKE Service LoadBalancer Integration — externalTrafficPolicy, NEG, session affinity - LoadBalancer Service → passthrough NLB, NEG-based vs legacy instance-group/target-pool, GKE subsetting, externalTrafficPolicy Local vs Cluster (giữ client IP vs phân phối lệch, health check port 10256), sessionAffinity ClientIP

  8. Connection Draining & Graceful Shutdown — Không cắt connection đang chạy - drainingTimeoutSec 0–3600s, khác biệt proxy (giữ TCP connection) vs passthrough (giữ connection tracking entry), phối hợp với preStop hook + terminationGracePeriodSeconds + readiness để đạt zero-downtime deploy

  9. SSL Policies & TLS Termination — Kiểm soát TLS version và cipher - TLS terminate ở đâu (GFE edge vs Envoy), frontend TLS vs backend TLS, minimum TLS version, profile COMPATIBLE/MODERN/RESTRICTED/FIPS/CUSTOM, cipher suite, gắn SSL policy vào target proxy, Google-managed cert

  10. Cloud Armor & Edge Security — WAF và DDoS protection - Enforce ở GFE edge (trước khi tới backend), preconfigured WAF rules (OWASP CRS, SQLi/XSS), rule priority, rate limiting, Adaptive Protection (ML DDoS), Network Edge Security cho passthrough, vì sao WAF L7 chỉ chạy với proxy LB

Lưu ý phạm vi: Chương này tập trung vào cơ chế datapath bên trong của Cloud Load Balancing — proxy vs passthrough, GFE vs Envoy vs Maglev, nơi connection được terminate và state sống ở đâu. Phần L7 expose ứng dụng trên GKE (Ingress controller, Gateway API, BackendConfig/FrontendConfig CRD) được trình bày ở Chương 21, xây trực tiếp trên các Application LB và NEG của chương này.


Chương 21: GKE Ingress & Gateway API — Expose Ứng Dụng Ra Bên Ngoài

Tại sao quan trọng: Ingress/Gateway là cửa ngõ expose ứng dụng externally. Cấu hình sai → SSL issues, 502 errors, lỗ hổng bảo mật. Hiểu cơ chế reconciliation và packet path là điều kiện tiên quyết để debug và thiết kế đúng.

Điều kiện tiên quyết: Chương 7, 20, Kubernetes Services

Mức độ sâu: 5/5

Chương 21 Full Index & Learning Paths

Các chủ đề con:

  1. Tổng Quan GKE Load Balancing — Gateway vs Ingress vs LoadBalancer Service - Ba cơ chế expose, GCP resources tương ứng, framework quyết định chọn đúng cơ chế

  2. GKE Ingress Controller Internals — Reconciliation, External vs Internal ALB, Packet Path - Controller reconciliation loop, annotation-based classification, mapping sang GCP resources, packet traversal từ client đến Pod

  3. BackendConfig & FrontendConfig CRDs — Cấu Hình Nâng Cao cho GKE Ingress - Binding mechanism, health check tuning, Cloud Armor, session affinity, CDN, SSL policies, connection draining

  4. Gateway API Architecture — GatewayClass, Gateway, HTTPRoute và GKE Implementation - Role-separated design, GatewayClass mapping sang GCP LB, controller model, TLS termination, traffic splitting, tại sao Gateway thay thế Ingress

  5. Container-native Load Balancing & NEG Internals — Traffic Thẳng Đến Pod - NEG với Pod IPs, NEG controller lifecycle, Pod Readiness Gates, health check tại Pod level, standalone NEGs

  6. Multi-cluster Ingress & Multi-cluster Gateway — Global Load Balancing Xuyên Cluster - Config cluster pattern, MultiClusterIngress/MultiClusterService CRDs, MCG với -mc GatewayClasses, MCS ServiceImport, failure modes


Chương 22: Cloud DNS & Service Discovery

Tại sao quan trọng: DNS failure là nguyên nhân phổ biến nhất gây outage trong hệ thống microservice. Hiểu resolution path giúp debug "connection refused" đúng layer và phòng tránh ndots:5 lookup storm, CoreDNS conntrack exhaustion.

Điều kiện tiên quyết: Chương 4, DNS fundamentals

Mức độ sâu: 5/5

Chapter 22 Full Index & Learning Paths

Các chủ đề con:

  1. DNS Resolution trong Pod — /etc/resolv.conf, ndots:5, và Lookup Storm - Kubelet cấu hình /etc/resolv.conf như thế nào, ndots:5 tạo ra bao nhiêu DNS queries cho một hostname ngắn, search domain expansion path, negative caching amplification, dnsPolicy options và khi nào override

  2. CoreDNS trong GKE — Plugin Chain và Corefile - Kiến trúc plugin pipeline, Corefile mặc định GKE với các plugin quan trọng (kubernetes, forward, cache, health, ready), kubernetes plugin watch-based sync, conntrack exhaustion vấn đề cốt lõi

  3. Kubernetes DNS Spec — Service Discovery - A/AAAA records cho Services và Pods, SRV records, format service.namespace.svc.cluster.local, hostname/subdomain fields, cross-namespace resolution, record propagation delay

  4. Headless Services & ExternalName Services - Headless (clusterIP: None): DNS trả về Pod IPs, per-Pod DNS records, client-side LB, tích hợp StatefulSet. ExternalName: CNAME mechanism, TLS gotcha, port mapping limitations

  5. NodeLocal DNSCache — Kiến Trúc, Conntrack Fix, và Cơ Chế Fallback - Tại sao tồn tại (conntrack exhaustion, không chỉ latency), DaemonSet trên mỗi node, link-local IP 169.254.20.10, cache TTL behavior, fallback hierarchy (cluster.local → kube-dns-upstream TCP, external → metadata server), GKE Autopilot mandatory

  6. Cloud DNS cho GKE — Private Zones và VPC Scope - Kiến trúc thay thế kube-dns bằng managed infrastructure, automatic zone management, ba DNS scope (cluster/VPC/additive-VPC) không thể đổi sau khi tạo cluster, split-horizon DNS, peering zones cho hybrid connectivity

  7. DNS Debugging & Performance Tuning - Workflow debug 6 layers từ Pod đến upstream, playbook theo symptom (timeout, NXDOMAIN storm, cross-namespace failure, slow external), CoreDNS metrics quan trọng, tuning ndots/cache size/max_concurrent


Chương 23: Cloud NAT — Cơ Chế Phân Bổ Port & Phòng Tránh Cạn Kiệt

Tại sao quan trọng: Port exhaustion của Cloud NAT là lỗi im lặng nguy hiểm — connections bị drop mà không có error message rõ ràng ở tầng application. Hiểu cơ chế phân bổ port là nền tảng cho capacity planning đúng đắn.

Điều kiện tiên quyết: Chương 19, NAT/SNAT fundamentals

Mức độ sâu: 5/5

Chapter 23 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc Cloud NAT & Cơ Chế SNAT - Distributed NAT trên Andromeda SDN, SNAT flow chi tiết, connection tracking, Endpoint-Independent Mapping, tại sao Cloud NAT không giảm băng thông

  2. Port Allocation & Port Exhaustion - Tại sao 64.512 ports/IP, static vs dynamic allocation, port math formula, 5-tuple constraints, TCP TIME_WAIT phantom shortage, Dynamic Port Allocation doubling, chẩn đoán NAT_ALLOCATION_FAILED

  3. Cloud NAT với GKE - Pod egress flow, ip-masq-agent, NAT scope cho VPC-native, capacity planning theo node, GKE Autopilot, cascade failure scenarios

  4. Cấu Hình, Giám Sát & Tối Ưu - TCP/UDP/ICMP timeout mechanics, NAT rules phân tách traffic, metrics quan trọng (port_usage, dropped_sent_packets_count), logging cost, connection pooling patterns


Chương 24: Private Service Connect — Mô Hình Hiện Đại Để Expose Service

Tại sao quan trọng: PSC là cách hiện đại để expose service mà không cần VPC peering. Hiểu PSC là điều kiện để thiết kế kiến trúc service multi-tenant đúng đắn.

Điều kiện tiên quyết: Chương 19, VPC Peering concepts

Mức độ sâu: 5/5

Chapter 24 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc & Components của PSC - Service attachment (producer), endpoint (consumer), backend, interface; cơ chế NAT cho phép overlapping IP ranges; data flow chi tiết ở tầng datapath; tại sao PSC đạt line-rate; endpoint vs backend

  2. PSC vs VPC Peering - Routing model (full route exchange vs single NAT'd IP), transitivity, blast radius bảo mật, IP coordination, scale limits, decision matrix khi nào dùng cái nào

  3. PSC cho Google APIs - Bundle all-apis vs vpc-sc, khác biệt bản chất với Private Google Access, global internal IP, DNS p.googleapis.com & Service Directory, pattern hybrid on-prem

  4. PSC cho Managed Services & GKE Control Plane - Mô hình tenant project, service attachment auto-created, Cloud SQL/Memorystore/AlloyDB, GKE control plane qua PSC (1000 vs 75 cluster), service connection policy

  5. Consumer vs Producer — IAM & Approval Workflow - Connection preference ACCEPT_AUTOMATIC vs ACCEPT_MANUAL, accept/reject list (5000/64), connection limit per consumer, vòng đời PENDING→ACCEPTED, IAM hai phía, org policy

  6. Global Access & DNS Automation - Global access cross-region (HA, consumer-side), DNS automation tự tạo A record/private zone, Service Directory, pattern multi-region HA

  7. Troubleshooting PSC - Năm connection state, cây quyết định cho PENDING, NAT subnet exhaustion, region mismatch, Connectivity Tests, flow logs & metrics, runbook tổng hợp


Chương 25: Cloud Router & BGP Internals — Control Plane của Dynamic Routing

Tại sao quan trọng: Cloud Router là control plane cho dynamic routing. BGP sai → routes không được advertise hoặc routes không chính xác được propagate âm thầm.

Điều kiện tiên quyết: Chương 19, BGP fundamentals

Mức độ sâu: 4/5

Chapter 25 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc Cloud Router — Distributed BGP Control Plane - Cloud Router không forward packet, chỉ là BGP control plane, cách nó lập trình routes vào Andromeda, redundancy nội tại, regional scope

  2. BGP Session Internals — Cơ Chế Thiết Lập Và Duy Trì Session - BGP FSM 6 states, OPEN/KEEPALIVE/UPDATE messages, eBGP vs iBGP, link-local 169.254.x.x, ASN configuration, hold timer, MD5 authentication

  3. Route Advertisement & Import — Điều Khiển Routes Đi Đâu Về Đâu - Advertisement modes (default vs custom), subnet routes, custom routes, custom learned routes, MED, best-path selection

  4. BGP Route Policies — Filtering và Modification Với CEL - Import/export policies, Common Expression Language, named sets, BGP communities, fail-open model

  5. BFD — Fast Failure Detection - BFD protocol mechanics, timer configuration, detection time formula, dampening system, BGP integration

  6. Cloud Router Với Hybrid Connectivity — HA VPN và Cloud Interconnect - BGP per tunnel, HA VPN SLA topology requirements, VLAN attachments, active-active vs active-passive failover

  7. Routing Modes — Regional vs Global - Regional vs global routing mode, inter-region cost, asymmetric routing, best-path selection modes

  8. Monitoring BGP Sessions - BGP session status, list routes, Cloud Monitoring metrics, troubleshooting playbooks


Chương 26: Cloud Interconnect & Cloud VPN — Hybrid Connectivity

Tại sao quan trọng: Hybrid connectivity là foundation của enterprise GCP deployments. Quyết định thiết kế ảnh hưởng trực tiếp đến latency, cost, và security posture.

Điều kiện tiên quyết: Chương 25, MPLS/WAN networking cơ bản

Mức độ sâu: 4/5

Chapter 26 Full Index & Learning Paths

Các chủ đề con:

  1. Cloud VPN — Cơ Chế Tunnel, IKE/ESP, HA Model - Classic vs HA VPN (single IP vs dual interface, 99.9% vs 99.99% SLA), IPsec tunnel mode mechanics (IKE Phase 1/2, ESP encapsulation, replay protection 4096-packet window), dynamic BGP routing over tunnel, MTU/MSS clamping cho AEAD ciphers, bandwidth ceiling 3 Gbps per tunnel

  2. Cloud Interconnect — Kết Nối Vật Lý, VLAN, BGP & MACsec - Dedicated vs Partner Interconnect (physical last-mile vs provider model), VLAN attachments (802.1Q logical connections, MTU 8896 jumbo frames), BGP over Interconnect (link-local 169.254/16, L2 vs L3 partner), redundancy topology cho 99.99% SLA (4 connections, 2 metros, edge availability domains), MACsec (IEEE 802.1AE, GCM-AES-256, hitless key rotation)

  3. Network Connectivity Center — Hub-and-Spoke Orchestration - Hub-and-spoke model, spoke types (VPC/hybrid/router appliance/gateway), site-to-site data transfer (route re-advertisement, GCP backbone as WAN), NCC vs VPC Peering mesh vs SD-WAN, giới hạn IPv4-only cho hybrid spokes

  4. Production Patterns & Monitoring - Active-active ECMP vs active-passive (MED/AS-path prepending), BFD fast failover (300ms vs 60s hold timer), monitoring metrics (tunnel health, BGP session, optical power, packet loss), operational runbook cho tunnel down/packet loss/BGP flapping


Chương 27: Bảo Mật Mạng — Firewall Policies, Cloud NGFW, Cloud Armor

Tại sao quan trọng: Network security là outer perimeter của mọi deployment GCP. Misconfigured firewall expose sensitive services hoặc block legitimate traffic. Hiểu ngăn xếp bảo mật nhiều lớp — từ stateful VPC rules đến L7 inspection, API boundary, edge WAF — là bắt buộc để thiết kế và vận hành production systems an toàn.

Điều kiện tiên quyết: Chương 19, security fundamentals

Mức độ sâu: 4/5

Chapter 27 Full Index & Learning Paths

Các chủ đề con:

  1. VPC Firewall Rules — Stateful Engine & Connection Tracking - Cơ chế stateful connection tracking trong Andromeda: 5-tuple, connection table, timeout/eviction, giới hạn theo machine type, ingress/egress asymmetry, priority resolution, implied rules

  2. Hierarchical Firewall Policies — Thi Hành Ở Tầng Tổ Chức - Evaluation order (org → folder → project → VPC), goto_next semantics, Secure Tags vs network tags, association model, tại sao HFP là nền tảng org-wide security governance

  3. Cloud NGFW — L7 Inspection, FQDN Filtering & TLS Interception - Ba tiers (Essentials/Standard/Enterprise), Firewall Endpoint deployment, FQDN vs URL filtering, TLS inspection/decryption mechanism, IPS signature-based detection, trade-off latency/throughput

  4. VPC Service Controls — Perimeter API & Chống Data Exfiltration - VPC-SC là API access boundary (không phải firewall), enforce tại GCP control plane, service perimeter design, access levels, ingress/egress rules, restricted VIP, dry-run mode

  5. Cloud Armor — WAF, OWASP CRS & Adaptive Protection - Tại sao chỉ hoạt động với proxy LB, security policy evaluation tại GFE edge, OWASP CRS 4.x sensitivity levels, rate limiting mechanics, Adaptive Protection ML model (baseline, confidence scoring, auto-deploy)

  6. Cloud IDS — Intrusion Detection Qua Packet Mirroring - Detect-only (không block), Palo Alto Networks App-ID/vulnerability/anti-spyware, packet mirroring mechanism, IDS Endpoint zonal deployment (5 Gbps), alert severity, integrate với Cloud NGFW cho IPS

  7. Network Intelligence Center — Firewall Insights - Phân tích shadowed rules, overly permissive rules, deny rules zero hits; yêu cầu Firewall Rules Logging; giới hạn TCP/UDP; ML forecasting cho proactive cleanup

  8. Secure Web Proxy — Egress Filtering Cho Workloads - Explicit HTTP/HTTPS proxy với default deny-all, deployment modes (explicit/PSC/next-hop), policy framework (sessionMatcher + applicationMatcher), mTLS client auth, GKE egress control pattern

  9. Private NAT — Egress An Toàn Trong Nội Bộ GCP - Private-to-private translation cho NCC spokes và hybrid NAT, gateway không trong data plane, blocking unsolicited inbound, overlapping IP constraint, phân biệt với Cloud NAT public internet egress


PHẦN IV: STORAGE & DATA SYSTEMS


Chương 28: Cloud Storage — Kiến Trúc, Tính Nhất Quán và Hiệu Năng

Tại sao quan trọng: GCS là universal data store của GCP. Hiểu consistency model và performance characteristics giúp tránh data races và slow reads ở scale thật.

Điều kiện tiên quyết: Object storage concepts, HTTP/S basics

Mức độ sâu: 4/5

Chapter 28 Full Index & Learning Paths

Các chủ đề con:

  1. Object Model, Buckets, Generations & Consistency — Cơ chế lưu trữ bên trong: bucket namespace, object immutability, generation numbers, metageneration; strong consistency model post-2021 với edge cases quan trọng (cached public objects, HMAC delays, batch non-atomicity)

  2. Storage Classes & Location Types — Cơ chế phân tầng Standard/Nearline/Coldline/Archive (SLA, minimum duration, retrieval fee); replication mechanics của regional/dual-region/multi-region; Autoclass; trade-off immutable location decision

  3. Lifecycle Management — Tiering & Deletion Automation — Rule evaluation engine, AND-logic conditions, SetStorageClass vs Delete vs AbortIncompleteMultipartUpload; conflict resolution; interaction với versioning và soft delete; Autoclass vs manual lifecycle

  4. Access Control: IAM, ACLs, Signed URLs & Requester Pays — Hai hệ thống song song IAM và ACLs; Uniform Bucket-Level Access (UBLA); IAM roles GCS; V4 Signed URL signing mechanics, expiry, revocation limitations; Requester Pays billing model

  5. VPC Service Controls Integration — GCS với VPC-SC perimeter ở API layer; restricted VIP (199.36.153.4/30); ingress/egress rules cho data exfiltration prevention; dry-run mode; interaction với Transfer Service và serverless

  6. Cloud Storage FUSE — POSIX Interface Trên Object Storage — Kiến trúc FUSE daemon; object name → directory translation; write path buffering; read sequential vs random; caching layers (stat/list/file); POSIX semantic gaps; GKE sidecar integration; anti-patterns

  7. Transfer Service — Bulk Migration & Scheduled Ingestion — Managed vs self-hosted agents; checksum và data integrity verification; incremental sync với overwrite conditions; bandwidth throttling; migration patterns (S3→GCS, on-prem); VPC-SC interaction

  8. Performance: Throughput Scaling & Parallel Operations — Auto-scaling request handling (baseline 1K write/5K read req/s); gradual ramp-up requirement; object naming và hotspot avoidance; HNS 8x baseline; parallel uploads với composite objects; parallel downloads; client-side bottlenecks


Chương 29: Persistent Disk & Hyperdisk — Block Storage

Tại sao quan trọng: Loại disk và cách sizing ảnh hưởng trực tiếp đến application performance. Giới hạn IOPS/throughput thường bị hiểu sai, dẫn đến bottleneck ẩn trong production.

Điều kiện tiên quyết: Compute Engine basics

Mức độ sâu: 4/5

Chapter 29 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc Nội Tại & Các Loại Persistent Disk — Persistent Disk là network block device trên Colossus; ba tầng replication & fault domains; bốn loại PD (standard/balanced/ssd/extreme); đặc tính IOPS/throughput; tại sao write latency cao hơn Local SSD; implications khi dùng pd làm database storage

  2. Mô Hình Hiệu Năng & Giới Hạn VM — Hai ceiling độc lập: per-disk (IOPS/GiB) và per-VM (machine type limit); công thức tính performance thực tế; baseline guarantees cho pd-balanced và pd-ssd; network egress như ceiling thứ ba; I/O queue depth requirement; sizing hướng đến performance target

  3. Hyperdisk & Provisioned Performance Model — Tách IOPS/throughput khỏi capacity; năm loại Hyperdisk (Balanced/Extreme/ML/Throughput/Balanced HA); machine type constraints; torn write protection & MySQL doublewrite buffer elimination; Storage Pools & Hyperdisk Exapools; khi nào Hyperdisk vs PD

  4. Multi-Writer Disks & Giới Hạn Concurrent Access — Chỉ Hyperdisk Balanced/Extreme hỗ trợ multi-writer; tại sao ext4/xfs không dùng được với shared block; SCSI Persistent Reservations cho I/O fencing; clustered file systems (OCFS2, GFS2, VMFS); Oracle RAC và SQL Server FCI patterns; shared performance budget

  5. Snapshot Internals & Cross-Region Backup — Changed block tracking tầng Colossus; full → incremental chain; deletion behavior không giảm storage ngay; global vs regional snapshot scope; cross-region copy tạo chain độc lập; crash-consistent vs application-consistent; snapshot scheduling best practices

  6. Regional PD & High Availability — Synchronous cross-zone replication, RPO=0; ba trạng thái replica (synced/catching-up/degraded); force-attach cho failover dưới 1 phút; initial sync time sau khi tạo; machine type constraints (E2/N1/N2/N2D only); Regional PD vs Hyperdisk Balanced HA; kết hợp với Stateful MIGs

  7. Encryption: Google-Managed & CMEK — Envelope encryption hai tầng (DEK + KEK); mã hóa trong VM trước khi data rời guest; CMEK với Cloud KMS — constraints quan trọng (không áp dụng cho disk cũ, không revert); CSEK cho key hoàn toàn off-cloud; key rotation mechanics; Confidential Hyperdisk cho in-use protection; khi nào thực sự cần CMEK


Chương 30: Filestore & Advanced Storage Options

Tại sao quan trọng: Filestore cung cấp shared NFS cho multi-reader workloads. Hiểu sai performance tier dẫn đến IO bottlenecks không thể giải quyết sau khi deploy.

Điều kiện tiên quyết: Chương 11, NFS basics

Mức độ sâu: 3/5

Chapter 30 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc Filestore & Service Tiers — NFS server model, data path, NFSv3 vs NFSv4.1, performance model (IOPS/throughput per tier), Basic HDD/SSD vs Zonal vs Regional vs Enterprise, single-client cap và nconnect, constraints và failure modes

  2. Filestore CSI Driver & Multishares — CSI driver architecture, dynamic provisioning (một PVC một instance), Multishares internal model (packing 80 shares/instance), StorageClass parameters, capacity management, limitations (no snapshot, StorageClass immutability, provisioning latency)

  3. Backup, Snapshot & Recovery — Cơ chế backup (differential copy, cross-region), snapshot (copy-on-write, child resource của instance), so sánh backup vs snapshot, recovery từng loại, scheduling automation, failure scenarios

  4. Regional HA & Cross-Zone — Synchronous 3-zone replication, transparent failover, IP stability, limitations trong zone failure, asynchronous cross-region replication, promote standby, performance trade-offs Regional vs Zonal


PHẦN V: IAM, SECURITY & COMPLIANCE


Chương 31: IAM Deep Dive — Model, Propagation, Conditions

Tại sao quan trọng: IAM là lớp access control duy nhất và phổ quát trong GCP. Hiểu propagation model, deny policies, và conditions là điều kiện tiên quyết để thiết kế least-privilege architecture và ngăn privilege escalation.

Điều kiện tiên quyết: Chương 1, resource hierarchy

Mức độ sâu: 5/5

Chapter 31 Full Index & Learning Paths

Các chủ đề con:

  1. IAM Policy Model — Allow Policy, Evaluation Order và PAB — Ba loại policy (Allow, Deny, PAB), cấu trúc allow policy, evaluation ba pha (PAB → Deny → Allow), effective policy từ hierarchy, etag-based optimistic locking

  2. Role Types — Basic, Predefined, Custom — Tại sao basic roles nguy hiểm trong production, predefined role expansion behavior, thiết kế custom roles, permission support levels, group-based bindings

  3. Policy Hierarchy & Propagation — Union inheritance model, không thể override từ level thấp, eventual consistency, propagation delay (30-60s), token caching gap trong incident response, debugging với Policy Troubleshooter

  4. IAM Conditions & CEL Expressions — Time-based / resource-based / request-based attributes, evaluation semantics (false vs cannot-evaluate), limitations với basic roles và allUsers, resource tags vs labels, anti-patterns

  5. IAM Deny Policies — Deny Before Allow — Cơ chế deny-before-allow, cấu trúc deny rule, deniable permissions và wildcard patterns, exceptionPrincipals, semantics fail-secure trong deny conditions, khi nào dùng deny vs revoke role

  6. Service Accounts — Key Management, Impersonation, Default SA — Ba loại SA, SA keys là anti-pattern, cơ chế impersonation/token exchange, delegation chains, nguy hiểm của default Compute SA với Editor role, org policy để disable auto-grant

  7. Audit Logging — Admin Activity, Data Access, Policy Denied — Bốn loại Cloud Audit Logs và defaults, cấu trúc log entry (principalEmail, delegationInfo, authorizationInfo), chiến lược enable Data Access logs có kiểm soát chi phí, forensic queries

  8. VPC Service Controls — Perimeter Orthogonal với IAM — VPC-SC enforce ở API layer (sau Org Policy, trước IAM), service perimeter, access levels, ingress/egress rules, restricted.googleapis.com vs private.googleapis.com, dry-run migration

  9. Organization Policies — Constraints và Inheritance — What vs who (Org Policy vs IAM), enforcement ở GCP API layer, constraint types (managed/list/boolean/custom), inheritance merge semantics, production constraints quan trọng, CEL trong custom constraints

  10. IAM Recommender — Least-Privilege Automation — 90-day observation window, ML co-occurrence analysis, bốn recommendation subtypes (REMOVE/REPLACE/REPLACE_CUSTOMIZABLE/SERVICE_AGENT), lateral movement insights, giới hạn (seasonal patterns, conditional bindings, manual apply required)


Chương 32: Secret Manager & Cloud KMS — Quản Lý Secrets & Mã Hoá

Tại sao quan trọng: Quản lý secrets và mã hoá dữ liệu là hai trụ cột bảo mật không thể thiếu trong production. Hiểu cơ chế lưu trữ secrets, versioning mechanics, và KMS key hierarchy là điều kiện để thiết kế encryption strategy đúng đắn.

Điều kiện tiên quyết: Chương 31, encryption fundamentals (AES-256, asymmetric cryptography)

Mức độ sâu: 4/5

Chapter 32 Full Index & Learning Paths

Các chủ đề con:

  1. Secret Manager — Model, Versioning & Replication Internals — Secret object vs Secret Version (immutable payload), version lifecycle states (ENABLED/DISABLED/DESTROYED), alias latest và rủi ro, replication automatic vs user-managed (data residency), CMEK integration, IAM resource-level, audit logging

  2. Secret Rotation & Các Patterns Truy Cập trong GKE — Rotation flow qua Pub/Sub (SECRET_ROTATE notification, không auto-rotate value), CSI driver architecture (SecretProviderClass, tmpfs mount, auto-refresh GKE 1.32+), init container pattern, so sánh kỹ thuật với environment variables

  3. Cloud KMS Key Hierarchy & Protection Levels — Key Ring (container, immutable, không xóa được), CryptoKey (immutable purpose/protection/algorithm), CryptoKeyVersion (key material thực), primary version mechanics, key purposes (ENCRYPT_DECRYPT/ASYMMETRIC_SIGN/MAC), protection levels (SOFTWARE/HSM/EXTERNAL)

  4. Envelope Encryption & CMEK — DEK/KEK model, encryption/decryption data flow chi tiết, tại sao KEK không bao giờ rời Cloud KMS, Google's internal key hierarchy ba tầng, CMEK vs Google-managed, IAM binding cho service agent, revoke access bằng disable key

  5. Key Rotation & Vòng Đời CryptoKeyVersion — Rotation tạo version mới không destroy version cũ, automatic rotation chỉ cho ENCRYPT_DECRYPT, manual rotation, version states, DESTROY_SCHEDULED (30 ngày mặc định từ Feb 2024), re-encryption requirement

  6. Cloud EKM — Keys Ngoài Hạ Tầng Google — EXTERNAL vs EXTERNAL_VPC protection levels, double-encryption model (internal + external material), data inaccessible khi external KMS down, HYOK use cases, khi nào thực sự cần EKM vs CMEK


Chương 33: VPC Service Controls & Organization Policies

Tại sao quan trọng: VPC SC là primary control ngăn data exfiltration. Org Policies cung cấp guardrails cho toàn tổ chức ở scale.

Điều kiện tiên quyết: Chương 31, 32

Mức độ sâu: 4/5

Chapter 33 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc VPC SC & Service Perimeter — Cơ chế enforce ở API layer, access policy, perimeter types (regular vs bridge), protected resources, restricted VIP 199.36.153.4/30

  2. Access Levels & Context-Aware Access — Basic access levels (IP/geo/device), custom access levels với CEL, tích hợp VPC SC và private IP access levels

  3. Ingress & Egress Rules — Cross-Perimeter Access — Cấu trúc ingressFrom/ingressTo, egressFrom/egressTo, IAM roles trong rules, pattern cross-perimeter phổ biến

  4. Dry-Run Mode — Test Trước Khi Enforce — Shadow enforcement, audit log format, workflow triển khai an toàn từ dry-run đến enforce

  5. Organization Policy Framework & Inheritance — Constraint framework, managed vs legacy constraints, inheritance merge semantics, exceptions với resource tags

  6. Custom Constraints & CEL Expressions — YAML syntax, CEL expressions, defensive coding, ví dụ thực tế production

  7. Policy Troubleshooter & Debugging Violations — VPC SC violation analyzer, audit log analysis, Org Policy troubleshooter, systematic debugging workflow


Chương 34: Binary Authorization — Bảo Mật Container Deployment

Tại sao quan trọng: Binary Authorization đảm bảo chỉ những image đã được xác nhận và ký số mới được deploy. Hiểu cơ chế attestation, enforcement path và continuous validation là bắt buộc để tránh bypass vô tình và misconfiguration.

Điều kiện tiên quyết: Chương 10 (GKE Admission Control), Chương 12 (GKE Security), Chương 31 (IAM), container basics

Mức độ sâu: 4/5

Chapter 34 Full Index & Learning Paths

Các chủ đề con:

  1. Policy Model, Attestors và Deployment Decision Logic — Policy singleton per project, defaultAdmissionRule vs clusterAdmissionRules, evaluation modes (ALWAYS_ALLOW/ALWAYS_DENY/REQUIRE_ATTESTATION), enforcement modes (block vs dry-run), attestors như trusted verifier, Artifact Analysis note/occurrence model, AND-logic của multiple attestors, multi-project separation of duties

  2. Attestation: Cơ Chế Ký Số Mật Mã và Artifact Analysis — Attestation là Artifact Analysis occurrence, payload format (docker-manifest-digest + docker-reference), PKIX vs PGP signing, Cloud KMS integration, cross-project attestation lookup, Cloud Build provenance, SLSA provenance, tag vs digest confusion (attestation scope binding), key rotation pattern

  3. Enforcement Path: Từ kubectl apply Đến Admission Decision — ValidatingAdmissionWebhook vị trí trong pipeline, failurePolicy: Fail semantics, image extraction từ Pod spec (containers/initContainers/ephemeralContainers), tag resolution vấn đề, allowlist check, rule matching (cluster identity format), attestation lookup algorithm, image digest pinning security model (tag là mutable pointer), Cloud Build automated attestation workflow, audit logging enforcement events

  4. Continuous Validation: Kiểm Soát Compliance Cho Running Pods — Tại sao deploy-time enforcement không đủ, CV architecture qua Cloud Asset Inventory feed, evaluation cycle (≥24h), platform policies vs project-singleton policies, 6 loại checks (simple signing/image freshness/SLSA/Sigstore/trusted directory/vulnerability), check set scoping per namespace, CV chỉ log không evict, Artifact Analysis vulnerability scanning integration, alerting trên violations

  5. Break-glass, Policy Exceptions và Anti-patterns — Break-glass annotation cơ chế (bypass + mandatory audit), audit log structure cho break-glass events, monitoring và alerting break-glass abuse, allowlist patterns (glob matching, * vs **), globalPolicyEvaluationMode system images exemption, 6 anti-patterns phổ biến (dry-run permanent, allowlist quá rộng, một attestor, không monitor break-glass, tag thay vì digest, default ALWAYS_ALLOW), governance process


PHẦN VI: MESSAGING & DISTRIBUTED SYSTEMS


Chương 35: Cloud Pub/Sub — Kiến Trúc & Delivery Semantics

Tại sao quan trọng: Pub/Sub là messaging backbone của GCP. Hiểu delivery semantics, ordering, failure modes giúp tránh duplicate processing và message loss.

Điều kiện tiên quyết: Distributed systems fundamentals, messaging patterns

Mức độ sâu: 5/5

Chapter 35 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc Pub/Sub & Message Lifecycle — Pub/Sub không phải Kafka: storage per-subscription, message lifecycle publish→store→deliver→ack→delete, multi-zone replication, sharding, tại sao at-least-once là design choice

  2. Delivery Semantics & Ack Deadline — At-least-once (lease-based, nguồn gốc duplicate), exactly-once (deduplication state, regional constraint, pull-only), ack deadline extension với ModifyAckDeadline, latency và throughput tradeoffs

  3. Pull Subscriptions — Unary Pull & StreamingPull — Unary Pull (request-response, max 1000 msgs), StreamingPull (persistent gRPC stream, server-push), connection lifecycle và server-side reset, lease management, competing consumers

  4. Push Subscriptions — HTTP Endpoint & Retry Mechanics — Push delivery model, push window slow-start algorithm, exponential backoff 100ms–60s, JWT authentication, ordering constraint (1 outstanding per key), push vs pull decision matrix

  5. Ordering Keys — Per-Key Guarantee & Regional Scope — Cơ chế "only one batch outstanding per key", regional scope requirement, single-region endpoint bắt buộc, throughput limit 1 MBps per key, hot key problem và giải pháp, ordering + exactly-once tradeoff

  6. Dead Letter Topics — Trigger Conditions & Processing — Trigger sau max delivery attempts (5–100, approximate), message được wrap với metadata attributes, IAM permissions bắt buộc cho Pub/Sub service account, DLT subscription patterns, monitoring

  7. Flow Control & Backpressure — Dual-watermark model (maxOutstandingMessages + maxOutstandingBytes), flow control không giải quyết persistent overload, metrics quan trọng (num_undelivered_messages, oldest_unacked_message_age), KEDA autoscaling dựa trên backlog

  8. Schemas & Consumer Scaling Patterns — Avro vs Protocol Buffers schema types, validation tại publish time, schema evolution với revision model, fan-out (nhiều subscriptions), competing consumers (nhiều instances/subscription), KEDA Pub/Sub scaler


Chương 36: Pub/Sub Regional Failure Behavior — Cơ Chế Chịu Lỗi Vùng

Tại sao quan trọng: Pub/Sub có global SLA nhưng sự cố vùng (regional failure) vẫn có thể ảnh hưởng đến việc truyền tải thông điệp. Hiểu rõ hành vi chịu lỗi vùng giúp thiết kế các consumers kiên cường và chống mất mát dữ liệu.

Chapter 36 Full Index & Learning Paths

Các chủ đề con:

  1. Pub/Sub Storage Model & Message Storage Policies - Mô hình lưu trữ vật lý của Pub/Sub, sao chép đa zone đồng bộ, Message Storage Policies và tuân thủ dữ liệu chéo vùng.

  2. Regional Endpoints và Ordering Keys - Cơ chế ghim luồng gửi tin nhắn vào một vùng duy nhất, routing logic của client library và cách giải phóng khóa hàng gửi (ResumePublish) khi gặp lỗi.

  3. Tác Động Của Sự Cố Vùng - Phân tách hành vi control plane vs data plane khi xảy ra lỗi vùng, các message chưa ACK bị kẹt, bão phân phối lại (redelivery storm) và khôi phục thứ tự.

  4. Pub/Sub + Dataflow: Exactly-Once Processing - Kiến trúc kết hợp để đạt trạng thái xử lý chính xác một lần, State Checkpointing, cửa sổ khử trùng lặp và tính năng Native Exactly-Once.

  5. Khử Trùng Lặp và Subscriber Failover - Cơ chế quản lý lease (modifyAckDeadline), tranh chấp lease khi có nhiều instance tiêu thụ song song và cấu hình Flow Control an toàn.

  6. Giám Sát Vận Hành & Khôi Phục Dữ Liệu - Thiết lập chỉ số giám sát độ trễ và tỷ lệ lỗi để phát hiện lỗi vùng, kỹ thuật tua ngược thời gian (Seek to Timestamp) và phát lại dữ liệu lỗi.


Chương 37: Eventarc — Event Routing & CloudEvents

Tại sao quan trọng: Eventarc là managed event routing layer của GCP. Hiểu trigger model, CloudEvents standard, và Eventarc Advanced là điều kiện tiên quyết để thiết kế event-driven architecture đúng: tránh duplicate processing, hiểu delivery semantics, và governance cho multi-team event mesh.

Điều kiện tiên quyết: Chương 35 (Cloud Pub/Sub), CloudEvents spec basics

Mức độ sâu: 3/5

Chapter 37 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc Eventarc & Trigger Model — Trigger là gì về mặt kỹ thuật, Pub/Sub làm universal transport, hai đường ingest event (Audit Log và Direct), filtering mechanics, destinations (Cloud Run/GKE/Workflows), IAM requirements

  2. CloudEvents Format & Delivery Guarantees — CloudEvents spec, HTTP binary mode, context attributes, GCP extension attributes, at-least-once semantics, retry behavior, dead letter handling, idempotency implications

  3. Eventarc Advanced — Bus, Enrollment & Pipeline — Tại sao Advanced tồn tại (governance gap), bus architecture (Envoy-based, FGAC), enrollment với CEL filtering, pipeline transformation, format conversion, Standard vs Advanced decision matrix


Chương 38: Cloud Tasks — Thực Thi Công Việc Bất Đồng Bộ Tại Quy Mô Production

Tại sao quan trọng: Cloud Tasks là managed task queue service dành cho công việc bất đồng bộ quy mô lớn. Explicit invocation model (khác hoàn toàn so với Pub/Sub) yêu cầu hiểu: dispatch rate limiting (tránh thundering herd), at-least-once semantics (handler phải idempotent), exponential backoff logic (failure recovery), và task deduplication window (24 giờ).

Điều kiện tiên quyết: Distributed systems fundamentals, HTTP/REST API, Chapter 35 (Cloud Pub/Sub)

Mức độ sâu: 4/5

Chapter 38 Full Index & Learning Paths

Các chủ đề con:

  1. Mô Hình Bên Trong Cloud Tasks — Task queue model vs event model, task anatomy, queue structure, explicit execution control, state machine, lifecycle

  2. Rate Limiting & Dispatch Configuration — max_dispatches_per_second (QPS), max_burst_size, max_concurrent_dispatches, cách chúng tương tác, công thức effective rate, preventing thundering herd, multi-queue scaling pattern

  3. Retry & Exponential Backoff — Retry trigger conditions, exponential backoff formula, max attempts, min/max backoff, max doublings, jitter mechanics, SLA (retry window ~4 days), failure modes, configurable backoff patterns

  4. Task Deduplication & Idempotency — At-least-once delivery, why deduplication window 24 hours after deletion, task ID semantics, idempotent handler patterns (upsert, check-then-act, trace ID), database constraints

  5. HTTP Targets & Authentication — HTTP target endpoints (Cloud Run, App Engine, GKE, Compute Engine), OIDC token authentication, service account setup, header injection (headers NOT for auth), timeout behavior, error handling

  6. Cloud Tasks vs Pub/Sub Decision Matrix — Explicit vs implicit invocation, when to use each (single handler → Tasks, multiple handlers → Pub/Sub), rate control tradeoffs, scheduled tasks, message size limits, hybrid patterns

  7. Failure Handling & Observability — Silent dropout behavior (tasks exceed max attempts → deleted), dead letter queue setup, Cloud Logging integration, key metrics (retry rate, max attempts, task age), alert patterns, troubleshooting procedures

  8. Operational Patterns — Pause, Resume, Purge — Queue lifecycle (create, pause, resume, delete), incident response (pause queue during outage), purge patterns, multi-queue scaling, rate adjustment, monitoring dashboard, load testing

  9. GKE Integration & Service-to-Service Tasks — HTTP targets pointing to GKE services, workload identity setup (KSA + GSA binding), dispatching tasks to GKE pods, network exposure (LoadBalancer), OIDC token validation in handler, service-to-service patterns, performance considerations


PHẦN VII: OBSERVABILITY & RELIABILITY ENGINEERING


Chương 39: Cloud Monitoring — Metrics, Alerting, SLOs

Tại sao quan trọng: Cloud Monitoring là single pane of glass cho GCP. Hiểu metrics model, alerting mechanics, SLO framework để build reliable services và respond quickly.

Điều kiện tiên quyết: Chương 14, SRE fundamentals

Mức độ sâu: 5/5

Chapter 39 Full Index & Learning Paths

Các chủ đề con:

  1. Metrics Data Model — GAUGE, DELTA, CUMULATIVE — Time series data model, metric kinds và semantics, value types (INT64/DOUBLE/DISTRIBUTION), alignment period, reduction, label cardinality, failure modes của CUMULATIVE reset

  2. Monitored Resources & Billing Model — Monitored resource types schema, gke_container vs k8s_container, free vs chargeable metrics, GMP billing cardinality explosion, chiến lược kiểm soát cost

  3. Managed Service for Prometheus (GMP) — Collector DaemonSet node-local scraping, push to Monarch, PodMonitoring/ClusterPodMonitoring CRDs, Rule Evaluator, global query scope, PromQL interface, GMP vs self-hosted trade-offs

  4. Alerting Architecture — Policies, Conditions & Notification Channels — Alerting policy structure, 4 condition types (threshold/absence/MQL/PromQL), evaluation engine & retest window, 6.5 phút minimum alert latency, multi-condition logic, incident lifecycle, missing data handling, notification channel reliability, uptime checks

  5. SLO Framework — SLI, Error Budget & Burn Rate Alerting — Request-based vs window-based SLI, availability/latency/quality SLIs, rolling vs calendar compliance periods, error budget mechanics, burn rate math (2% budget/1h = 14.4x), multi-window alerting strategy, select_slo_burn_rate MQL, anti-patterns

  6. Dashboards, Phương Pháp Observability & Alert Best Practices — USE method (Utilization/Saturation/Errors), RED method (Rate/Errors/Duration), Four Golden Signals, dashboards as code với Terraform, alert design principles, alert fatigue reduction

  • Alert best practices: false positive reduction

Chương 40: Cloud Logging — Kiến Trúc, Định Tuyến & Quản Lý Chi Phí

Tại sao quan trọng: Chi phí Cloud Logging có thể tăng đột biến nếu không hiểu kiến trúc routing. Data Access Audit Logs mặc định bật cho BigQuery, VPC Flow Logs chưa được sampling có thể ngốn hàng trăm USD/tháng, và log duplication qua nhiều sinks tính phí nhân đôi. Hiểu Log Router pipeline, sự khác biệt giữa ingestion cost và storage cost, và cơ chế field exclusion là điều kiện tiên quyết để vận hành logging hiệu quả.

Điều kiện tiên quyết: Chương 39 (Cloud Monitoring)

Mức độ sâu: 4/5

Chapter 40 Full Index & Learning Paths

Các chủ đề con:

  1. Log Types & Ingestion Pipeline — Bốn nhóm log (Platform, User-written, Security/Audit, Component), cơ chế ingestion qua Cloud Logging API, timestamp validation, structured vs unstructured logging, log entry schema (LogEntry proto), auto-extraction fields, quota và rate limits

  2. Audit Logs Deep Dive — Admin Activity (miễn phí, không tắt được, vào _Required), Data Access (tính phí, mặc định tắt trừ BigQuery), System Event (miễn phí, tự động), Policy Denied (tính phí storage, có thể exclude); tại sao BigQuery là ngoại lệ, cách bật Data Access có chọn lọc per-service, tính toán volume trước khi bật, tích hợp Security Command Center

  3. Log Router & Sinks Architecture — Log Router như component logic tại mỗi resource level, hierarchy traversal (project→folder→org), _Required sink (không modify được, luôn nhận audit logs), _Default sink (có thể modify/disable), cơ chế sink evaluation song song, non-intercepting vs intercepting aggregated sinks, destinations (log bucket/BigQuery/GCS/Pub/Sub), sink errors và monitoring

  4. Retention, Storage & Cost Management — Cấu trúc chi phí hai tầng (ingestion cost khi ghi vào bucket vs storage cost sau 30 ngày), field exclusion trên bucket (trim fields trước khi ghi, giảm storage không giảm ingestion), _Required bucket 400 ngày miễn phí, custom bucket 1–3650 ngày, Log Views cho access control, chiến lược giảm chi phí thực tế (VPC Flow Logs sampling/aggregation, container log severity filter, load balancer log sampling, tránh duplication)

  5. Log-based Metrics & Alerting — Counter metrics (đếm matching entries mỗi 60 giây) và distribution metrics (extract numeric value, histogram), timing model (update mỗi phút, data gap khi không có entries), labels và cardinality constraints, system-defined vs user-defined metrics, log-based metric alerting vs direct log alerting, tích hợp SLO framework

  6. Logging Query Language & Debug Patterns — LQL syntax: field paths, toán tử (= equality vs : contains), boolean logic, hàm đặc biệt (log_id(), source(), sample(), SEARCH()), indexed vs non-indexed fields; debug patterns: trace ID correlation, timestamp-based correlation, GKE pod debugging, audit trail forensics, OOM kill diagnosis, Log Analytics SQL cho complex queries


Chương 41: Cloud Trace, Profiler, và Error Reporting — Chẩn đoán hệ thống phân tán

Tại sao quan trọng: Distributed tracing, continuous profiling, và error aggregation là ba công cụ chuyên biệt để chẩn đoán latency, performance bottleneck, và error epidemic trong microservices. Trace cho biết khi nào có vấn đề, profiler cho biết tại sao, error reporting cho biết tần suấtphạm vi. Ba công cụ này phải hoạt động như một hệ thống unified observability.

Điều kiện tiên quyết: Chương 39, 40, OpenTelemetry basics, distributed systems fundamentals

Mức độ sâu: 3/5

Chapter 41 Full Index & Learning Paths

Các chủ đề con:

  1. Cloud Trace Fundamentals — Cơ chế Distributed Tracing — Trace ID generation, span lifecycle, parent-child relationships, trace collection architecture, OTLP protocol, batching & compression, head-based vs tail-based sampling, sampling decision framework, retention & quotas, cost model, production patterns

  2. Trace Propagation & W3C Standards — Context Di chuyển — W3C Trace Context headers (traceparent, tracestate), X-Cloud-Trace-Context legacy format, HTTP/gRPC/message queue propagation, baggage mechanism, clock skew detection, context propagation failures, cross-origin requests, third-party API integration, async workflow correlation

  3. OpenTelemetry Integration — Instrumentation & OTLP Export — OTEL architecture, TracerProvider, sampler types, span processors, OTLP protocol (gRPC vs HTTP), compression, auto-instrumentation vs manual instrumentation, GCP Cloud Trace exporter, authentication, performance overhead, batch configuration, failure handling

  4. Cloud Profiler: Continuous Profiling — CPU/Heap/Goroutine Profiles — Statistical sampling mechanism, CPU profiling (on-CPU time), heap profiling (allocation tracking), goroutine profiling (concurrency analysis), contention profiling (lock waiting), flame graph interpretation, profiler agent architecture, profile selection, overhead analysis (<1%), production patterns, constraints & limitations

  5. Error Reporting — Grouping, Fingerprinting, Notifications — Fingerprint generation (error type + stack trace + service), error group lifecycle (Open/Acknowledged/Resolved/Muted), affected user detection, error events & sampling (1000 sample limit), error deduplication, notifications (email/Slack/PagerDuty), issue tracker integration, error budget tracking

  6. Traces, Logs, Metrics — Correlation & Investigation Workflows — Three pillars of observability, correlation through trace context, request IDs vs trace IDs, query patterns (latency spike → error rate → root cause), structured logging, baggage for business context, observability as system, sampling coordination, cost optimization

Chap 42: SRE Practices trên Google Cloud Platform

Tại sao quan trọng: SRE principles aplikasi di GCP memerlukan understanding both people dan systems. Error budgets, incident response, toil reduction adalah practical skills cho production reliability.

Điều kiện tiên quyết: Chap 39–41, SRE Book concepts, distributed systems fundamentals

Mức độ sâu: 5/5

Chapter 42 Full Index & Learning Paths

Các chủ đề con:

  1. SLI, SLO, Error Budget — Định lượng độ tin cậy - SLI measurement, SLO target, error budget consumption, compliance periods, rolling vs calendar

  2. Error Budget Policy — Quản lý risk deployment - Policy thresholds, feature freeze vs code freeze, automated enforcement, communication, quarterly review

  3. Incident Response — Levels, Roles, Runbooks - Severity levels (L1-L4), incident commander, SME, communication lead, on-call rotation, machine-readable runbooks

  4. Postmortems — Blameless Culture & Root Cause Analysis - Swiss cheese model, 5 Whys, timeline analysis, action items, postmortem timing, measuring effectiveness

  5. Resilience Patterns — Graceful Degradation, Canaries, Circuit Breakers - Graceful degradation tiers, canary deployment (5%-25%-50%-100%), circuit breaker states, Istio integration, anti-patterns

  6. Timeout & Retry Management — Ngăn chặn cascading failures - Timeout hierarchies, exponential backoff, jitter, retry budget, idempotency, application/service mesh implementation

  7. Load Shedding & Request Prioritization - Server-side rejection, priority-based shedding, probabilistic shedding, upstream shedding, Cloud Load Balancer rate limiting

Chap 43: GKE Production Debugging Methodology — Structured Approach to Troubleshooting

Tại sao quan trọng: Debugging production issues requires systematic approach. "Random kubectl exec" là antipattern. Build mental model để structured debugging, understand failure cascade, correlate across layers.

Điều kiện tiên quyết: Semua GKE chapters (5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18)

Mức độ sâu: 5/5

Chapter 43 Full Index & Learning Paths

Các chủ đề con:

  1. Pod Lifecycle Debugging — Pod status state machine, kubelet restart policy, container states (waiting/running/terminated), debugging Pending pods (scheduling vs image pull vs volume), CrashLoopBackOff (exit codes, restart backoff logic), OOMKilled (container limit vs node pressure), init container failures, resource request tuning, probe configuration, GKE-specific crictl debugging

  2. Service Connectivity Debugging — DNS resolution architecture (kube-dns vs Cloud DNS), ClusterIP virtual routing (iptables vs eBPF), load balancing mechanics, NetworkPolicy enforcement, service endpoint discovery, cross-node connectivity, MTU issues, firewall rules, port-forward as debugging tool, packet capture analysis

  3. Node Debugging — Node health conditions (Ready, MemoryPressure, DiskPressure, PIDPressure), kubelet heartbeat mechanism, container runtime (containerd) debugging, system resource monitoring (memory, disk, CPU, PID), node taint and drain operations, GCP Compute Engine instance issues, zone maintenance, node auto-repair, crictl tools

  4. Control Plane Debugging — API server request path (auth → authz → admission → etcd), etcd performance and slow operations, scheduler latency (filtering/scoring), admission webhook timeouts, webhook configuration debugging, GKE managed control plane metrics, Cloud Audit logs, cluster control plane status checks

  5. Cross-Cutting Debugging — Distributed tracing architecture (trace/span model), trace ID propagation, Cloud Trace integration with GKE, logs + metrics + traces correlation, request path debugging (client → LB → pod), end-to-end timeline analysis, OpenTelemetry instrumentation, service mesh auto-tracing (Istio), error recording in traces, SLO definition via latency percentiles

  • Cross-cutting debugging:
    • Request tracing: LB → node → pod dengan Cloud Trace
    • Correlate: access logs + app logs + traces
    • gcloud container operations list

PHẦN VIII: PLATFORM AUTOMATION & CI/CD


Chương 44: Terraform & CI/CD Platform trên GCP

Tại sao quan trọng: Infrastructure as Code là non-negotiable cho production. Hiểu GCP-specific Terraform patterns, state management, Cloud Build internals, Artifact Registry và SLSA attestation để ngăn chặn state drift, destructive applies và supply chain attacks.

Điều kiện tiên quyết: Chương 31 (IAM), Chương 34 (Binary Authorization), Terraform fundamentals, Docker/containers

Mức độ sâu: 4/5

Chapter 44 Full Index & Learning Paths

Các chủ đề con:

  1. Terraform State với GCS Backend — Cơ chế lưu state dưới dạng GCS objects, state locking qua conditional writes (x-goog-if-generation-match), encryption options (Google-managed/CMEK/CSEK), workspace vs directory separation, bucket configuration, orphaned lock recovery, state migration bootstrap workflow

  2. Terraform Modules & IaC Patterns cho GCP — Module resolution và scope, terraform-google-modules ecosystem (network/gke/project-factory), project factory pattern, landing zone structure, remote state references, directory-based environment management, version pinning, anti-patterns (thin wrapper, hardcode env, output leakage)

  3. Terraform Security trên GCP — Ba mô hình service account (JSON key anti-pattern, impersonation, Workload Identity), least privilege design theo layer, sensitive state management, CMEK cho state bucket, gcloud terraform vet, scheduled drift detection, separation of duties trong CI/CD pipelines

  4. Cloud Build Internals — Triggers, Worker Pools, Caching — Ephemeral worker lifecycle, build step execution và waitFor parallelism, trigger types (CSR/GitHub/GitLab/Pub/Sub/webhook/scheduled), filter mechanics, private worker pools (VPC peering, NO_PUBLIC_EGRESS, CIDR sizing), kaniko layer caching, GCS custom cache, substitution variables (built-in/custom/dynamic), org policy allowedIntegrations

  5. Artifact Registry — Repository Model & Vulnerability Scanning — Repository types và modes (standard/remote/virtual), regional storage architecture, Artifact Analysis vulnerability scanning mechanism (on-push continuous vs on-demand 48h), SBOM generation, cleanup policies (delete/keep/most-recent), IAM model cấp repository, VPC Service Controls integration

  6. SLSA Attestation & Binary Authorization Integration — SLSA framework L1-L3, Cloud Build SLSA L3 provenance generation, in-toto Statement format, DSSE signing với Google-managed key (non-forgeable), Binary Authorization policy verification flow, custom attestors (vulnerability scan gate), end-to-end attestation workflow, limitations (images field vs explicit push, private pool vs hosted)


Chương 45: Cloud Build & Artifact Registry — CI/CD Pipeline

Tại sao quan trọng: Secure CI/CD pipeline là critical security control. Understand Cloud Build execution model và Artifact Registry security prevent supply chain attacks.

Điều kiện tiên quyết: Chương 32, Docker/containers, CI/CD fundamentals

Mức độ sâu: 4/5

Chapter 45 Full Index & Learning Paths

Các chủ đề con:

  1. Cloud Build: Execution Model & Architecture — Build steps, workers, ephemeral execution environment, build network isolation, workspace model, build metadata, execution guarantees, failure modes

  2. Service Accounts & Security: Least Privilege — Default service account, custom service accounts, least privilege, IAM roles, per-step service account binding, cross-project deployments, service account audit

  3. Private Worker Pools: Network Isolation & VPC Integration — Service producer network, VPC peering model, static IP ranges, firewall configuration, private pool architecture, trade-offs with default pools

  4. Build Triggers & Configuration Management — Build trigger types (CSR, GitHub, GitLab, scheduled, webhook), cloudbuild.yaml schema, build steps, substitution variables, build options, caching strategies

  5. Artifact Registry: Architecture & Repository Types — Repository types (Docker, Maven, npm, Python, Go, APT/YUM, generic), regional storage, access control, cost model, virtual repositories, remote repositories

  6. Container Vulnerability Scanning & SBOM — On-push scanning, continuous scanning, SBOM generation, cleanup policies, protection rules, artifact retention, continuous monitoring

  7. Build Provenance & SLSA Attestation — SLSA framework levels, build provenance format (v1.0), SLSA Level 3 requirements, provenance verification, attestation, build audit trail

  8. Cloud Deploy & Managed Delivery Strategies — Delivery pipeline model, deployment strategies (all-at-once, canary, blue-green), approval gates, manual review, rollback mechanism, Cloud Build integration

  9. Organization Policies & Supply Chain Security — allowedIntegrations policy, Cloud Build integration points, service account policies, supply chain audit trail, security checklist, common misconfigurations


Chap 46: Cloud Deploy & GitOps — Progressive Delivery

Tại sao quan trọng: Cloud Deploy cung cấp managed CD với built-in approval, rollback, audit trail. Config Sync implement GitOps model: continuous reconciliation thay vì push-based deployment. Hiểu cơ chế bên trong để thiết kế safe deployment pipelines.

Điều kiện tiên quyết: Chap 45, Kubernetes Deployments

Mức độ sâu: 4/5

Chapter 46 Full Index & Learning Paths

Các chủ đề con:

  1. Cloud Deploy Internal Model — Pipeline, Targets, Releases, Rollouts - Resource hierarchy: DeliveryPipeline → Target → Release → Rollout → Phase → Job → JobRun; render phase, deploy phase, GCS artifact storage, promotion mechanics, failure modes

  2. Skaffold Integration — Render & Deploy Pipeline - Tại sao Skaffold tồn tại trong Cloud Deploy, render phase (skaffold diagnose + render), deploy phase (skaffold apply), profiles cho multi-environment, Helm/Kustomize rendering, execution environment

  3. Progressive Delivery Strategies — Canary & Blue-Green - Canary: 3 loại (automated/custom-automated/custom), traffic splitting mechanics GKE (Service Networking vs Gateway API) và Cloud Run, phases, analysis; Blue-green: parallel environments, cutover, rollback; khi nào dùng chiến lược nào

  4. Verification, Hooks & Cloud Deploy Automation - Pre/post-deploy hooks, execution environments (Cloud Build vs kubernetesCluster), verify jobs, idempotency requirement; Automation: promote, canary advance, rollback repair; AutomationRun lifecycle

  5. Governance — Approval, Rollback & Deploy Policies - IAM role separation (deployer vs approver), approval flows qua Pub/Sub, rollback mechanics (tạo Rollout mới với Release cũ), deploy policies (time-based freeze windows, action restrictions), notifications, audit trail

  6. Config Sync — GitOps Engine Bên Trong - GitOps vs CD pipeline, kiến trúc Config Sync (Reconciler Manager, per-sync Reconcilers, ResourceGroup Controller, Admission Webhook), reconciler pipeline (fetch→render→parse→apply), RootSync vs RepoSync (scope, permissions), sources (Git/OCI/Helm), drift detection (reactive vs proactive)

  7. Multi-Cluster GitOps & Policy Controller - Repository structure patterns (monorepo/Kustomize overlays/multi-repo delegation), Fleet integration, Policy Controller với GitOps (ConstraintTemplates qua Config Sync, policy bundles), namespace sameness, Cloud Deploy + Config Sync phân chia concerns


PHẦN IX: ADVANCED PRODUCTION PATTERNS


Chap 47: GKE Service Mesh — Cloud Service Mesh (Managed Istio)

Tại sao quan trọng: Service mesh provides mTLS, observability, traffic management at infrastructure level. Hiểu Istio/Envoy mechanics để debug connection failures và tune performance.

Điều kiện tiên quyết: Chap 7, 21, microservices patterns

Mức độ sâu: 4/5

Chapter 47 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến Trúc CSM — Managed Istio, Data Plane vs Control Plane - Cloud Service Mesh lifecycle, Istiod components (Pilot/Citadel/Galley), Envoy sidecar role, GKE integration model, CSM vs self-managed Istio

  2. Sidecar Injection — Init Container, iptables Interception - MutatingWebhook injection mechanics, istio-init container iptables REDIRECT rules, traffic interception path, CNI plugin alternative, troubleshooting injection failures

  3. mTLS & SPIFFE/SVID — Zero-Trust Service Identity - SPIFFE/SVID identity model, Istiod CA certificate lifecycle, PERMISSIVE vs STRICT PeerAuthentication, AuthorizationPolicy, mTLS migration strategy, debugging cert failures

  4. Traffic Management — VirtualService, DestinationRule, Circuit Breaking - Istio traffic management API, weighted canary routing, retry/timeout policies, circuit breaker với Envoy outlier detection, ServiceEntry, Gateway ingress/egress

  5. Envoy xDS API — Cách Istiod Push Config - xDS protocol (CDS/EDS/LDS/RDS/SDS), SotW vs Delta xDS, push model, config ACK/NACK, debugging với istioctl proxy-config, scaling xDS

  6. Observability & Distributed Tracing - Istio RED metrics, response_flags decode, Envoy access logs, trace header propagation (B3/W3C), sampling strategies, CSM dashboard SLOs, Cloud Trace integration

  7. Sidecar Performance — Overhead, Latency, Resource Tuning - CPU/memory overhead thực tế, latency per-hop analysis, concurrency tuning, connection pool sizing, ambient mesh (sidecarless) khi nào cần


Chap 48: Multi-Cluster Architecture & Networking

Tại sao quan trọng: Multi-cluster là standard pattern để production GKE deployments. Networking across clusters adds complexity requiring specific patterns.

Điều kiện tiên quyết: Chap 21, 47, Chap 18 Fleet

Mức độ sâu: 4/5

Chapter 48 Full Index & Learning Paths

Các chủ đề con:

  1. Multi-Cluster Use Cases & Deployment Models - HA/DR, geo-distribution, scale beyond limits, deployment models (Active-Active/Active-Passive/Hub-Spoke/Mesh)

  2. Multi-Cluster Services (MCS) Deep Dive - ServiceExport/ServiceImport, DNS integration, endpoint propagation, locality-aware load balancing, eventual consistency

  3. Multi-Cluster Ingress (MCI) — Global Load Balancing - Config cluster pattern, NEGs, global HTTP(S) load balancer, health checking cross-cluster, traffic routing

  4. Gateway API cho Multi-Cluster Routing - Gateway API abstraction, HTTPRoute, cross-cluster backends, traffic splitting, advanced routing

  5. Workload Identity Across Clusters - Identity sameness, principal definitions, security implications, service account impersonation

  6. Cross-Cluster Service Mesh & Trust Federation - mTLS federation, certificate distribution, SPIFFE identity, authorization policies

  7. Network Isolation & VPC Peering - VPC peering topology, non-transitivity, firewall rules, route exchange, mesh vs hub-spoke

  8. DNS Peering Across Clusters - Cloud DNS peering, per-cluster zones, service discovery, failover patterns, TTL implications


Chap 49: GKE AI/ML Infrastructure — GPU, TPU, Large-Scale Workloads

Tại sao quan trọng: AI/ML là dominant workload pattern. GPU/TPU infrastructure có unique characteristics để maximize utilization và minimize cost.

Điều kiện tiên quyết: Chap 6, 9, GPU/accelerator fundamentals

Mức độ sâu: 4/5

Chapter 49 Full Index & Learning Paths

Các chủ đề con:

  1. GPU Node Pool Architecture & Device Plugin — GPU node pool creation, NVIDIA device plugin mechanism, GKE-managed vs user-managed drivers, taints/tolerations for GPU workloads, GPU quota and zone constraints, driver installation and lifecycle

  2. GPU Memory Strategies - MIG, Time-Slicing, Tradeoffs — Multi-Instance GPU hardware partitioning, GPU time-slicing with instruction-level preemption, NVIDIA MPS (legacy), memory isolation gaps, isolation vs utilization tradeoffs, scheduling fragmentation, composite MIG+time-slicing strategies

  3. TPU Architecture - Types, Topology, Multi-host Model — TPU v4/v5p/v6e/Ironwood generations, 3D topology constraints (AxBxC), multi-host vs single-host slices, atomic scheduling, inter-chip interconnect (ICI) latency, pod-to-TPU mapping rules, failure modes and recovery

  4. Gang Scheduling & Batch Reservation - ProvisioningRequest, Kueue — Why gang scheduling matters for distributed training, ProvisioningRequest API for atomic resource allocation, Kueue job queueing system, Dynamic Workload Scheduler, cost implications, failure modes

  5. GPU Inter-GPU Communication - NCCL Fast Socket, GPUDirect, RDMA — NCCL collective operations, NCCL Fast Socket with Andromeda optimization, GPUDirect-TCPX/TCPXO for A3 High/Mega, GPUDirect RDMA for A3 Ultra/A4, network scaling laws, performance bottlenecks at scale

  6. High-Performance Networking - InfiniBand, A3 Clusters, Compact Placement, H4D — A3 machine series with GPUDirect networking, sub-microsecond latency requirements, compact placement policy for topology-aware scheduling, H4D HPC-optimized machines, multi-NIC Pods, Dataplane V2 requirement

  7. Data Loading Optimization - Hyperdisk ML, Parallelstore, Volume Populator — Data loading as critical path for AI/ML, Hyperdisk ML for model weights, Parallelstore for distributed filesystem with sub-ms latency, GKE Volume Populator for automatic data staging, cost vs performance tradeoffs

  8. LLM Serving Patterns - vLLM, TGI, Triton, Optimization — vLLM with PagedAttention and continuous batching, Text Generation Inference (TGI), NVIDIA Triton with TensorRT-LLM, tensor parallelism strategies, batch size tuning for latency vs throughput, custom metrics for autoscaling

  9. Dynamic Resource Allocation (DRA) - Next-Gen GPU Scheduling — DRA vs device plugins, DeviceClass blueprint for hardware categories, ResourceClaim/ResourceClaimTemplate flexible filtering, ResourceSlice device inventory, DRANET for networking resource allocation, migration from device plugins

  10. Cost Optimization - Spot VMs, Preemption Handling, Billing Models — Spot VMs up to 91% discount with arbitrary preemption window, graceful shutdown and checkpoint/resume patterns, Pod disruption budgets, hybrid on-demand+Spot strategies, cost modeling for batch vs serving workloads


Chap 50: GKE Large-Scale Design — 1000+ Nodes

Tại sao quan trọng: GKE clusters > 1000 nodes memiliki operational characteristics berbeda. Architecture decisions at creation time mempengaruhi scalability ceiling.

Điều kiện tiên quyết: Chap 5, 6, 8, 9

Mức độ sâu: 5/5

Chapter 50 Full Index & Learning Paths

Các chủ đề con:

  1. GKE Scalability Limits — Hard Limits at Scale - API server, etcd, scheduler, networking, controller manager limits. Layers và bottleneck analysis. Approaching limits safely.

  2. Node Pool Planning — Sizing, Density, Optimization - Node pool mechanics, Pod density constraints, resource planning, bin-packing strategies, sizing formula.

  3. API Server at Scale — Request Routing, Watch Connections, Latency - Streaming list response encoding, watch bottleneck, latency SLOs, tuning parameters.

  4. etcd Scalability — Object Storage, Compaction, Spanner Migration - Revision history, compaction mechanics, Spanner-based etcd, performance optimization.

  5. Controller Manager Scalability — Work Queues, Workers, Reconciliation - Reconciliation loop, work queue depth, controller latency, optimization patterns.

  6. Scheduler Performance at Large Scale — Algorithm, Latency, Topology - Predicate evaluation, latency curves, Topology Aware Scheduling, scheduling optimization.

  7. IP Planning for Large Scale — CIDR Sizing, Pod Density, Exhaustion - VPC secondary CIDR, pod IP allocation, subnet sizing, IP exhaustion prevention.

  8. Service Mesh Scalability — xDS Protocol, Sidecar Overhead, Endpoint Limits - xDS push model, sidecar memory overhead, 260K endpoint limit, control plane scaling.

  9. NodeLocal DNSCache — DNS Performance, Caching Strategy, When Mandatory - DNS bottleneck, CoreDNS caching, when mandatory, operational considerations.

  10. Node Pool Strategy — Single vs Multiple Pools, Blast Radius, Flexibility - Single pool vs multiple pools, trade-offs, blast radius analysis, hybrid approach.

  11. Workload Distribution — Topology Spread, Affinity, Bin-Packing - Pod placement strategy, TSC constraints, anti-affinity, bin-packing optimization.

  12. Large-Scale Upgrade — Surge Sizing, Concurrency, Disruption Budget - Upgrade mechanics, surge nodes, PDB constraints, upgrade strategy.

  13. Network Policy Scalability — eBPF Limits, Policy Compilation, Endpoint Cardinality - eBPF map limits, policy compilation overhead, endpoint cardinality.

  14. Metrics Cardinality Management — Cardinality Explosion, High-Cardinality Patterns, Cost Control - Cardinality sources, limit enforcement, scrape filtering, cost optimization.

  15. Logging at Scale — Log Volume, Sampling, Exclusion, Cost Management - Log volume calculation, exclusion strategies, sampling, cost control.


Chap 51: Cost Optimization Engineering — Systematic Approach

Chapter 51 Full Index & Learning Paths

Điều kiện tiên quyết: Semua service chapters

Mức độ sâu: 3/5

Các chủ đề con:

  1. Committed Use Discounts (CUDs): Resource-based vs Flexible Spend
  2. Sustained Use Discounts: Automatic Savings for GCE
  3. Spot VMs: Preemption Mechanics, Pricing, Workload Fit
  4. GKE Cost Allocation: Namespace-level Breakdown & Chargeback
  5. Rightsizing: Machine Type Selection, VPA Recommendations, Insights
  6. Idle Resource Detection: Recommender API & Automation
  7. Egress Costs: Inter-region Pricing, Optimization Strategies
  8. Storage Tier Automation: Lifecycle Policies & Archive Storage
  9. Cloud Billing Exports: BigQuery Analysis & Anomaly Detection
  10. Budget Alerts: Programmatic Controls & Automation
  11. Cost Monitoring Dashboards & FinOps Patterns

Chap 52: Disaster Recovery — Kiến Trúc, Chiến Lược và Thực Thi

Tại sao quan trọng: DR planning cho production systems là critical. GCP provides bao nhiêu options với different cost/complexity tradeoffs. RTO/RPO targets quyết định entire architecture.

Chapter 52 Full Index & Learning Paths

Điều kiện tiên quyết: Chap 9, 28–30 (storage fundamentals)

Mức độ sâu: 4/5

Các chủ đề con:

  1. RTO vs RPO: Định Nghĩa, Trade-offs, và Cost Implications — Mục tiêu phục hồi thực sự ý nghĩa gì, vì sao chúng quyết định kiến trúc toàn bộ, relationship giữa RTO/RPO và operational cost.

  2. Multi-Region Architecture: Active-Active vs Active-Passive — Sự khác biệt cơ bản, distributed consensus problem, split-brain scenarios, data consistency challenges, vì sao active-active không phải luôn là câu trả lời tốt nhất.

  3. Backup Strategies for GKE: Concepts, Architecture, Restore Workflows — Backup for GKE internals, incremental backup mechanics, snapshot integration, cross-cluster restoration, disaster scenario workflows.

  4. Persistent Disk Snapshots & Cross-Region Replication — Snapshot mechanics (copy-on-write), incremental snapshot chains, async replication, RPO guarantees, recovery procedures.

  5. Cloud Storage Geo-Redundancy: Multi-Region, Dual-Region, Turbo Replication — Replication mechanisms, eventual consistency windows, Turbo Replication RPO guarantees, cross-bucket replication patterns.

  6. Database Disaster Recovery: Cloud SQL Replicas & Spanner Global Instances — High availability architecture, read replicas vs HA replicas, replication lag, failover mechanics, Spanner multi-region topology.

  7. Configuration Backup: GitOps, Terraform State, Resource Manifests — Why configuration is critical to DR, GitOps as recovery source-of-truth, Terraform state replication, secret management in recovery.

  8. DNS Failover: Health Checks, Weighted Routing, Recovery Procedures — Health check mechanisms, TTL impact on failover latency, global load balancing failover, DNS propagation realities.

  9. DR Testing, Chaos Engineering, and Failure Validation — Regular recovery drills methodology, chaos at region level, synthetic monitoring, recovery metrics validation.

  10. DR Runbooks, Incident Response, and Step-by-Step Procedures — Runbook structure, incident escalation, decision trees, automation boundaries, post-recovery validation.


PHẦN X: ADVANCED DEBUGGING & INCIDENT MANAGEMENT


Chap 53: Production GKE Debugging Framework

Chapter 53 Full Index & Learning Paths

Tại sao quan trọng: Framework debugging toàn diện cho GKE production — từ hypothesis-driven methodology, telemetry model, pod/service/node/control-plane debugging cho đến cross-layer correlation, incident timeline reconstruction và root cause analysis.

Điều kiện tiên quyết: Chap 39–43

Mức độ sâu: 5/5

  1. Phương Pháp Debugging Khoa Học & Mô Hình Telemetry GKE — Hypothesis-driven debugging (hypothesis → test → validate), internal model của Kubernetes Events, Cloud Logging, Cloud Monitoring, Cloud Trace, signal selection matrix và chiến lược thu thập thông tin.

  2. Pod Debugging: Pending, Crashed & Hung States — Scheduling pipeline và filter predicates, exit code taxonomy, CrashLoopBackOff exponential backoff, OOM kill mechanics với cgroup v1 vs v2, hung pod với liveness/readiness probe semantics, ImagePullBackOff.

  3. Service Connectivity Debugging: DNS, Routing & Network Policy — CoreDNS resolution pipeline, ndots problem, NodeLocal DNSCache, iptables vs eBPF (GKE Dataplane V2), conntrack exhaustion, Network Policy enforcement mechanics, 6-step connectivity debugging.

  4. Infrastructure Debugging: Node & Control Plane — Kubelet heartbeat model (kube-node-lease), node conditions và eviction thresholds, PLEG stall, node-problem-detector, API server request lifecycle, etcd watch fan-out, admission webhook timeouts, giới hạn GKE managed control plane.

  5. Cross-Layer Correlation, GKE Dashboard & Incident Timeline — Clock skew và sampling bias, trace context propagation, Cloud Logging tích hợp với traces, đọc GKE Observability Dashboard signals, xây dựng incident timeline từ nhiều signal source, phân biệt causality vs correlation.

  6. Root Cause Analysis: Five Whys, Fishbone & Action Planning — Five Whys mechanics và giới hạn trong distributed systems, Ishikawa diagram categories cho Kubernetes, proximate vs root cause, framework action items 3 tầng (immediate/tactical/strategic), detection gap analysis, blameless postmortem culture.


Chap 54: Incident Response & Post-Mortems

Chapter 54 Full Index & Learning Paths

Tại sao quan trọng: Incident response là kỹ năng phân biệt SRE giỏi với SRE xuất sắc. Một quy trình có cấu trúc — từ severity classification, incident command system, đến mitigation engineering và blameless postmortem — giúp giảm MTTR và giới hạn blast radius một cách có hệ thống, thay vì phụ thuộc vào phản xạ cá nhân dưới áp lực.

Điều kiện tiên quyết: Chap 39–43, Chap 53

Mức độ sâu: 4/5

  1. Phân Loại Severity & Incident Command System — Severity như cơ chế kích hoạt (không phải nhãn mô tả), static/dynamic severity qua Cloud Monitoring MQL, bốn vai trò Incident Commander/Operations Lead/Communications Lead/Planning Lead, live incident document, ngưỡng escalate và cơ chế handoff.

  2. Vòng Đời Incident: Detection → Triage → Mitigation → Resolution — Incident như state machine với exit criteria từng trạng thái, alerting symptom-based, funnel triage Cloud Monitoring → Cloud Logging → Cloud Trace/Error Reporting, nguyên tắc không trì hoãn mitigation để chờ root-cause đầy đủ.

  3. Mitigation Engineering: Rollback, Feature Flag, Circuit Breaker — Ba tầng mitigation khác nhau: rollback (Cloud Deploy/Cloud Run — tầng deployment artifact), feature flag/kill switch (Firebase Remote Config — tầng runtime configuration), circuit breaker (Cloud Service Mesh outlier detection — tầng network call), giới hạn thật của từng cơ chế.

  4. Incident Communication: Status Page & Stakeholder Management — Kênh nội bộ vs kênh bên ngoài, case study Personalized Service Health/Cloud Service Health của Google Cloud, cấu trúc bản cập nhật stakeholder, vì sao im lặng gây hoảng loạn hơn tin xấu.

  5. Blameless Postmortem: Timeline Reconstruction & Root Cause Analysis — Blameless như cơ chế thu thập thông tin trung thực, tiêu chí bắt buộc viết postmortem, cấu trúc 5 bước (create → capture facts → analyze root cause → plan → execute), timeline reconstruction ở tầng tổ chức kết hợp system event với quyết định con người.

  6. Action Item Governance & Knowledge Sharing — Vì sao action item chết trong document nhưng sống trong issue tracker, ownership và diffusion of responsibility, neo priority vào error budget policy, incident readout và runbook như living document.


PHẦN XI: SPECIAL TOPICS & ADVANCED CONCEPTS


Chap 55: Kubernetes API Machinery Deep Dive

Tại sao quan trọng: Hiểu cơ chế bên trong API server, informer pattern, watch mechanism là điều kiện tiên quyết để debug các hành vi phức tạp của control plane mà tài liệu vận hành thông thường không giải thích tới.

Điều kiện tiên quyết: Chap 5

Mức độ sâu: 5/5

Chapter 55 Full Index & Learning Paths

Các chủ đề con:

  1. API Server Request Pipeline: Authentication → Authorization → Admission → Storage — Authenticator chain và union semantics, authorizer chain fail-secure, admission hai pha (mutating/validating), reinvocation policy, AdmissionReview protocol, conversion và optimistic concurrency ở storage layer

  2. Watch Mechanism: Lan Truyền State Hiệu Quả Không Cần Polling — Giao thức list-watch, ngữ nghĩa resourceVersion theo từng loại request, watch bookmark, chunking limit/continue, lỗi 410 Gone và cơ chế relist

  3. Informer Pattern: List-Watch, Local Cache, Resync Intervals — Kiến trúc Reflector, DeltaFIFO, Indexer/ThreadSafeStore, SharedInformerFactory, cơ chế resync tự sửa sai, và vì sao mọi read path nên qua local cache

  4. Controller-Runtime Framework: Reconciliation Loop Patterns — Manager, split-client (cache read/direct write), Reconciler interface, For/Owns/Watches, Predicate filtering, rate-limited workqueue, leader election

  5. API Priority and Fairness: Flow Schema, Priority Level, Seat & Queue — Nominal concurrency shares, seat calculation cho list/watch request, shuffle sharding vào hàng đợi, hành vi 429 khi hàng đợi đầy

  6. Etcd Consistency: Linearizability, MVCC & Watch Caching — Raft consensus, linearizable vs serializable read, mô hình revision MVCC, compaction, và watch cache của API server xếp chồng lên watch stream etcd

  7. Resource Versioning: Optimistic Locking & Conflict Resolution — resourceVersion như optimistic concurrency token, 409 Conflict, JSON Patch/Merge Patch/Strategic Merge Patch, Server-Side Apply field ownership

  8. Custom Resource Definitions: Cơ Chế Mở Rộng API Server — Dynamic REST endpoint, structural schema bắt buộc, pruning/defaulting, versioning và conversion webhook, subresource status/scale, CEL validation, so sánh với aggregated API server


Chap 56: Kubernetes Advanced RBAC & Authorization Patterns

Tại sao quan trọng: RBAC design cho production scale requires careful planning. Understand aggregation, impersonation, conditions prevent privilege creep.

Điều kiện tiên quyết: Chap 12, Chap 31

Mức độ sâu: 4/5

Chapter 56 Full Index & Learning Paths

Các chủ đề con:

  1. RBAC Authorization Pipeline & Mental Model - Authorization decision rendering, fail-secure defaults, privilege escalation prevention, authorization modes chain

  2. ClusterRole Aggregation: Composing Roles - Label-based role composition, immutability constraints, dynamic role merging, scaling implications

  3. ClusterRoleBinding vs RoleBinding Semantics - Scope semantics, namespace override behavior, binding immutability, verb mappings, subject matching

  4. Service Account Impersonation & Delegation - Impersonate verb, delegation chains, privilege escalation prevention, audit implications, safe patterns

  5. Group Binding Strategies & Semantics - System groups, custom groups, LDAP/OIDC integration, group membership semantics, delegation via groups

  6. RBAC Conditions: Attribute-Based Access Control - CEL expressions, condition evaluation, request attributes, fine-grained authorization, limitations vs webhooks

  7. Least Privilege RBAC: Scoping & Risk Reduction - Role reduction heuristics, namespace isolation, verb minimization, blast radius limiting, privilege scoping dimensions

  8. RBAC for Multi-Tenancy & Namespace Isolation - Cross-namespace access prevention, tenant-bound RBAC, escape prevention, audit per tenant, safe patterns

  9. RBAC Audit Logging & Compliance - Audit record structure, authorization decision tracking, impersonation audit, compliance patterns, retention strategies


Chap 57: GKE with Windows Server Containers

Tại sao quan trọng: Windows containers trên GKE là niche nhưng quan trọng đối với enterprise .NET workloads.

Điều kiện tiên quyết: Chap 5, 6, Windows fundamentals

Mức độ sâu: 3/5

Chapter 57 Full Index & Learning Paths

Các chủ đề con:

  1. GKE Windows Node Pool Creation & Mental Model - Mixed-OS cluster architecture, LTSC vs SAC channel, containerd requirement, VPC-native constraint, OS labels/taints

  2. Windows CNI Considerations — Mô Hình Networking Khác Biệt - HNS, VFP, win-bridge/win-overlay, kube-proxy kernelspace mode, vì sao Dataplane V2 không hỗ trợ Windows

  3. Image Pulling — Windows Image Registry Optimization - Vì sao image Windows lớn hơn Linux nhiều lần, version-matching constraint, chiến lược pre-pull và layer caching

  4. Resource Requests — CPU/Memory Sizing Trên Windows - Job Objects thay cgroups, memory overhead baseline, nguyên tắc sizing thực tế

  5. Pod Disruption — Graceful Termination Trên Windows - CTRL_SHUTDOWN_EVENT thay SIGTERM, IHostApplicationLifetime pattern, tác động tới node drain latency

  6. Monitoring — Windows-Specific Metrics & Observability Gaps - Giới hạn cAdvisor trên Windows, windows_exporter, .NET-specific observability


Chương 58: Confidential Compute trên GKE — AMD SEV & Intel TDX

Tại sao quan trọng: Confidential Computing là pattern đang nổi lên cho các workload nhạy cảm, nơi ngay cả hypervisor và nhà cung cấp cloud cũng không thể đọc được dữ liệu đang xử lý trong bộ nhớ. Hiểu đúng cơ chế mã hóa bộ nhớ phần cứng, mô hình attestation, và overhead hiệu năng thực tế là điều kiện bắt buộc trước khi đưa vào production cho các ngành quy định chặt (tài chính, y tế, chính phủ).

Điều kiện tiên quyết: Chương 6, Chương 32

Mức độ sâu: 3/5

Chapter 58 Full Index & Learning Paths

Các chủ đề con:

  1. Trusted Execution Environment & Mô Hình Mối Đe Dọa - TEE thu hẹp trusted computing base thế nào, ba trụ cột isolation/encryption/attestation, ranh giới bảo vệ thực sự, khác biệt với Shielded VM
  2. AMD SEV / SEV-ES / SEV-SNP — Cơ Chế Mã Hóa Bộ Nhớ - AMD Secure Processor, ASID-bound key, C-bit, VMSA, Reverse Map Table chống replay/remap, vòng đời khóa gắn với node lifecycle
  3. Intel TDX — Trust Domain Extensions - SEAM mode, TDX Module, Trust Domain, MRTD/RTMR, DCAP quote verification, so sánh kiến trúc với AMD SEV
  4. Confidential GKE Nodes — Kiến Trúc & Triển Khai - Ba cấp độ cấu hình cluster/node pool/workload qua ComputeClass, ma trận machine type, giới hạn kỹ thuật, vTPM cho Pod
  5. Remote Attestation, Confidential Space & Quản Lý Khóa - Mô hình attester/verifier/relying party, passport model, cấu trúc attestation report, Confidential Space, secure key release với Cloud KMS
  6. Performance Overhead & Quyết Định Áp Dụng Production - Nguồn gốc overhead ở tầng CPU/memory controller, benchmark thực tế theo workload, khung quyết định 5 bước, anti-pattern

Chap 59: Managed Prometheus Ở Quy Mô Lớn — Tối Ưu Hóa & Xử Lý Sự Cố

Tại sao quan trọng: Managed Service for Prometheus (GMP) là giải pháp Prometheus có khả năng mở rộng, nhưng high cardinality có thể làm chi phí và latency tăng đột biến nếu không hiểu đúng cơ chế Monarch bên dưới.

Điều kiện tiên quyết: Chap 39, kiến thức Prometheus cơ bản

Mức độ sâu: 4/5

Chapter 59 Full Index & Learning Paths

Các chủ đề con:

  1. Kiến trúc Monarch & GMP ở Quy Mô Toàn Cầu — Kiến trúc ba tầng root mixer/zone mixer/leaf, cách fan-out quyết định chi phí query, khác biệt managed collection vs self-deployed collection.
  2. PodMonitoring CRD: Cấu Hình Nâng Cao & Kiểm Soát Scrape — Pipeline ScrapeEndpoint, metricRelabeling, ScrapeLimits (sampleLimit, labelLimit...), protected target label, RBAC PodMonitoring vs ClusterPodMonitoring.
  3. Rule Evaluator: Recording Rules & Alert Evaluation Ở Quy Mô — Cơ chế single-replica Rule Evaluator, khác biệt scope Rules/ClusterRules/GlobalRules, quyền multi-project, cạm bẫy aggregate away label cluster/namespace.
  4. High Cardinality & Label Explosion: Cơ Chế Vỡ Ở Tầng Nào — Giải phẫu bốn tầng cardinality explosion: metric descriptor quota, time series explosion, Monarch Quarantiner, active time series billing.
  5. Chi Phí Ingestion & Mô Hình Billing Active Time Series — Mô hình billing theo sample, công thức cardinality × tần suất scrape, ví dụ tính toán thực tế, trang Metrics Management để cost attribution.
  6. Tối Ưu Hóa PromQL: Query Nào Đắt, Vì Sao — Query timeout 120s, ratio query, long-range lookback, kỹ thuật offset và recording rule để giảm chi phí query trên Monarch.
  7. Thay Thế Vai Trò Của Thanos: Self-Deployed Collector & Retention — Vì sao GMP không hỗ trợ federation, self-deployed collector với local aggregation, ví dụ di trú kiến trúc Thanos, chính sách retention/downsample thực tế của Monarch.
  8. Troubleshooting Runbook: Query Timeout, Cardinality, Ingestion Failure — Runbook có cấu trúc từ metric up, phân nhánh query-side/ingestion-side, danh mục lỗi thực tế kèm nguyên nhân gốc và lệnh chẩn đoán.

Chap 60: Advanced Cloud Armor WAF Configuration

Tại sao quan trọng: Cloud Armor là Web Application Firewall của GCP, thực thi ngay tại Google Front End. Tuning rule đúng cách ngăn cả false positive/negative lẫn tấn công DDoS, và hiểu evaluation pipeline là điều kiện để reason chính xác về mọi hành vi của policy.

Điều kiện tiên quyết: Chap 20 (Cloud Load Balancing), Chap 27 (Network Security), OWASP Top 10 fundamentals

Mức độ sâu: 3/5

Chapter 60 Full Index & Learning Paths

Các chủ đề con:

  1. Security Policy Internal Model & Evaluation Pipeline - Enforcement tại GFE, các loại policy (backend/edge/network edge/internal), cấu trúc rule, priority và first-match-wins, action, preview mode, giới hạn body inspection

  2. Preconfigured WAF Rules: OWASP ModSecurity CRS - CRS 4.22, signature matching, sensitivity/paranoia level, evaluatePreconfiguredWaf(), opt-in/opt-out signature, request field exclusion, triết lý tuning giảm false positive

  3. Custom Rules Match Language (CEL) - Request attributes, operator và function (inIpRange, matches RE2, decode), JA3/JA4 fingerprint, evaluateThreatIntelligence/evaluateAddressGroup, giới hạn 5 subexpression

  4. Rate Limiting: Throttle & Rate-Based Ban - Bộ đếm per-key, throttle vs rate_based_ban, enforce_on_key, threshold, ban_duration, bản chất xấp xỉ và per-region của distributed counter

  5. Adaptive Protection & Always-On DDoS Defense - ML baseline per-backend-service, cấu trúc alert, manual vs evaluateAdaptiveProtectionAutoDeploy(), always-on L3/L4 DDoS, network edge security policy

  6. Managed Rules, Threat Intelligence & Named IP Lists - Cloud Armor Standard vs Enterprise, threat intelligence categories, named IP address lists bên thứ ba, address groups, hierarchical policy, mô hình cập nhật tự động

  7. Logging, Field Masking & Signed Cookies - Cấu trúc request log, verbose logging và rủi ro rò rỉ dữ liệu nhạy cảm (matchedFieldValue 16 byte), bảo vệ dữ liệu trong log, signed cookies/URLs cho custom domain


PHẦN XII: SPECIAL PRODUCTION RUNBOOKS & TROUBLESHOOTING


Chap 61: GKE Troubleshooting Runbook — Các Vấn Đề Phổ Biến & Giải Pháp

Tại sao quan trọng: Common GKE issues require specific debugging steps. Pre-written runbooks mengurangi MTTR (Mean Time To Recovery) trong production incidents.

Điều kiện tiên quyết: Chap 39–43 (Observability), Chap 53–54 (Debugging Framework)

Mức độ sâu: 4/5

Chapter 61 Full Index & Learning Paths

Các chủ đề con:

  1. Pod Creation Failures: Troubleshooting Checklist - Resource quota, insufficient node capacity, CNI failures, admission webhook rejection, PVC provisioning issues

  2. Scheduling Failures: Resolving Pending Pods - Insufficient CPU/memory, cluster autoscaler blocking, node affinity conflicts, PVC binding, daemonset resource consumption

  3. Networking Issues: Connectivity Test Procedures - Network policies blocking, service endpoints missing, DNS resolution failures, dataplane V2 issues, external load balancer configuration

  4. Storage Issues: Volume Attachment Failures - Storage class not found, PV quota exceeded, fsGroup permission timeout, filesystem type mismatch, PD attachment limits

  5. Control Plane Issues: API Server & etcd Health - etcd oversize, API server CPU/memory overload, watch cache exhaustion, admission webhook timeout, large cluster scaling limits

  6. Node Issues: NotReady Diagnosis - Kubelet crash, container runtime not running, disk pressure, memory pressure, network connectivity issues, auto-repair verification

  7. Workload Identity Failures: Token Exchange Debugging - KSA-to-GSA binding, metadata server health, token exchange quota limits, GSA missing permissions, token format issues

  8. Autoscaling Issues: HPA/CA Troubleshooting - Pod resource requests not set, metrics server unavailable, cluster autoscaler scale-up blocked, slow scaling delays, ComputeClass issues


Chap 62: GKE Cluster Upgrade Runbook — Quy Trình Zero-Downtime

Tại sao quan trọng: Nâng cấp cluster có thể gây gián đoạn nghiêm trọng nếu thực hiện thiếu cẩn trọng. Một runbook đã được kiểm chứng là yếu tố bắt buộc để vận hành production an toàn, đặc biệt khi GKE buộc nâng cấp cluster khi minor version hết hỗ trợ.

Điều kiện tiên quyết: Chap 15 (GKE Upgrade Mechanics), Chap 54 (Incident Response & Postmortems)

Mức độ sâu: 4/5

Chapter 62 Full Index & Learning Paths

Các chủ đề con:

  1. Pre-upgrade Validation: Kiểm Tra Tương Thích Trước Khi Nâng Cấp - Version skew policy, cơ chế phát hiện deprecated API dựa trên lời gọi thực tế, admission webhook/CRD conversion compatibility, giới hạn cửa sổ quan sát 30 ngày

  2. Node Surge Strategy: Sizing Cho Rollout Node Pool Ổn Định - Thuật toán surge upgrade 5 bước, ý nghĩa maxSurge/maxUnavailable per-zone, công thức sizing capacity, ảnh hưởng quota Compute Engine

  3. PDB Configuration: Đảm Bảo Disruption Budget Trong Lúc Upgrade - Eviction API và disruption controller, minAvailable vs maxUnavailable, giới hạn cứng 60 phút, 4 pattern Recommender phát hiện

  4. Control Plane Upgrade Window: Giám Sát Và Rollback Trigger - Zonal vs regional control plane, ràng buộc không-nhảy-cóc minor version, two-step upgrade với soak duration, scope maintenance exclusion

  5. Node Pool Upgrade Execution: Giám Sát Và Health Check Trong Lúc Rollout - Operation lifecycle, NodeReady health check gate, retry với backoff tăng dần, 5 phase của blue-green upgrade, autoscaled blue-green

  6. Post-upgrade Validation: Xác Nhận Hệ Thống Thực Sự Khỏe Sau Nâng Cấp - Bốn lớp xác thực (version, node label, webhook, Workload Identity), vì sao Pod Running không đủ làm tiêu chí thành công

  7. Rollback Procedures: Quy Trình Khẩn Cấp Khi Upgrade Đi Sai Hướng - Giới hạn cấu trúc của control plane rollback, node pool rollback là rolling operation ngược, cửa sổ rollback thật của blue-green


Chap 63: GCP Network Troubleshooting Methodology

Tại sao quan trọng: Network issues sulit untuk debug. Systematic approach và tool knowledge essential để resolve quickly. GCP networking layers (VPC, routing, firewall, DNS, NAT) complex, cần structured debugging.

Điều kiện tiên quyết: Chap 3 (VPC), 19–27 (networking services)

Mức độ sâu: 4/5

Chapter 63 Full Index & Learning Paths

Các chủ đề con:

  1. Connectivity Tests: GCP Native Diagnostic Tool - Using gcloud connectivity tests, route analysis, firewall rule evaluation, endpoint reachability verification

  2. VPC Flow Logs: Packet-Level Network Debugging - Flow log structure, packet direction analysis, sampling configuration, querying dalam BigQuery

  3. Firewall Rule Debugging: Evaluation Order & Matching - Rule priority, source/destination matching, stateful inspection, firewall rule audit logging

  4. Route Troubleshooting: Destination Matching & Recursive Lookup - Route selection algorithm, CIDR overlap resolution, next-hop types, custom routes vs system routes

  5. DNS Debugging: Resolution Path & TTL Issues - Cloud DNS vs CoreDNS, query resolution trace, TTL caching, FQDN resolution failures, DNS peering issues

  6. NAT Issues: Port Exhaustion Diagnosis & Prevention - Cloud NAT port allocation, port exhaustion symptoms, connection tracking limits, NAT gateway selection

  7. Load Balancer Debugging: Backend Health & Traffic Distribution - Health check failures, backend endpoint selection, traffic distribution algorithms, session affinity

  8. Service Mesh Networking: Envoy Sidecar Traffic Flow - Envoy proxy configuration, traffic interception, mTLS certificate issues, sidecar injection failures


B. SUMMARY & COVERAGE VALIDATION

Coverage Statistics:

  • Total chapters: 63 main chapters
  • Total sub-topics: 800+ detailed sub-topics
  • Total parts: 12 major sections
  • Estimated pages: 1,500–2,000 pages (if printed)
  • Estimated study time: 8–12 months for deep mastery

Covered Domains (100% Coverage):

✅ GKE Architecture & Internals (Chapters 5–18) ✅ GCP Networking Foundation & Services (Chapters 2–4, 19–27) ✅ Storage & Persistence (Chapters 28–30) ✅ IAM, Security, Compliance (Chapters 31–34) ✅ Messaging & Distributed Systems (Chapters 35–38) ✅ Observability & SRE (Chapters 39–43) ✅ CI/CD & Automation (Chapters 44–46) ✅ Service Mesh & Multi-Cluster (Chapters 47–48) ✅ Advanced Workloads (Chapters 49–50) ✅ Cost Optimization & DR (Chapters 51–52) ✅ Debugging & Incident Response (Chapters 53–54) ✅ Advanced Deep-Dives (Chapters 55–60) ✅ Production Runbooks (Chapters 61–63)

Advanced "Kill Content" - Staff/Principal Level Topics:

  1. "GKE Packet Path Anatomy: Từ Container veth pair đến Internet" — Complete packet trace mọi layer
  2. "Cluster Autoscaler Decision Engine: Tại sao scale-up chậm 45 giây" — Latency breakdown, node provisioning mechanics
  3. "etcd vs Spanner: GKE Control Plane State Storage" — Backend comparison, consistency implications
  4. "Cloud NAT Port Exhaustion: Silent Killer Production" — 5-tuple exhaustion, monitoring strategy
  5. "Workload Identity Token Exchange: Mỗi bước chi tiết" — JWT validation, STS exchange, security boundaries
  6. "Private Cluster Leak Assumptions: 5 Cách Traffic Exposes" — metadata server, DNS leaks, edge cases
  7. "Pub/Sub Regional Failure: Ordering & Failover Mechanics" — Region-scoped constraints, failover detection
  8. "Binary Authorization dalam Production: Gaps & Vectors" — Break-glass abuse, attestation replay, workarounds
  9. "SLO Error Budget Burn Rate: Multi-Window Alerting Math" — 2% budget trong 1h = 100x burn rate
  10. "Production GKE Upgrade Runbook: Zero-Downtime Playbook" — Proven procedures, validation gates, rollback

C. RECOMMENDED READING SEQUENCE

Phase 1: Foundation (Weeks 1–4)

  • Chap 1: Resource Hierarchy
  • Chap 2: Jupiter Fabric & Andromeda
  • Chap 3: VPC Model
  • Chap 4: Cloud DNS
  • Chap 31: IAM Deep Dive

Phase 2: GKE Essentials (Weeks 5–12)

  • Chap 5: Control Plane Internals
  • Chap 6: Node Lifecycle
  • Chap 7: Networking Internals
  • Chap 8: Scheduler
  • Chap 9: Autoscaling
  • Chap 10: Admission Control

Phase 3: Production Operations (Weeks 13–24)

  • Chap 39: Cloud Monitoring
  • Chap 40: Cloud Logging
  • Chap 42: SRE Practices
  • Chap 43: Debugging Methodology
  • Chap 54: Incident Response

Phase 4: Advanced Topics (Weeks 25–52)

  • Chap 32–34: Security & Secrets
  • Chap 35–38: Messaging Systems
  • Chap 44–46: Automation & CI/CD
  • Chap 47–50: Advanced Workloads
  • Chap 55–63: Deep-Dives & Runbooks