Skip to content

Chaos Engineering — Kiểm tra độ mạnh mẽ một cách chủ động

Tại sao chaos engineering quan trọng

You believe your system is resilient. Timeouts work. Retries work. Circuit breakers work.

But do they? Only one way to know: test them.

Real scenario:

Company invests $500K in disaster recovery infrastructure.
Never tested in production.

Disaster happens → DR doesn't work → $2M in losses

Why? Infrastructure worked in staging, but:
  - DNS failover not actually cutting over traffic
  - Cross-region replication laggy
  - Backup restore took 2 hours, not 5 minutes as assumed

Chaos engineering answers: What happens when things actually break?

Instead of hoping your resilience patterns work, you actively test them by injecting failures.

Internal Model: Chaos engineering levels

Level 1: Dependency failures

Inject failures into external dependencies:

E-commerce checkout:

Test 1: Database timeout
  - Inject: All database calls timeout after 100ms
  - Expected: Checkout shows graceful error (cached data)
  - Actual: Checkout returns 500, no fallback
  - Finding: Circuit breaker missing!

Test 2: Cache unavailable
  - Inject: Cache service returns 5xx
  - Expected: Fall back to database
  - Actual: Fall back works
  - Confidence: +1

Level 2: Pod/node failures

Crash containers, kill nodes:

GKE cluster:

Test 1: Kill 1 pod
  - Inject: Delete random pod
  - Expected: Service continues, new pod spins up
  - Actual: Service continues, new pod replaces within 10s
  - Confidence: +1

Test 2: Kill 50% of pods
  - Inject: Delete half the replicas
  - Expected: Remaining pods handle load, new ones scale up
  - Actual: Remaining pods saturated, request queue exceeds 1000
  - Finding: HPA threshold too high!

Level 3: Compute resource failures

Saturate CPU, memory, disk:

Test 1: CPU saturation
  - Inject: Each pod maxes CPU to 100%
  - Expected: HPA scales, adds new pods
  - Actual: HPA adds 1 pod, but takes 2 minutes (slow startup)
  - Finding: HPA too slow, add pre-warming

Test 2: Memory leak
  - Inject: Simulate memory leak (gradual increase)
  - Expected: OOMKill kills pod, new one replaces
  - Actual: Pod OOMKilled, but recovery takes 30s
  - Finding: Document RTO is 30s, acceptable

Level 4: Network failures

Introduce latency, packet loss, partition:

Test 1: High latency
  - Inject: Add 500ms latency to all network calls
  - Expected: Timeouts kick in, fallbacks work
  - Actual: Everything times out, no fallback
  - Finding: Timeout thresholds misconfigured

Test 2: Network partition
  - Inject: Block all traffic between region A and B
  - Expected: Failover to region C
  - Actual: Both regions try to coordinate, split-brain
  - Finding: Quorum-based decision making needed

Level 5: Cascading failures

Combine multiple failures:

Test 1: Database down + cache down
  - Inject: Database 5xx + cache timeout
  - Expected: Graceful degradation
  - Actual: Requests queue indefinitely
  - Finding: Load shedding missing!

Test 2: Primary region down + slow secondary
  - Inject: Failover to secondary, secondary latency 2s
  - Expected: Service stays available with latency SLO breach
  - Actual: Clients timeout, abandon requests
  - Finding: Client timeout too aggressive for failover scenario

Designing chaos experiments

Step 1: Hypothesis

Before you run chaos, state what you expect:

Hypothesis: "When database is down, service falls back to cache and remains available"

Test:
  - Inject: Database 5xx errors
  - Expected outcome: Error rate < 1%, latency < 500ms
  - Expected behavior: Metrics show cache hits increase

Step 2: Scope — what to break?

Start small, escalate:

Week 1: Single pod crashes
Week 2: Database connection timeout
Week 3: Network latency to dependencies
Week 4: Multiple failures together

Each test validates one aspect. Combining tests comes later.

Step 3: Blast radius — what's at risk?

Experiment: Inject latency on database calls
Blast radius: Only staging environment (not prod)
Or: Production but canary traffic only (5%)
Duration: 5 minutes max (can revert fast)

Never start chaos experiment on 100% production traffic. Start with:

  • Staging environment (safe, no customers)
  • Canary in production (5–10% traffic)
  • Then expand if results good

Step 4: Observability — how to know if it worked?

Before running experiment, decide what to measure:

Experiment: Kill 50% of pods
Metrics to watch:
  - Error rate (should remain < 1%)
  - Latency p99 (should remain < 500ms)
  - Pod count (should recover to 100%)
  - Request queue depth (should not exceed 1000)

If any metric breaches, experiment fails

Step 5: Execution and rollback

T-0: Establish baseline metrics (5 min no injection)
T-0min: Inject failure
T+5min: Observe metrics
T+10min: Stop injection, verify recovery
T+15min: Metrics back to baseline

Automated rollback if metrics breach threshold:

if ErrorRate > 2% {
    StopChaosExperiment()
    Alert("Chaos broke system!")
}

Implementing chaos on GCP

Option 1: Pod disruption (GKE native)

Kubernetes PodDisruptionBudget allows controlled pod kills:

yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api-pdb
spec:
  minAvailable: 2  # Always keep 2 pods running
  selector:
    matchLabels:
      app: api
---
# Chaos experiment: Kill pods respecting PDB
$ kubectl delete pod -l app=api -n production --grace-period=0 --force
# Kubernetes kills 1 pod at a time, respecting minAvailable=2

Option 2: Istio fault injection

Inject faults into service mesh traffic:

yaml
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: api
spec:
  hosts:
  - api
  http:
  - fault:
      delay:
        percentage: 10  # 10% of requests
        fixedDelay: 500ms
      abort:
        percentage: 5   # 5% of requests
        httpStatus: 500
    route:
    - destination:
        host: api
        port:
          number: 8080

This injects latency and errors without touching application code.

Option 3: Chaos Mesh

Chaos Mesh is open-source chaos engineering platform. Can run on GKE:

bash
# Install Chaos Mesh
helm repo add chaos-mesh https://charts.chaos-mesh.org
helm install chaos-mesh chaos-mesh/chaos-mesh -n chaos-testing --create-namespace

Define experiments:

yaml
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: kill-one-pod
  namespace: default
spec:
  action: kill
  mode: one
  selector:
    namespaces:
      - default
    labelSelectors:
      app: api
  scheduler:
    cron: "0 10 * * *"  # Run every day at 10 AM

Chaos Mesh provides Web UI for experiments, automatic scheduling, metrics integration.

Option 4: Traffic shaping with Cloud Load Balancer

Use GCP Cloud Load Balancer to inject failures:

Can't directly inject failures, but can:
  - Simulate high latency (use advanced routing)
  - Route to slower backend (simulates degradation)
  - Rate limit (simulates overload)

Option 5: Manual injection in staging

bash
# Simulate database timeout in staging
$ kubectl set env deployment/api -n staging \
    DB_TIMEOUT="100ms"  # Force all DB calls to timeout

# Run load test against staging
$ load-test --target staging.example.com --rps 1000

# Observe behavior
$ kubectl logs -f deployment/api -n staging | grep -i error

# Rollback
$ kubectl set env deployment/api -n staging \
    DB_TIMEOUT=""  # Reset to normal

Chaos experiment templates

Template 1: Database failure

yaml
Experiment: "Database completely unavailable"

Setup:
  - Staging environment only
  - Baseline: Run load test for 5 min, capture metrics

Injection:
  - Block all connections to database
  - Mechanism: Firewall rule, iptables, or kill database pod
  - Duration: 5 minutes

Expected outcomes:
  - Service returns graceful error (5xx or fallback data)
  - Fallback metrics show cache hits increase
  - No cascading failures to upstream services

Measure:
  - Error rate (should be < 2%)
  - Cache hit rate (should increase by > 50%)
  - Latency p99 (should remain < 1s even with 5xx errors)

Rollback:
  - Remove firewall rule / restart database
  - Verify recovery within 1 minute

Template 2: Slow service (cascading impact)

yaml
Experiment: "Downstream service becomes slow"

Setup:
  - Canary traffic (5% of production)
  - Measurement window: 10 minutes

Injection:
  - Add 1 second latency to all downstream service calls
  - Mechanism: Istio fault injection
  - Percentage: 100% (all calls slow, not just sample)

Expected outcomes:
  - Upstream service applies timeout (cancels slow calls)
  - Circuit breaker opens after 5 failed calls
  - Fallback to cached data
  - Error rate stays low (< 1%)

Measure:
  - Timeout count (should be > 0)
  - Circuit breaker state transitions
  - Fallback usage
  - End-to-end latency increase

Rollback:
  - Remove Istio fault injection config
  - Verify latency returns to normal

Template 3: Node failure (GKE)

yaml
Experiment: "Entire node crashes"

Setup:
  - Production, limited to 1 region first
  - Precondition: PodDisruptionBudget in place (minAvailable > 0)

Injection:
  - Cordon node (mark as unschedulable)
  - Drain pods from node (graceful shutdown)
  - After drain, force stop node (simulate hardware failure)

Expected outcomes:
  - Pods reschedule to other nodes
  - HPA doesn't trigger (already have minAvailable)
  - Service stays available
  - Pod startup completes within 30 seconds

Measure:
  - Pod rescheduling time
  - Service availability during failure
  - Request errors during transition
  - Node recovery time

Rollback:
  - Uncordon node
  - Verify node rejoins cluster

Template 4: Cascading regional failure

yaml
Experiment: "Entire region becomes unavailable"

Setup:
  - Production, multi-region setup
  - Precondition: Cross-region replication working

Injection:
  - Block all traffic to primary region (iptables, firewall)
  - Duration: 10 minutes or until failover completes
  - Measurement: When does traffic actually failover?

Expected outcomes:
  - Traffic automatically routes to secondary region
  - Failover completes within 2 minutes
  - Data consistency maintained (no corruption)
  - Clients don't see errors (seamless failover)

Measure:
  - Failover time (from partition to secondary active)
  - Error rate during failover (should be 0%)
  - Data consistency check after failover
  - Time to recovery (restore primary)

Rollback:
  - Unblock traffic to primary
  - Verify replication catches up
  - Failover back to primary (if desired)

Common mistakes in chaos experiments

Mistake 1: Chaos without observability

Run chaos experiment, but no metrics collected.

Result: Don't know what happened.

Fix: Before running chaos, set up dashboards monitoring the specific system.

Mistake 2: Chaos on 100% production immediately

Run experiment on 100% prod traffic.

Result: Real customers affected, might violate SLO.

Fix: Start with staging. Then canary (5%). Then expand.

Mistake 3: Experiment too aggressive

Kill 90% of pods at once.

Result: Service completely down, takes hours to recover.

Fix: Start small (1 pod, then 10%, then 50%).

Mistake 4: No automated rollback

Experiment breaks system, manual intervention required.

Result: Incident, manager yells "don't do chaos experiments again!"

Fix: Implement automated rollback if metrics breach.

Mistake 5: Experiment never documented

Run chaos experiment, find issues, fix them, forget what you learned.

Result: Month later, run same experiment, discover same issues.

Fix: Document experiment, findings, and fixes. Use as runbook.

Measuring chaos effectiveness

Metric 1: Issues found

Q1 Chaos Experiments:
  Experiment 1: Missing circuit breaker → found and fixed
  Experiment 2: Timeout thresholds too aggressive → adjusted
  Experiment 3: Cache fallback works → confidence +1
  
Total issues found: 2 critical, 1 enhancement

If chaos experiments find 0 issues, either:

  • System is very robust (good!)
  • Experiments too shallow (bad!)

Metric 2: Real incident prevention

Without chaos: 3 production incidents per quarter (average)
After 6 months of chaos experiments: 0 incidents

Correlation: Chaos experiments prevented known failure modes

Metric 3: RTO/RPO validation

Chaos experiment: Regional failure
Measured RTO: 45 seconds
SLO RTO target: 60 seconds

Status: Compliant
Confidence: Medium (tested in prod)

Summary

Chaos engineering validates resilience:

Test dependencies fail — circuit breakers, fallbacks work
Test compute fails — pod/node failures don't cascade
Test resource exhaustion — autoscaling responds, load shedding works
Test network issues — timeouts, retries work
Test cascading failures — multiple failures don't break system
Start in staging — escalate to production gradually
Measure everything — can't improve what you don't measure
Automate rollback — experiment shouldn't break production
Document findings — turn discoveries into improvements


References