Skip to content

Timeout & Retry Management — Ngăn chặn cascading failures

Tại sao timeout & retry quan trọng

Service A calls Service B, which calls Service C.

Service C becomes slow (p99 latency 10 seconds).

Without timeout management:

  • Service B waits 10 seconds per request
  • Thread pool exhausted (all threads waiting)
  • Service B becomes slow
  • Service A waits, threads exhaust
  • Cascading failure: A, B, C all down

With proper timeout/retry:

  • Service B times out Service C after 2 seconds
  • Service B retries with backoff
  • Service C recovers
  • Failure contained, doesn't cascade

Timeout & retry are network resilience fundamentals.

Internal Model: Timeout hierarchies

Why hierarchy?

If timeouts are not coordinated:

Bad: Each layer uses same timeout
  Client: 10s timeout
  Service A: 10s timeout
  Service B: 10s timeout
  
What happens:
  - Call takes 3s in Service B
  - Service A waits 3s
  - Service A adds 2s overhead
  - Total 5s → Client waits 5s
  - Seems fine
  
  But edge case:
  - Call takes 9s in Service B
  - Service A starts request, waits 9s
  - Service A finishes at 9s (within timeout!)
  - But client timeout is 10s, so still ok
  
  Problem: Service A might send response at 9.5s
  Client receives at 9.5s
  But client timeout is 10s
  This is too late for client! Client might have given up.

Solution: Timeout hierarchy

User-facing request timeout: 200ms
  ├─ API Gateway: 180ms (20ms buffer)
  │   ├─ Service A: 150ms (30ms buffer)
  │   │   ├─ Service B: 100ms (50ms buffer)
  │   │   └─ Database: 80ms (20ms buffer)
  │   └─ Service C: 150ms (30ms buffer)
  │       └─ Cache: 50ms (100ms buffer)

Each layer has progressively shorter timeout going deeper. The buffers accumulate upward.

Calculating hierarchies

Formula:

Timeout = Parent timeout - (Network latency + Processing overhead)

Example:
  Client timeout: 200ms (user facing)
  Network latency Client→API: 5ms
  API processing: 10ms
  Overhead: 5ms
  
  API→downstream timeout: 200 - (5 + 10 + 5) = 180ms

Anti-pattern: Same timeout everywhere

If all services use 30 second timeout:

  • Service X slow → everything waits 30 seconds
  • Recovery takes 30+ seconds
  • Failed requests stay in queue for 30 seconds

If timeouts are hierarchical:

  • Service X slow → quick failover upstream
  • Recovery faster
  • Failed requests cleaned up faster

Timeout values — what should they be?

Depends on:

  1. SLO latency target (e.g., 95% requests < 200ms)
  2. Expected latency (e.g., p99 latency 50ms normally)
  3. Acceptable degradation (at what point is service "unacceptably slow?")

Common pattern:

Expected p99 latency: 50ms
SLO latency target: 200ms
Timeout should be: ~150-200ms (somewhere between p99 and SLO)

Why not timeout at p99 latency?

  • p99 means 1% of requests exceed it
  • If timeout = p99, then 1% of requests timeout (bad)

Why not timeout at SLO?

  • SLO is the target availability
  • If timeout = SLO, then timeout = SLO breach (defeats purpose)

Typical rule: timeout = (p99 + SLO) / 2 or just p99 + 2*stddev.

Internal Model: Retry strategy

When to retry?

Not all failures should retry:

Should retry:
  - Network timeout (service might recover)
  - 5xx error (server error, not client fault)
  - Connection refused (service temporarily down)

Should NOT retry:
  - 4xx error (client error, won't fix by retrying)
  - 404 Not Found (won't reappear by retrying)
  - 400 Bad Request (input was bad, retrying won't help)

Naive retry problem

Request times out → Service slow
Retry immediately → More load on slow service
More requests fail → More retries
Result: Thundering herd, system collapse

Solution: Exponential backoff

Attempt 1: Fail, wait 100ms
Attempt 2: Fail, wait 200ms
Attempt 3: Fail, wait 400ms
Attempt 4: Fail, wait 800ms
Attempt 5: Give up

Total wait: 100 + 200 + 400 + 800 = 1500ms
Backoff spreads load, gives service chance to recover

Exponential backoff with jitter

Problem with pure exponential backoff:

If 1000 requests all fail at same time:
  All retry at T + 100ms
  All fail again
  All retry at T + 300ms (100 + 200)
  All fail again
  Result: Synchronized thundering herd

Solution: Add jitter (randomness)

Wait = base * (2 ^ attempt) + jitter

Attempt 1: Wait = 100 * 2^1 + random(0, 50) = 200-250ms
Attempt 2: Wait = 100 * 2^2 + random(0, 100) = 400-500ms
Attempt 3: Wait = 100 * 2^3 + random(0, 200) = 800-1000ms

Result: Requests spread out, no thundering herd

Retry budget

Unlimited retries can still cause problems:

Original request: 1000 requests
Retry once: 1000 retries = 2000 total requests
Retry twice: 1000 + 1000 + 1000 = 3000 total requests
Retry 5x: 6000 total requests (6x amplification!)

If system was overloaded, 6x amplification kills it

Solution: Retry budget

Max retries per request: 3
Max total request amplification: 2x (i.e., original + 1 retry on average)

If 1000 requests initially:
  Expected total: 1000 * 2 = 2000 (not 6000)

Implement:

Max retries: 3
Retry only on: 5xx, timeout, connection refused
NOT on: 4xx errors

Implementing timeout & retry on GCP

In application code

Go example with timeout and retry:

go
import "github.com/cenkalti/backoff"

func CallServiceWithRetry(ctx context.Context, url string) ([]byte, error) {
    // Set timeout for entire operation
    ctx, cancel := context.WithTimeout(ctx, 200*time.Millisecond)
    defer cancel()
    
    // Exponential backoff
    backoffConfig := backoff.NewExponentialBackOff()
    backoffConfig.InitialInterval = 100 * time.Millisecond
    backoffConfig.MaxInterval = 1 * time.Second
    backoffConfig.MaxElapsedTime = 0 // retry until context timeout
    
    backoff.RetryNotify(
        func() error {
            resp, err := http.DefaultClient.Do(req.WithContext(ctx))
            if err != nil {
                return err  // Will retry
            }
            if resp.StatusCode >= 500 {
                return fmt.Errorf("5xx error: %d", resp.StatusCode)  // Will retry
            }
            if resp.StatusCode >= 400 {
                return backoff.Permanent(fmt.Errorf("4xx error"))  // Won't retry
            }
            return nil
        },
        backoffConfig,
    )
}

In service mesh (Istio) on GKE

Define retry policy declaratively:

yaml
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: my-service
spec:
  hosts:
  - my-service
  http:
  - retries:
      attempts: 3
      perTryTimeout: 100ms  # Each attempt: 100ms timeout
      retryOn: "5xx,reset,connect-failure,retriable-4xx"
    route:
    - destination:
        host: my-service
        port:
          number: 80

Istio handles backoff, jitter, timeout automatically.

In Cloud Functions (serverless)

Google Cloud Functions have implicit timeout (9 minutes max).

For transient errors:

go
func CallWithRetry(ctx context.Context) error {
    var lastErr error
    for attempt := 0; attempt < 3; attempt++ {
        err := DoWork(ctx)
        if err == nil {
            return nil
        }
        
        // Check if retryable
        if !IsRetryable(err) {
            return err
        }
        
        // Exponential backoff
        backoffDuration := time.Duration(math.Pow(2, float64(attempt))*100) * time.Millisecond
        select {
        case <-time.After(backoffDuration):
            // Continue
        case <-ctx.Done():
            return fmt.Errorf("context cancelled")
        }
        
        lastErr = err
    }
    return lastErr
}

Failure modes and anti-patterns

Anti-pattern 1: Retry forever

go
for {
    err := CallService()
    if err == nil {
        break
    }
}

If service is down, this loops forever, consuming resources.

Fix: Max retries + backoff

go
for attempt := 0; attempt < 3; attempt++ {
    err := CallService()
    if err == nil { break }
    if attempt < 2 {
        time.Sleep(exponentialBackoff(attempt))
    }
}

Anti-pattern 2: Immediate retry (no backoff)

go
for attempt := 0; attempt < 3; attempt++ {
    err := CallService()
    if err == nil { break }
}

No delay between retries. If service is slow to recover, retries don't help.

Fix: Add backoff

go
for attempt := 0; attempt < 3; attempt++ {
    err := CallService()
    if err == nil { break }
    if attempt < 2 {
        time.Sleep(time.Duration(100 * math.Pow(2, float64(attempt))) * time.Millisecond)
    }
}

Anti-pattern 3: Synchronized timeouts

All requests timeout at exactly same time (due to synchronized clock):

Request batch 1: All timeout at T+30s
Batch 2: All timeout at T+60s
Batch 3: All timeout at T+90s

Service receives burst of requests every 30s
Can't handle burst
Keeps failing

Fix: Add jitter to timeout

Timeout = 30000ms + random(0, 5000)ms
Requests now timeout between 30-35s
Spreads out requests to service

Anti-pattern 4: Timeout too long

Timeout: 60 seconds
User: "Why is my request taking 1 minute?"

Long timeout = long user wait for failure.

Fix: Timeout should match SLO target (e.g., 200ms for user-facing).

Anti-pattern 5: Retry on non-idempotent operations

POST /transfer → Transfer money from A to B
Retry → Transfers twice (money transferred twice)

If retry causes duplicate operation, data corrupts.

Solution: Ensure operations are idempotent (same request = same result, no side effects).

Or mark requests with unique ID:

POST /transfer with X-Idempotency-Key: uuid-123
Even if retried, server deduplicates (checks key)
Only processes once

Testing timeout and retry

Test 1: Service slow, request times out

Start mock slow service (delays 500ms)
Call with 200ms timeout
Verify timeout error raised
Verify no connection leaked

Test 2: Service fails, retry succeeds

Start mock service that fails 2 times, then succeeds
Call with retry (max 3 attempts)
Verify succeeds after retries
Verify backoff delays respected

Test 3: Cascading failure prevented

Simulate:
  - Service A calls Service B
  - Service B calls Service C
  - Service C becomes slow

Without timeout hierarchy:
  A slow → B slow → C slow → all timeout
  
With timeout hierarchy:
  C slow, B times out quickly
  A times out even faster
  Failure contained

Summary

Timeout and retry management prevent cascading failures:

Timeout hierarchy ensures timeouts decrease going deeper
Retry only on transient errors (5xx, timeout, connection refused)
Exponential backoff prevents thundering herd
Jitter spreads out retry storms
Retry budget limits request amplification
Idempotency allows safe retries
Test regularly to verify timeouts and retries work


References