NodeLocal DNSCache — Kiến Trúc, Conntrack Fix, và Cơ Chế Fallback
Tại Sao NodeLocal DNSCache Tồn Tại
Nhiều người nghĩ NodeLocal DNSCache là một performance optimization — thêm cache gần client để giảm latency. Điều này đúng nhưng chưa đủ. Bài toán cốt lõi mà NodeLocal DNSCache giải quyết không phải là latency mà là conntrack table exhaustion — một failure mode ở kernel level dẫn đến DNS packets bị drop silently.
Để hiểu tại sao, cần hiểu cơ chế hoạt động của DNS resolution thông qua kube-proxy iptables.
Internal Model — Vấn Đề Với Standard CoreDNS
Standard DNS Flow Và conntrack
Trong cluster không có NodeLocal DNSCache, flow DNS của một Pod như sau:
Pod → UDP query → 10.96.0.10:53 (ClusterIP of kube-dns)
↓
iptables DNAT rule
↓
CoreDNS Pod A hoặc B (DNAT đến endpoint thực)
↓
Response về PodBước quan trọng: khi gói tin UDP đi qua iptables DNAT, Linux kernel tạo một conntrack entry để track state của "connection" UDP này. Conntrack (connection tracking) là cơ chế kernel dùng để track state của network connections, cho phép return packets được forward đúng về nguồn.
Vấn đề với UDP conntrack: UDP là connectionless protocol. Kernel không có cách biết khi nào một UDP "session" kết thúc — vì vậy nó giữ conntrack entry cho đến khi timeout. Timeout mặc định cho UDP: 30 giây. Với DNS, đây nghĩa là mỗi DNS query = 1 conntrack entry tồn tại 30 giây.
Conntrack Table Exhaustion
Conntrack table có giới hạn kích thước. Mặc định trên nhiều Linux distributions: 65,536 entries. Có thể tune cao hơn nhưng có ceiling do memory.
Với cluster lớn:
- 1000 Pods × 10 DNS queries/giây/Pod = 10,000 entries/giây vào conntrack table
- Mỗi entry tồn tại 30 giây
- Peak entries = 10,000 × 30 = 300,000 entries — vượt xa giới hạn mặc định
Khi conntrack table đầy:
- Kernel drop packets mới không thể tạo entry cho
- DNS query bị drop → client timeout sau 5 giây và retry
- Retries tạo thêm pressure → positive feedback loop
- Symptom: DNS latency spike đột ngột, "DNS lookup timeout" ngẫu nhiên
Đây không phải vấn đề của CoreDNS — CoreDNS có thể đang handle bình thường. Vấn đề ở kernel networking layer. NodeLocal DNSCache loại bỏ hoàn toàn conntrack overhead cho DNS bằng cách loại bỏ nhu cầu DNAT.
NodeLocal DNSCache Architecture
DaemonSet Trên Mỗi Node
NodeLocal DNSCache là một DaemonSet — mỗi node có một pod chạy CoreDNS trong chế độ caching agent. Pod này lắng nghe trên link-local IP 169.254.20.10 (khi tích hợp với Cloud DNS) hoặc một link-local IP khác tùy cấu hình.
Node A:
├── NodeLocal DNS Pod → lắng nghe trên 169.254.20.10:53
├── Pod-1 (app) → DNS queries đến 169.254.20.10:53
├── Pod-2 (app) → DNS queries đến 169.254.20.10:53
└── Pod-3 (app) → DNS queries đến 169.254.20.10:53
Node B:
├── NodeLocal DNS Pod → lắng nghe trên 169.254.20.10:53
├── Pod-4 (app) → DNS queries đến 169.254.20.10:53
└── Pod-5 (app) → DNS queries đến 169.254.20.10:53Link-local IP 169.254.20.10 là địa chỉ trong range 169.254.0.0/16 — range này không routable ngoài subnet, không thể bị confuse với bất kỳ real IP nào trong cluster. Đây là lý do nó được chọn: safe, không conflict, luôn "local" với node.
Tại Sao Link-Local Loại Bỏ conntrack
Khi Pod gửi query đến 169.254.20.10:53:
- Không có DNAT:
169.254.20.10là IP thực của NodeLocal DNS pod trên cùng node — không cần iptables DNAT - Không có conntrack: Vì không có DNAT, kernel không cần track connection state
- Local network: Gói tin không rời node, đi thẳng từ Pod network namespace đến NodeLocal DNS pod network namespace
Kết quả: Zero conntrack entries cho DNS queries. Loại bỏ hoàn toàn vấn đề conntrack exhaustion.
Kubelet Configuration
Khi NodeLocal DNSCache được enable, kubelet cấu hình /etc/resolv.conf của mỗi Pod để trỏ đến 169.254.20.10 thay vì ClusterIP của kube-dns:
# Trước NodeLocal DNSCache
nameserver 10.96.0.10
# Sau NodeLocal DNSCache
nameserver 169.254.20.10Pods không cần biết về sự thay đổi này — họ vẫn gửi DNS queries theo cách cũ, chỉ là nameserver address khác.
Cache Behavior và TTL
Caching Logic
NodeLocal DNSCache cache responses theo TTL của DNS record:
- Responses có TTL ≤ 30 giây: Cache theo TTL đó
- Responses có TTL > 30 giây: Cache tối đa 30 giây (cap at 30s)
- NXDOMAIN responses: Cache 5 giây
Con số 30 giây cap là balance giữa cache effectiveness và freshness. Với Kubernetes Services có TTL 30 giây từ CoreDNS kubernetes plugin, NodeLocal DNSCache không extend TTL thêm — cached response vẫn expire sau 30 giây tối đa.
Negative Caching
NXDOMAIN responses được cache 5 giây. Đây là giá trị quan trọng cho ndots:5 lookup storm pattern: mỗi NXDOMAIN query cần 5 giây để expire khỏi cache. Sau 5 giây, nếu Pod retry query tương tự, một NXDOMAIN query mới sẽ được gửi đến upstream.
5 giây là TTL ngắn vì NXDOMAIN trong Kubernetes thường là do search domain expansion — api.company.com.default.svc.cluster.local sẽ luôn là NXDOMAIN nhưng cache quá lâu có thể delay trong edge cases khi domain thực sự được tạo.
Memory Usage
Mỗi NodeLocal DNS pod cache tối đa 10,000 entries theo mặc định. Với ~30MB memory per 10,000 entries khi đầy, tổng memory overhead của DaemonSet trên một node ~30-50MB. Có thể điều chỉnh qua CoreDNS cache size config.
Fallback Mechanism — Cache Miss Handling
Lookup Hierarchy
Khi NodeLocal DNS cache miss một query, nó forward đến đúng upstream dựa trên query type:
Cho cluster.local queries (cluster-internal DNS):
NodeLocal DNS Cache (miss) → kube-dns-upstream Service → CoreDNS podsNodeLocal DNSCache forward đến một Service đặc biệt kube-dns-upstream (khác với kube-dns) qua TCP thay vì UDP. TCP removes conntrack timeout issue vì TCP connection được close sau response — không có 30-giây orphan entries.
Cho external queries (non-cluster.local):
Khi tích hợp với Cloud DNS for GKE:
NodeLocal DNS Cache (miss) → 169.254.169.254:53 (metadata server / Cloud DNS)Khi không có Cloud DNS (standard CoreDNS):
NodeLocal DNS Cache (miss) → forward plugin → upstream from /etc/resolv.conf của nodeTại Sao TCP Cho Upstream Queries
TCP vs UDP là sự khác biệt quan trọng:
- UDP conntrack: Entries timeout sau 30 giây, tạo pressure trên conntrack table
- TCP conntrack: Entries bị xóa khi connection close (FIN/RST)
NodeLocal DNSCache nói chuyện với upstream CoreDNS qua TCP. Mỗi cache miss → một TCP connection ngắn → conntrack entry bị xóa ngay sau response. Tổng pressure trên conntrack table cho upstream queries giảm đáng kể so với Pod-to-CoreDNS UDP trực tiếp.
GKE Integration — Hai Chế Độ
NodeLocal DNSCache + CoreDNS (Không Có Cloud DNS)
Đây là setup mặc định cho Standard clusters với kube-dns:
Pod → 169.254.20.10 (NodeLocal) →
Cache hit: return immediately
Cache miss (cluster.local): → kube-dns-upstream Service → CoreDNS (TCP)
Cache miss (external): → upstream DNS (node's /etc/resolv.conf)NodeLocal DNSCache + Cloud DNS for GKE
Khi cluster dùng Cloud DNS for GKE, NodeLocal DNSCache thay đổi upstream:
Pod → 169.254.20.10 (NodeLocal) →
Cache hit: return immediately
Cache miss (cluster.local): → Cloud DNS via metadata server 169.254.169.254:53
Cache miss (external): → Cloud DNS via metadata server 169.254.169.254:53Cloud DNS handles cả internal và external resolution. NodeLocal DNSCache chỉ là caching layer phía trước Cloud DNS.
Theo tài liệu GKE: "When you use NodeLocal DNSCache with Cloud DNS, NodeLocal DNSCache uses the link-local IP 169.254.20.10 as the nameserver endpoint. When NodeLocal DNSCache can't resolve the query from the cache, it forwards the request to the local metadata server that runs on the same node."
Auto-Enable Trong GKE Autopilot
GKE Autopilot bật NodeLocal DNSCache mặc định và không cho phép disable. Trong Standard clusters, từ phiên bản GKE 1.34.1-gke.3720000+, NodeLocal DNSCache cũng được bật mặc định nhưng có thể disable.
Lý do Autopilot enforce NodeLocal DNSCache: Autopilot là shared infrastructure, density Pod/node cao hơn Standard, nguy cơ conntrack exhaustion cao hơn nhiều. NodeLocal DNSCache là prerequisite để Autopilot scale được.
Constraints và Failure Modes
OOMKill Và Stale iptables Rules
NodeLocal DNS pod bị OOMKill là failure mode nguy hiểm nhất. Khi pod bị kill:
- Process dừng → Pod IP giải phóng
- Nhưng iptables rules redirect DNS traffic đến
169.254.20.10vẫn còn - Packets DNS bị drop (không có listener ở
169.254.20.10nữa) - Tất cả Pods trên node mất DNS resolution
DaemonSet sẽ restart pod, nhưng trong thời gian chờ (thường 5-30 giây), tất cả DNS queries từ Pods trên node đó fail. Đây là node-level DNS outage.
Mitigation: Set memory limits hợp lý cho NodeLocal DNS pods. GKE manage điều này tự động, nhưng nếu tự quản lý, không nên set limits quá thấp. Mặc định GKE set limits đủ cao để tránh OOM trong production.
Link-Local Address Conflict
169.254.20.10 là link-local address được assign statically cho NodeLocal DNS. Nếu có ứng dụng nào trên node đang dùng địa chỉ này (hiếm nhưng có thể), conflict xảy ra.
Trong thực tế, link-local range 169.254.0.0/16 hiếm khi được dùng cho application traffic trên GKE nodes.
NodeLocal DNSCache Không Hỗ Trợ Tất Cả dnsPolicy
NodeLocal DNSCache chỉ affect Pods với dnsPolicy: ClusterFirst hoặc ClusterFirstWithHostNet. Pods với dnsPolicy: Default hoặc None không đi qua NodeLocal DNSCache — họ dùng DNS server được chỉ định trong pod spec hoặc node DNS config.
Scale Ceiling Và Memory
Mặc dù NodeLocal DNSCache loại bỏ conntrack bottleneck, mỗi pod vẫn có giới hạn:
- Max concurrent upstream queries: cấu hình qua
max_concurrenttrong forward plugin - Memory limit: quyết định cache size và pod stability
- CPU: với DNS query rate rất cao (>50,000 qps/node), NodeLocal DNS pod có thể become CPU bottleneck
Latency Improvement — Định Lượng
NodeLocal DNSCache giảm latency DNS theo hai cách:
Loại bỏ iptables DNAT overhead: DNAT lookup trong iptables là O(rules) operation. Với cluster lớn có nhiều Services, iptables chains dài → DNAT latency tăng. NodeLocal DNS bypass hoàn toàn iptables cho DNS.
Loại bỏ network hop: Trong standard CoreDNS setup, DNS query có thể DNAT đến CoreDNS pod ở node khác — thêm một network hop qua fabric. NodeLocal DNS luôn local với node — không có cross-node hop.
Trong thực tế, improvement thường là:
- p50: 0.5ms → 0.1ms (5x improvement)
- p99: 10ms → 2ms (5x improvement)
- p99.9 (tail latency): 100ms+ → 5ms (20x+ improvement do loại bỏ conntrack drop + retry cycle)
Tail latency improvement là significant nhất — với standard setup, conntrack drops tạo 30-second timeout spikes. Với NodeLocal DNSCache, những spikes này gần như biến mất.
GKE-Specific Configuration
# Kiểm tra NodeLocal DNSCache có được enable không
kubectl get pods -n kube-system -l k8s-app=node-local-dns
# Xem config của NodeLocal DNS
kubectl get configmap -n kube-system node-local-dns -o yaml
# Kiểm tra resolv.conf trong Pod (phải thấy 169.254.20.10)
kubectl exec -it <pod> -- cat /etc/resolv.conf
# Enable NodeLocal DNSCache trên Standard cluster (nếu chưa bật)
# Phải thực hiện khi tạo cluster hoặc qua gcloud update
gcloud container clusters update CLUSTER_NAME \
--update-addons NodeLocalDNS=ENABLED
# Metrics từ NodeLocal DNS (yêu cầu Prometheus)
kubectl port-forward -n kube-system <node-local-dns-pod> 9253:9253
curl http://localhost:9253/metrics | grep coredns_dns