Skip to content

Config Sync Architecture — GitOps Engine Bên Trong

Tại Sao Quan Trọng Trong Production

Config Sync là GitOps engine của GKE fleet — nhưng "GitOps" là một abstract khái niệm mà nhiều người hiểu khác nhau. Để vận hành Config Sync hiệu quả trong production với hàng chục clusters và hàng nghìn Kubernetes objects, bạn cần hiểu cơ chế cụ thể: reconciler pods làm gì, pipeline xử lý config như thế nào từ source đến cluster, khi nào sync fail và tại sao.

Không hiểu internals, bạn sẽ gặp các tình huống khó debug: tại sao config đã push lên Git nhưng 10 phút vẫn chưa apply, tại sao một Deployment bị rollback sau khi team vừa apply thủ công bằng kubectl, tại sao có ResourceConflict errors mà không rõ conflict từ đâu.

Internal Model — Config Sync Hoạt Động Như Thế Nào

Hai Resource Type: RootSync và RepoSync

Config Sync sử dụng hai loại Custom Resource là điểm khởi đầu của mọi sync operation:

RootSync quản lý synchronization ở mức cluster-scoped:

  • Chạy với quyền cluster-admin hoặc quyền cao tương đương
  • Có thể sync mọi loại resource — cluster-scoped (ClusterRole, Namespace, CustomResourceDefinition) và namespace-scoped (Deployment, Service, ConfigMap)
  • Mỗi cluster có tối đa một số giới hạn RootSync objects (thực tế thường dùng một RootSync primary)
  • Thường được sở hữu và quản lý bởi platform/cluster admin team

RepoSync quản lý synchronization ở mức namespace-scoped:

  • Chạy với quyền trong một namespace cụ thể
  • Chỉ có thể sync namespace-scoped resources trong namespace đó
  • Nhiều RepoSync có thể tồn tại, mỗi cái trong một namespace khác nhau
  • Thường được sở hữu bởi application team — đây là cơ chế delegation: cluster admin tạo RepoSync permission cho team, team tự quản lý config trong namespace của mình

Sự phân chia này phản ánh Kubernetes RBAC model: cluster-scoped operations cần cluster-admin, namespace-scoped operations chỉ cần namespace admin. Config Sync ánh xạ trực tiếp lên model này.

Reconciler Pod — Đơn Vị Thực Thi

Với mỗi RootSync hoặc RepoSync object, Config Sync operator tạo ra một Reconciler Pod tương ứng. Pod này là thực thể thực sự pull config và apply lên cluster.

Mỗi reconciler pod bao gồm nhiều containers làm việc theo pipeline:

[Source Container] → [Hydration Container] → [Sync Container]
       ↓                      ↓                      ↓
  Pull từ Git/OCI     Render Helm/Kustomize     Apply lên cluster
  Verify signature    Resolve references        Server-side apply
  Cache locally       Output hydrated configs   Status update

Source Container (git-sync / gceimage-sync):

  • Pull content từ source (Git, OCI image)
  • Maintain local cache (emptyDir volume hoặc PVC)
  • Retry với exponential backoff khi source không accessible
  • Verify cryptographic signature nếu configured
  • Output: raw config files trong shared volume

Hydration Container (hydration-controller):

  • Chỉ active khi source là Helm chart hoặc có Kustomize overlay
  • Render templates thành pure Kubernetes manifests
  • Output: "hydrated" configs trong shared volume
  • Nếu không cần hydration (plain YAML), bước này bị skip

Sync Container (reconciler):

  • Đây là container quan trọng nhất
  • Đọc hydrated configs từ shared volume
  • So sánh desired state (từ source) với actual state (từ cluster API server)
  • Apply sự khác biệt bằng server-side apply (không phải kubectl apply truyền thống)
  • Cập nhật RootSync/RepoSync status với kết quả
  • Maintain resource group — danh sách tất cả resources do sync này quản lý

Reconciliation Loop Chi Tiết

Reconciler không chỉ pull và apply một lần — nó chạy liên tục theo vòng lặp:

Polling Loop:
1. Source Container check new commits/tags (mỗi {syncPeriod}, default 15s)
2. Nếu có thay đổi: pull, verify, cache local
3. Hydration Container process nếu cần
4. Sync Container so sánh desired vs actual
5. Apply resources cần update (server-side apply)
6. Update RootSync/RepoSync .status
7. Quay lại bước 1

Drift Remediation xảy ra tại bước 4-5: nếu ai đó kubectl edit một resource mà Config Sync quản lý, sync container sẽ phát hiện drift và overwrite lại về desired state từ source. Đây là hành vi intentional — Config Sync là source of truth, manual changes sẽ bị revert.

Interval mặc định giữa các sync cycles là 15 giây (configurable qua spec.git.period hoặc tương đương). Có nghĩa là drift correction có thể mất tối đa 15 giây để có hiệu lực, cộng với thời gian pull từ source.

Server-Side Apply — Tại Sao Không Phải kubectl apply

Config Sync sử dụng Server-Side Apply (SSA) thay vì client-side apply truyền thống. Đây là quyết định thiết kế quan trọng với implications lớn:

Client-side apply (kubectl apply cũ):

  • Client tính toán patch dựa trên last-applied-configuration annotation
  • Không biết về concurrent changes từ sources khác
  • Có thể gây ra "strategic merge patch conflicts"
  • Manager ownership không được track rõ ràng

Server-side apply (SSA):

  • API server tính toán merge, không phải client
  • Mỗi "applier" đăng ký field manager — Config Sync dùng manager configsync.gke.io/{root-sync-name}
  • Khi field X được owned bởi Config Sync, bất kỳ attempt nào từ other managers để set field X sẽ create conflict
  • Conflict phải được resolved explicitly — không tự động overwrite

Hệ quả thực tế: nếu platform team dùng Config Sync để manage một Deployment, và application team chạy kubectl apply để update image tag, SSA sẽ raise conflict vì image field đang bị hai managers cùng claim. Config Sync status sẽ báo ResourceConflict.

Để giải quyết: hoặc force override (Config Sync dùng force: true để reclaim ownership), hoặc configuration phải sử dụng field ownership splitting hợp lý (platform manages infrastructure, app team manages application config riêng).

Resource Group — Inventory Management

Config Sync maintain một ResourceGroup object (Kubernetes CRD) cho mỗi RootSync/RepoSync. ResourceGroup là danh sách tất cả resources hiện đang được managed bởi sync đó.

Điều này giải quyết một vấn đề cơ bản: nếu bạn xóa một resource khỏi Git, Config Sync cần biết resource đó đã từng được managed để delete nó khỏi cluster. Không có ResourceGroup inventory, Config Sync sẽ không biết "resource này trong cluster có phải do mình tạo ra không, hay do người khác tạo".

Khi config được sync:

  1. Config Sync compare ResourceGroup (inventory) với current source content
  2. Resources trong inventory nhưng không còn trong source → DELETE trên cluster
  3. Resources trong source nhưng chưa trong inventory → CREATE trên cluster
  4. Resources trong cả hai nhưng khác nhau → UPDATE trên cluster

ResourceGroup được update sau mỗi sync cycle thành công. Nếu apply fail, ResourceGroup không được cập nhật, đảm bảo idempotency cho lần retry tiếp theo.

Pruning — Xử Lý Resources Bị Xóa Khỏi Source

Khi resource bị xóa khỏi Git/OCI source, behavior phụ thuộc vào loại resource:

Namespace-scoped resources (Deployment, Service, ConfigMap...): Config Sync sẽ DELETE resource này từ cluster.

Cluster-scoped resources (Namespace, ClusterRole...): Tương tự, sẽ bị DELETE nếu bị xóa khỏi source.

Protected resources: Một số resources được đánh dấu configsync.gke.io/prevent-deletion: "true" annotation để tránh bị xóa khi bị remove khỏi source (useful cho Namespaces quan trọng).

Pruning là một trong những lý do tại sao phải hiểu rõ config ownership trong team. Nếu namespace-X được managed bởi Config Sync và ai đó vô tình xóa nó khỏi Git repo, namespace và tất cả resources trong đó sẽ bị DELETE khỏi cluster.

Conflict Detection Và Error Handling

ResourceConflict

Xảy ra khi: hai managers (Config Sync và kubectl/other tools) cùng own một field của cùng resource.

Status trong RootSync/RepoSync:

yaml
status:
  conditions:
  - type: Syncing
    status: "False"
    reason: ResourceConflict
    message: "KNV1020: resource conflict detected for Deployment/payments/payment-api"

Giải pháp:

  • Dùng --force để Config Sync reclaim ownership: spec.override.resources[] hoặc annotation trên RootSync
  • Hoặc remove field từ Config Sync source, để field đó do other manager quản lý
  • Hoặc xóa field manager annotation bằng server-side apply với force flag

ParseError

Xảy ra khi YAML/JSON trong source không valid hoặc không thể được parsed thành Kubernetes objects.

Config Sync không apply bất cứ thứ gì nếu một file bị parse error, ngay cả những files khác valid. Đây là "all-or-nothing" model cho safety: partial apply có thể tạo ra inconsistent state.

ApplyError

Xảy ra khi API server reject một apply request (validation error, admission webhook reject, quota exceeded...).

Khác với ParseError, ApplyError cho phép một số resources được apply thành công trong khi một số khác fail. Config Sync sẽ retry failed resources trong sync cycle tiếp theo.

NetworkError

Xảy ra khi reconciler không thể reach source (Git server, OCI registry) hoặc cluster API server. Reconciler sẽ dùng last successfully fetched content từ local cache và tiếp tục apply, không dừng lại. Điều này đảm bảo cluster tiếp tục được maintained với state cuối cùng known-good từ source, ngay cả khi source temporarily unavailable.

Hierarchical Namespace Controller Integration

Config Sync có thể tích hợp với Hierarchical Namespace Controller (HNC) — một extension cho phép namespaces có parent-child relationships và tự động propagate resources từ parent xuống child namespaces.

Khi kết hợp, pattern điển hình là:

  • RootSync manage HNC parent namespace với shared config (default NetworkPolicy, ResourceQuota defaults, common RBAC)
  • HNC tự động propagate config này xuống child namespaces
  • Team's RepoSync manage tenant-specific config trong namespace của họ

Đây là cơ chế tự nhiên để implement multi-tenant config hierarchy mà không cần duplicate config cho mỗi tenant namespace.

Constraints Và Giới Hạn Thực Tế

Sync Period minimum: Mặc dù configurable, sync period không nên set quá thấp (< 5s) vì sẽ tăng load trên Git server và API server. 15s là default hợp lý cho hầu hết use cases.

Large repo performance: Với repo có hàng nghìn files, initial sync (first time fetch full repo) có thể mất vài phút. Subsequent syncs chỉ fetch git diff, nhanh hơn nhiều. OCI source thường nhanh hơn Git cho large configs vì không cần full clone.

Resource count per sync: Không có hardcoded limit nhưng hiệu suất degraded với hàng chục nghìn resources trong một sync. Pattern tốt hơn là chia nhỏ thành nhiều RootSync/RepoSync.

RBAC cho RepoSync: Để tạo RepoSync trong một namespace, user cần permissions để create RepoSync CRD trong namespace đó. Thường cluster admin tạo RepoSync và delegate cho team, hoặc dùng Config Sync itself để tạo RepoSync objects (meta-GitOps).

Cross-namespace reference limitation: RepoSync không thể reference resources outside namespace của nó. Nếu Deployment cần reference ConfigMap từ namespace khác, đó là design problem — ConfigMap đó nên được sync bởi RootSync hoặc RepoSync riêng trong namespace đó.

Deletion protection for Namespaces: Xóa namespace khỏi source sẽ trigger namespace deletion, kéo theo tất cả resources trong đó. Đây là breaking change. Luôn dùng configsync.gke.io/prevent-deletion: "true" annotation cho namespaces quan trọng.

Anti-Pattern: Manual Changes Trên Cluster

Hiểu lầm phổ biến nhất là nghĩ rằng có thể "temporarily" fix một issue bằng kubectl edit hoặc kubectl apply trực tiếp lên cluster, và sau đó "remember to update Git later".

Đây là anti-pattern nguy hiểm vì:

  1. Sync cycle tiếp theo (15s sau) sẽ revert manual change
  2. Nếu Git không được cập nhật, "fix" sẽ mất sau mỗi sync
  3. Nếu nhiều người làm điều này, Git và cluster diverge → không ai biết "source of truth" thực sự là gì
  4. ResourceConflict errors bắt đầu xuất hiện → team tắt Config Sync để "fix" → loop vicious

Mental model đúng: cluster là read-only từ góc độ human operators. Mọi thay đổi đi qua Git. Nếu cần emergency fix, fix trên Git và đợi sync (15s) hoặc trigger manual sync.

References