Cloud Router Với Hybrid Connectivity — HA VPN và Cloud Interconnect
Tại Sao Quan Trọng Trong Production
Mọi enterprise hybrid connectivity deployment trên GCP đều dùng Cloud Router — nhưng cách Cloud Router integrate với HA VPN và Cloud Interconnect có những đặc thù quan trọng mà nếu không hiểu, bạn sẽ thiết kế sai topology và không đạt được SLA cam kết.
Điều nguy hiểm nhất: hai tunnel VPN hoặc hai VLAN attachments KHÔNG đồng nghĩa với high availability nếu chúng chia sẻ cùng physical infrastructure. SLA của hybrid connectivity đến từ topology vật lý, không đến từ số lượng tunnels hay BGP sessions.
Cloud Router Với HA VPN
Cơ Chế Tích Hợp Cơ Bản
Mỗi HA VPN tunnel có một BGP session riêng với Cloud Router. Một HA VPN gateway với 2 tunnels = 2 BGP sessions trên Cloud Router. Routes được học và advertised độc lập qua mỗi tunnel.
On-Prem VPN Gateway Cloud Router
│ │
├── Tunnel 1 ──────────── BGP Session 1
│ (169.254.0.1/30) │ (routes via tunnel 1)
│ │
└── Tunnel 2 ──────────── BGP Session 2
(169.254.1.1/30) │ (routes via tunnel 2)
│
VPC Dynamic Routes
(learned from both sessions)Tại sao cần BGP per tunnel (không phải per gateway)?
Khi tunnel bị down, nếu chỉ có 1 BGP session cho cả 2 tunnels, BGP không biết tunnel nào bị down. Với BGP per tunnel, khi tunnel bị down → BGP session tương ứng bị tear down → routes qua tunnel đó bị withdraw → traffic tự động reroute sang tunnel còn lại.
SLA 99.9% vs 99.99% — Topology Quyết Định Tất Cả
Đây là điểm nhiều người hiểu nhầm nhất về HA VPN.
Không đạt được 99.99% với "2 tunnels, 1 gateway":
Nếu bạn có 2 tunnels nhưng chúng đều đi vào cùng một VPN gateway (hoặc cùng một physical interface), thì gateway đó là single point of failure. Khi gateway fail, cả hai tunnels down → không còn redundancy.
Để đạt 99.99% SLA với HA VPN:
- Phía GCP: Dùng HA VPN gateway (không phải Classic VPN) — HA VPN gateway có 2 external interfaces ở 2 independent failure domains trong region
- Phía on-prem: Hai on-prem VPN gateways riêng biệt (hoặc một gateway với 2 independent interfaces)
- Topology: Mỗi external interface GCP side kết nối đến một on-prem gateway riêng
[Correct HA VPN for 99.99% SLA]
On-Prem GW 1 (interface A) ──── GCP HA VPN IF-0 ──── Tunnel 1
On-Prem GW 1 (interface A) ──── GCP HA VPN IF-1 ──── Tunnel 2
On-Prem GW 2 (interface B) ──── GCP HA VPN IF-0 ──── Tunnel 3
On-Prem GW 2 (interface B) ──── GCP HA VPN IF-1 ──── Tunnel 4
Cloud Router: 4 BGP sessions (1 per tunnel)Với topology này, không có single point of failure — một GCP interface down, một on-prem gateway down, hay bất kỳ single tunnel down đều không gây outage.
SLA 99.9% có thể đạt được với:
- 1 HA VPN gateway (2 interfaces) kết nối đến 1 on-prem gateway (2 interfaces)
- 2 tunnels, mỗi tunnel qua một interface pair
SLA thấp hơn hoặc không đảm bảo:
- Classic VPN (single tunnel)
- 2 tunnels cùng through same GCP interface
BGP Sessions Với HA VPN — Technical Details
Khi bạn tạo HA VPN tunnel và configure BGP peer, Cloud Router tạo một BGP interface trên tunnel đó với link-local address trong 169.254.0.0/16.
# Xem BGP sessions và trạng thái
gcloud compute routers get-status ROUTER_NAME \
--region=REGION \
--format='json(result.bgpPeerStatus)'Output cho thấy per-tunnel BGP session status, bao gồm:
- Peer IP (link-local)
- Session state (ESTABLISHED, IDLE, etc.)
- Uptime
- Routes received/sent
Khi tunnel bị down:
- VPN gateway detect tunnel failure (DPD — Dead Peer Detection hoặc physical signal)
- BFD detect forwarding path failure (nếu BFD enabled — khuyến khích)
- BGP session cho tunnel đó bị torn down
- Cloud Router withdraw routes learned qua tunnel đó
- Traffic failover sang tunnel khác (nếu có route qua tunnel backup)
- Failover time: với BFD ~5-10 giây, không có BFD ~60-180 giây
Cloud Router Với Cloud Interconnect
Dedicated Interconnect — Kiến Trúc Vật Lý
Dedicated Interconnect cung cấp kết nối vật lý 10G hoặc 100G giữa on-prem network và Google network tại một colocation facility. Cấu trúc:
On-Prem Network
│
└── On-Prem Router (ở colocation)
│
│ [Physical fiber cross-connect]
│
┌─────────▼─────────────┐
│ Google Edge (MSEE) │ ← Google-managed edge router
│ (Meet-Me Router) │ trong colocation facility
└─────────┬─────────────┘
│
┌─────────▼─────────────┐
│ VLAN Attachment │ ← Logical connection tạo trong GCP
└─────────┬─────────────┘
│
┌─────────▼─────────────┐
│ Cloud Router │ ← BGP session termination
└─────────┬─────────────┘
│
VPC NetworkVLAN Attachment là một logical resource đại diện cho một VLAN trên Interconnect circuit. Mỗi VLAN attachment:
- Gắn với một Interconnect circuit (physical)
- Được assign một VLAN ID (802.1Q)
- Có bandwidth allocation (share bandwidth với other attachments trên same circuit)
- Kết nối đến một Cloud Router (BGP termination point)
BGP Sessions Qua Interconnect
Với Dedicated Interconnect, BGP session được thiết lập giữa Cloud Router và on-prem router qua VLAN attachment. Link-local addresses được dùng làm BGP peering addresses (tương tự HA VPN).
Một VLAN attachment = một BGP session trên Cloud Router.
Tại sao cần nhiều VLAN attachments:
- Redundancy: một attachment fail → traffic failover sang attachment khác
- Bandwidth scaling: nhiều attachments = nhiều bandwidth (shared circuit hoặc multiple circuits)
- Isolation: tách traffic types (production vs management)
Redundancy Cho 99.99% SLA Với Interconnect
Tương tự HA VPN, SLA đến từ topology vật lý.
99.9% SLA (single-region, 2 attachments):
On-Prem Router
├── Circuit 1 → Google Edge 1 → VLAN Attachment 1 → Cloud Router
└── Circuit 2 → Google Edge 1 → VLAN Attachment 2 → Cloud RouterNếu Google Edge 1 fail → cả hai attachments down. Đây là single point of failure, chỉ đạt 99.9%.
99.99% SLA (2 metro locations, 4 attachments):
Location A:
On-Prem Router A → Circuit 1 → Google Edge A1 → VLAN Attachment 1
On-Prem Router A → Circuit 2 → Google Edge A2 → VLAN Attachment 2
Location B (khác metro):
On-Prem Router B → Circuit 3 → Google Edge B1 → VLAN Attachment 3
On-Prem Router B → Circuit 4 → Google Edge B2 → VLAN Attachment 4
Cloud Router: 4 BGP sessions, routes qua tất cả 4 attachmentsKhông có single physical failure nào (1 router, 1 edge, 1 circuit) có thể gây outage toàn bộ.
Partner Interconnect — Khác Biệt Quan Trọng
Partner Interconnect dùng một service provider là intermediary. Hai model:
Layer 2 Partner (L2):
- Provider mang L2 Ethernet từ on-prem đến Google colocation
- On-prem router có BGP session TRỰC TIẾP với Cloud Router (qua L2 circuit của provider)
- Behavior tương tự Dedicated Interconnect về BGP
Layer 3 Partner (L3):
- Provider terminate BGP ở phía họ, và có BGP riêng với Google
- On-prem router không có BGP session với Cloud Router
- Cloud Router chỉ thấy routes từ partner's BGP infrastructure
- Hạn chế quan trọng: MED không được hỗ trợ qua L3 Partner Interconnect. Bạn không thể dùng MED để control traffic routing với L3 Partner. Thay vào đó phải dùng AS path length.
Active-Active vs Active-Passive Failover
Active-Active — Tất Cả Tunnels/Attachments Đều Mang Traffic
Trong active-active, Cloud Router học routes từ tất cả tunnels/attachments và distributes traffic qua chúng:
Cách thiết lập: Advertise cùng routes từ tất cả peers với cùng MED và AS path. Cloud Router sẽ coi tất cả là equal-cost và ECMP (Equal-Cost Multi-Path) routing.
Hành vi khi một link fail: Routes qua link đó bị withdraw. Traffic còn lại được phân phối sang các links healthy.
Ưu điểm: Tận dụng toàn bộ bandwidth, không có "wasted" standby capacity.
Nhược điểm: Khi fail, traffic phải converge — trong khoảng thời gian convergence (BFD: ~5-10s), có disruption. Asymmetric routing có thể xảy ra nếu cả hai bên không đồng ý về best path.
Active-Passive — Một Link Chính, Một Link Backup
Trong active-passive, một link chứa toàn bộ traffic (active), link kia không chứa traffic bình thường nhưng BGP session vẫn ESTABLISHED (passive/standby).
Cách thiết lập:
- Dùng MED: Active link advertise với MED thấp (preferred), passive link advertise với MED cao.
- Dùng AS path prepending: Passive link prepend AS path nhiều lần để route có AS path length dài hơn (kém preferred hơn).
Hành vi khi active link fail: Routes qua active link bị withdraw. Passive link routes (với MED cao/AS path dài) trở thành best path và được install.
Ưu điểm: Predictable routing, không asymmetric, dễ troubleshoot.
Nhược điểm: Lãng phí 50% bandwidth (passive link không chứa traffic trong normal operation). Convergence khi active link fail vẫn tốn thời gian.
Ví Dụ Cấu Hình Active-Passive Với MED
# Trên Cloud Router, cấu hình BGP peer cho active link
gcloud compute routers update-bgp-peer ROUTER_NAME \
--peer-name=PEER_ACTIVE \
--advertisement-mode=DEFAULT \
--region=REGION
# Cấu hình BGP peer cho passive link (advertise với MED cao)
# Thực hiện ở phía on-prem: on-prem router set MED=100
# cho routes advertised qua passive linkPhía GCP: on-prem router advertise routes qua active link với MED=0 (hoặc không set, default 0), qua passive link với MED=100. Cloud Router chọn MED thấp hơn → routes qua active link được install.
Khi active link fail:
- Routes qua active link bị withdraw
- Routes qua passive link (MED=100) là best remaining → được install
- Traffic chuyển sang passive link
Lưu ý với standard vs legacy best-path selection: Để MED-based traffic engineering hoạt động nhất quán qua nhiều Cloud Router BGP tasks, nên dùng standard best-path selection mode.
Constraints Và Anti-Patterns Thực Tế
Anti-Pattern: "2 Tunnels Là HA"
Biểu hiện: Team tạo 2 VPN tunnels trên cùng một Classic VPN gateway và cho rằng đã có HA.
Vì sao sai: Classic VPN gateway là single instance. Khi nó fail, cả 2 tunnels down. Không có HA.
Cách đúng: Dùng HA VPN với 2 external interfaces ở independent failure domains.
Anti-Pattern: Không Có BFD Với Interconnect
Biểu hiện: Dedicated Interconnect được triển khai không bật BFD, dựa vào BGP hold timer (180s) cho failure detection.
Vì sao sai: Với 99.99% SLA Interconnect, hold timer 180s nghĩa là potential 3 phút blackhole khi physical failure. Điều này ảnh hưởng trực tiếp đến SLA của application.
Cách đúng: Enable BFD ở cả Cloud Router side và on-prem router side cho tất cả Interconnect BGP sessions.
Anti-Pattern: Dùng L3 Partner Interconnect Với MED-Based Traffic Engineering
Biểu hiện: Team chọn L3 Partner Interconnect và cố dùng MED để control traffic giữa active và passive links.
Vì sao sai: MED không được hỗ trợ qua L3 Partner Interconnect. MED set bởi on-prem không được passed đến Cloud Router qua L3 partner's BGP.
Cách đúng: Với L3 Partner Interconnect, dùng AS path prepending để control traffic engineering. Hoặc chuyển sang L2 Partner hoặc Dedicated Interconnect nếu MED cần thiết.