Shared ingress on Envoy Gateway
Every new customer needs a load balancera route.
The evaluation, routing design and rollout were mine.
Our product runs one server per customer, about 270 across five AWS regions. Customers' devices connect on fourteen ports, most of them custom TCP, not HTTP.
Why it was needed
Each server had its own AWS load balancer, so cost and DNS records grew with every customer, and an AWS account limit was close. Sharing one was hard: every customer uses the same port numbers, and customers' firewalls allow our IP address, so it could not change.
How it was done
Tested the main ingress options against two requirements, and wrote down why each one failed.
Routed by server name with TLSRoute, because TCPRoute cannot tell customers apart on the same port.
Proved all fourteen ports on staging with our architect's upstream NATS fix, so every protocol sends a name to route on.
Rolled out in batches over three months behind one Elastic IP, with a switch back to per-customer load balancers.
Problems on the way
The first batch of 20 customers could not reach their servers.
Envoy's default circuit breaking on the shared gateway. Fix: Turned it off with a BackendTrafficPolicy before the next batch.
Half of the second batch had not updated their DNS.
Their own domains still pointed at the old load balancer. Fix: Checked each domain first and moved the late ones one by one.
Months later, one stale route took down port 443 for a region.
Two routes claimed one hostname, so Envoy rejected the whole listener. Fix: Alerts on the gateway's rejections and on applications left out of sync.
Sharing moved the risk. It did not remove it.
Results
- N → 1load balancers per region, behind one Elastic IP
- 4options rejected, each for a written reason
- 4production clusters moved in batches