Worked case: a multi-tenant insurance API on EKS
Prerequisites: 03-networking, 05-identities, 06-data
07 / Insurance API: a worked EKS design
Scroll the diagram sideways for readable labels.
Objective and assumptions synthesis
This is an illustrative design exercise, not a claim about Tony’s deployed infrastructure. Assume HTTP quote and underwriting APIs, tenant-scoped records, uploaded documents and asynchronous tasks. Demand varies by business hours; the team wants bounded operations and observable recovery.
[S25] [S19]Request and task flow synthesis
Clients reach an ALB. Quote and underwriting services run as replicated Deployments in private workload subnets. The API checks product authorization and tenant scope before reading PostgreSQL or issuing document access. Longer tasks use a durable queue and workers with scoped AWS identity.
[S05] [S17] [S25]Engineering decision and trade-off synthesis
Separate latency-sensitive API capacity from retryable worker capacity. Keep an appropriate reliable baseline for serving traffic; consider interruption-tolerant compute only for tasks that can safely replay. This inference rests on different failure tolerance, not on a universal claim that workers should use Spot.
[S26] [S09]Failure case synthesis
Suppose demand increases and HPA adds replicas while the database connection budget stays fixed. More Pods can amplify pool pressure and timeouts. For this exercise, bound pool sizes and task concurrency, apply backpressure and choose scaling signals that reflect useful throughput. Kubernetes cannot infer the database’s safe concurrency from replica count.
[S08] [S07]Validate and improve synthesis
Measure completed quotes, error rate, tail latency, oldest task age and database saturation before and during a load test. Require idempotency for retried document or underwriting tasks. A useful next improvement is a failure drill with one unavailable dependency, evaluating whether the API fails quickly and recovers cleanly.
[S27] [S09]Check yourself: Why might adding API replicas increase timeouts?
Each replica can add downstream connections and concurrent work. If the database or another dependency is saturated, extra callers increase contention rather than useful capacity.