Part 7 · Applications

Worked case: a multi-tenant insurance API on EKS

Prerequisites: 03-networking, 05-identities, 06-data

07 / Insurance API: a worked EKS design

Clients reach an ALB and quote or underwriting API Pods in private subnets. APIs use managed PostgreSQL, S3 and a durable queue for worker tasks.

Scroll the diagram sideways for readable labels.

Reasoned design example, not a description of a deployed customer platform. Arrows denote selected dependencies; the queue box represents asynchronous dispatch from the API to workers. [S05] [S17] [S19] [S25]

Objective and assumptions synthesis

This is an illustrative design exercise, not a claim about Tony’s deployed infrastructure. Assume HTTP quote and underwriting APIs, tenant-scoped records, uploaded documents and asynchronous tasks. Demand varies by business hours; the team wants bounded operations and observable recovery.

[S25] [S19]

Request and task flow synthesis

Clients reach an ALB. Quote and underwriting services run as replicated Deployments in private workload subnets. The API checks product authorization and tenant scope before reading PostgreSQL or issuing document access. Longer tasks use a durable queue and workers with scoped AWS identity.

[S05] [S17] [S25]

Engineering decision and trade-off synthesis

Separate latency-sensitive API capacity from retryable worker capacity. Keep an appropriate reliable baseline for serving traffic; consider interruption-tolerant compute only for tasks that can safely replay. This inference rests on different failure tolerance, not on a universal claim that workers should use Spot.

[S26] [S09]

Failure case synthesis

Suppose demand increases and HPA adds replicas while the database connection budget stays fixed. More Pods can amplify pool pressure and timeouts. For this exercise, bound pool sizes and task concurrency, apply backpressure and choose scaling signals that reflect useful throughput. Kubernetes cannot infer the database’s safe concurrency from replica count.

[S08] [S07]

Validate and improve synthesis

Measure completed quotes, error rate, tail latency, oldest task age and database saturation before and during a load test. Require idempotency for retried document or underwriting tasks. A useful next improvement is a failure drill with one unavailable dependency, evaluating whether the API fails quickly and recovers cleanly.

[S27] [S09]
Keep this: Keep the API stateless, bound downstream concurrency and make asynchronous work safe to retry.
Check yourself: Why might adding API replicas increase timeouts?

Each replica can add downstream connections and concurrent work. If the database or another dependency is saturated, extra callers increase contention rather than useful capacity.

Sources & further reading

  1. [S05] Service

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service selection, stable access and EndpointSlices

    Read the linked primary source for implementation details and current constraints.

  2. [S07] Resource management

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Requests, limits, scheduling, CPU throttling and memory enforcement

    Read the linked primary source for implementation details and current constraints.

  3. [S08] Horizontal Pod Autoscaling

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metric ratio algorithm, CPU utilization relative to requests, metric pipeline

    Read the linked primary source for implementation details and current constraints.

  4. [S09] Disruptions

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: PDB limits eviction; involuntary failure and rollout limitations

    Read the linked primary source for implementation details and current constraints.

  5. [S17] EKS Pod Identity

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service account role associations, agent, SDK support and Linux EC2 scope

    Read the linked primary source for implementation details and current constraints.

  6. [S19] VPC and subnet considerations

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Cluster subnets, private endpoint access, upgrade IP headroom

    Read the linked primary source for implementation details and current constraints.

  7. [S25] Crossuite migration to EKS

    AWS / Crossuite · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: HPA, Karpenter, ALB, RDS PostgreSQL, CloudWatch and reported outcomes

    Read the linked primary source for implementation details and current constraints.

  8. [S26] Karpenter best practices

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Pending workload capacity, diversification and interruption handling

    Read the linked primary source for implementation details and current constraints.

  9. [S27] Kubernetes observability

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metrics, logs, traces and resource Metrics API distinction

    Read the linked primary source for implementation details and current constraints.