Part 8 · Applications

Published cases: healthcare SaaS and ML search ranking

Prerequisites: 04-eks-model, 03-networking, 06-data

08 / Crossuite: elasticity across two layers

ALB routes to EKS Pods; HPA scales replicas, Karpenter capacity. RDS stores data; CloudWatch observes.

Scroll the diagram sideways for readable labels.

Published architecture elements, simplified for teaching; not the complete topology. [S25]

Booking.com: de-risk the migration

Inference moved first, then ranking API and models, then hybrid serving. Dedicated EKS capacity was benchmarked.

Scroll the diagram sideways for readable labels.

Published migration phases and result; not a transferable latency guarantee. [S24]

Crossuite: reported architecture and result verified

AWS’s Crossuite case describes a healthcare application moving from manually managed Kubernetes to EKS. Its architecture uses ALB, HPA for services, Karpenter for EC2 capacity, RDS PostgreSQL and CloudWatch. The report places the production switch in May 2024 and reports 30% lower overall costs and 99.9% uptime. A headline benefits panel displays a different uptime figure; this atlas uses the narrative’s 99.9% and treats it as a customer report, not an independently verified SLA.

[S25]

Crossuite: engineering lesson synthesis

Inference: replica scaling and node provisioning solve different bottlenecks, and external database dependencies must survive the migration. Reuse the two-layer capacity pattern only after measuring startup delay and dependency limits. The reported savings lack a controlled comparison for your workload.

[S25]

Booking.com: reported architecture and result verified

AWS’s Booking.com case describes a dedicated EKS environment for search ranking and model-serving flexibility. Migration proceeded through inference separation, ranking API and model movement, then hybrid serving. The company benchmarked instance types and the report states approximately 40 ms latency for 99.9% of requests. The same article places some other model workloads on SageMaker; it does not claim every ML workload runs on EKS.

[S24]

Booking.com: engineering lesson synthesis

Inference: test the riskiest interface before moving the whole serving system. A phased migration and workload-specific instance tests make the outcome reviewable. The published latency is not an EKS service guarantee, nor evidence that another model or dataset reaches the same result.

[S24]

Decision, pitfall and next improvement synthesis

For your own migration, define a baseline, representative peak load, acceptable regression and a reversible traffic cutover. Record the old and new compute/network/storage conditions. Avoid attributing revenue growth or every cost change solely to Kubernetes; the case descriptions report a broader modernization.

[S25] [S24]
Keep this: Transfer the decision pattern; do not copy the published performance number as your own forecast.
Check yourself: Which result can be safely copied from these cases?

The architecture and evaluation patterns are useful hypotheses. Their cost and latency figures require your own measurement under your workload and operating conditions.

Sources & further reading

  1. [S24] Booking.com search ranking on Amazon EKS

    AWS / Booking.com · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Dedicated ranking cluster, phased migration and reported latency

    Read the linked primary source for implementation details and current constraints.

  2. [S25] Crossuite migration to EKS

    AWS / Crossuite · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: HPA, Karpenter, ALB, RDS PostgreSQL, CloudWatch and reported outcomes

    Read the linked primary source for implementation details and current constraints.