Published cases: healthcare SaaS and ML search ranking
Prerequisites: 04-eks-model, 03-networking, 06-data
08 / Crossuite: elasticity across two layers
Scroll the diagram sideways for readable labels.
Booking.com: de-risk the migration
Scroll the diagram sideways for readable labels.
Crossuite: reported architecture and result verified
AWS’s Crossuite case describes a healthcare application moving from manually managed Kubernetes to EKS. Its architecture uses ALB, HPA for services, Karpenter for EC2 capacity, RDS PostgreSQL and CloudWatch. The report places the production switch in May 2024 and reports 30% lower overall costs and 99.9% uptime. A headline benefits panel displays a different uptime figure; this atlas uses the narrative’s 99.9% and treats it as a customer report, not an independently verified SLA.
[S25]Crossuite: engineering lesson synthesis
Inference: replica scaling and node provisioning solve different bottlenecks, and external database dependencies must survive the migration. Reuse the two-layer capacity pattern only after measuring startup delay and dependency limits. The reported savings lack a controlled comparison for your workload.
[S25]Booking.com: reported architecture and result verified
AWS’s Booking.com case describes a dedicated EKS environment for search ranking and model-serving flexibility. Migration proceeded through inference separation, ranking API and model movement, then hybrid serving. The company benchmarked instance types and the report states approximately 40 ms latency for 99.9% of requests. The same article places some other model workloads on SageMaker; it does not claim every ML workload runs on EKS.
[S24]Booking.com: engineering lesson synthesis
Inference: test the riskiest interface before moving the whole serving system. A phased migration and workload-specific instance tests make the outcome reviewable. The published latency is not an EKS service guarantee, nor evidence that another model or dataset reaches the same result.
[S24]Decision, pitfall and next improvement synthesis
For your own migration, define a baseline, representative peak load, acceptable regression and a reversible traffic cutover. Record the old and new compute/network/storage conditions. Avoid attributing revenue growth or every cost change solely to Kubernetes; the case descriptions report a broader modernization.
[S25] [S24]Check yourself: Which result can be safely copied from these cases?
The architecture and evaluation patterns are useful hypotheses. Their cost and latency figures require your own measurement under your workload and operating conditions.