Depth 2: fundamentals, applications, production decisions and further improvements for a software practitioner. Fourteen topics, twenty-two editable SVG illustrations, two published customer cases and one hypothetical insurance API scenario. Stable API examples; no claim to exhaustive ecosystem or cluster-version coverage. No AWS deployment or measured performance result.
1
Fundamentals
Control loops, objects, traffic paths and EKS ownership.
2
Applications
Identity, persistent data, insurance design and published customer cases.
3
Best practices
Safe releases, capacity, isolation and troubleshooting.
4
Further improvements
Upgrade compatibility, cost and measured learning experiments.
Part 1 · Fundamentals
Kubernetes: APIs, control loops and execution
Prerequisites: Start here / basic technical literacy
01 / The control plane coordinates
Scroll the diagram sideways for readable labels.
Simplified coordination view. API access also carries status updates; networking and storage plugins are omitted from the arrows. [S01]
Desired state is a promise to keep
Scroll the diagram sideways for readable labels.
Illustrative replica-loss sequence. ReplicaSet maintains the count; readiness and application correctness are separate. [S02][S04]
Learning objective and mental model verified
Trace a declared workload to running containers. The API server exposes and validates the cluster API; etcd holds cluster state. Controllers reconcile objects, the scheduler places unassigned Pods, and node kubelets coordinate with the container runtime. CNI and CSI integrations supply networking and storage interfaces.
A controller observes resources, compares actual state with the desired specification and requests changes. Different controllers handle separate responsibilities. Reconciliation is repeated, rather than a one-time deployment script.
Suppose a quote API should have three replicas and one Pod disappears. Its ReplicaSet creates another Pod. Scheduling, image retrieval, startup and readiness still need to succeed. The application must preserve durable state outside that replaceable process.
For an illustrative service, declare the desired count through a Deployment rather than manually creating replacement Pods. Keep spare placement capacity and test recovery delay. A controller cannot repair invalid credentials, an unavailable database or insufficient subnet addresses merely by trying again. This inference follows from the separate scheduling and execution responsibilities; it is not a recovery-time guarantee.
Keep this: Kubernetes continually tries to converge on your declared state; convergence and business correctness are different promises.
Check yourself: Does deleting a Pod necessarily lose customer records?
Only if the records depended on that Pod’s ephemeral state. Replica replacement restores process capacity; durable data and recovery are separate concerns.
Part 2 · Fundamentals
Workloads, configuration and ownership
Prerequisites: 01-control-loops
02 / Choose the owner of the Pods
Scroll the diagram sideways for readable labels.
A workload choice map, not a guarantee that a StatefulSet replicates data. Database consistency remains the database or operator’s responsibility. [S03][S35][S36]
Learning objective and mental model verified
Choose between Deployment for interchangeable replicas, StatefulSet for stable Pod identities and storage association, DaemonSet for node-local facilities, and Job or CronJob for tasks that complete. A StatefulSet does not implement database replication or application consistency for you.
A ConfigMap stores non-confidential settings and can supply environment variables or mounted files. A Secret represents sensitive values. Base64 encoding does not make a secret confidential; control API access, storage protection and exposure through logs or manifests.
Use a Deployment for HTTP quote servers, a Job for an export and a DaemonSet for a node log agent on compatible EC2 compute. If running a database inside Kubernetes, choose an operator and storage recovery plan deliberately; stable identities alone do not satisfy the data contract.
Use labels and selectors consistently so the right controller and Service identify the intended Pods. Updating a Deployment template creates a new ReplicaSet revision; editing a generated Pod is a short-lived change that its owner may replace. Treat runtime configuration changes as releases and verify the application reload behavior.
Build a small inventory mapping every workload to its owner, dependencies, configuration source and restart behavior. This makes lifecycle ownership reviewable before a rollout or recovery exercise.
Keep this: Choose the workload controller for the lifecycle you need, then separate configuration from the image.
Check yourself: Does StatefulSet mean “highly available database”?
No. It provides workload identity and ordering/storage association primitives; database replication, consistency, failover and backup must still be designed.
Part 3 · Fundamentals
Networking: DNS, Services, routing and VPC IPs
Prerequisites: 01-control-loops, 02-objects
03 / Internal service discovery
Scroll the diagram sideways for readable labels.
DNS resolution precedes the illustrated request. Service forwarding depends on the cluster dataplane implementation; EndpointSlices describe the backing endpoints. [S05]
ALB control path ≠ request path
Scroll the diagram sideways for readable labels.
Logical routing example with IP targets. In instance-target mode traffic instead reaches a node’s NodePort. Auto Mode manages its load balancing integration separately. [S20][S30][S15]
Learning objective and mental model verified
A Service selects backend Pods and offers stable access while Pod endpoints change. EndpointSlices describe the backend endpoints. Cluster DNS provides service discovery. The forwarding implementation may use kube-proxy or another dataplane; a Service object is not itself a running proxy process.
On standard EKS, the AWS Load Balancer Controller can create an ALB for Ingress and an NLB for a LoadBalancer Service. The diagram uses ALB IP targets, where request traffic reaches Pod IPs. Instance targets take a different path through node ports. Auto Mode has managed load-balancing integration with its own configuration contract.
The Amazon VPC CNI assigns VPC addresses to Pods on AWS infrastructure. Plan workload and cluster subnet capacity, including temporary upgrade and scale-out needs. A private API endpoint requires an administration path from the VPC or a connected network; it does not automatically provide internet egress for workloads.
For the illustrative quote API, expose only the application route through an ALB and keep worker nodes in private subnets. Check DNS, ready endpoints, ALB target health, security groups, enforced policy and outbound dependencies independently. An empty Service selector or IP shortage can break traffic even while EC2 instances appear healthy.
Track available subnet addresses and target-registration delay alongside CPU and memory. Test the complete request path rather than treating “Ingress exists” as proof of reachability.
Keep this: Separate the configuration path from the request path, and treat IP space as a capacity constraint.
Check yourself: Does a client HTTP request pass through the Kubernetes API server?
Normally no. The API configures desired routing; application requests travel through the selected network and load-balancer dataplane.
Part 4 · Fundamentals
EKS responsibility boundaries and compute choices
Prerequisites: 01-control-loops
04 / Managed does not mean ownerless
Scroll the diagram sideways for readable labels.
Responsibility summary, not a full service contract. Self-managed and Hybrid Nodes are additional options. EKS control plane reliability and application reliability are different layers. [S14][S15][S16][S37]
Learning objective and mental model verified
EKS supplies a managed Kubernetes control plane. AWS documents control-plane operation across three Availability Zones. Workload replicas, dependency resilience and application recovery still require design; a highly available API server does not make a single application replica highly available.
Managed node groups simplify EC2 node provisioning and lifecycle operations. Self-managed nodes give more control and more maintenance. Auto Mode extends managed operation into compute autoscaling, networking, load balancing and block storage. Hybrid Nodes cover supported non-AWS compute environments. These are ownership choices, not application availability guarantees.
Fargate provides per-Pod compute without managing EC2 worker instances. EKS Fargate does not support DaemonSets, privileged containers or GPUs; it uses private subnets and IP load-balancer targets. EBS cannot be mounted to Fargate Pods. EFS support has separate provisioning constraints.
For a typical API with custom node agents, a managed node group may be easier to integrate than Fargate. Auto Mode is a candidate when supported defaults fit and reducing node operations matters. This is a conditional judgment: validate storage classes, IAM, agent support and disruption behavior before choosing.
“Managed” is not a universal promise of zero maintenance. Write a responsibility matrix for node releases, add-ons, identity, backups, alerts and workload compatibility. Review it whenever compute mode changes.
Keep this: Select how much infrastructure AWS operates, without giving away application ownership.
Check yourself: Will Auto Mode make an application with one replica resilient to every node replacement?
No. It manages infrastructure lifecycle; your workload still needs suitable replication, disruption tolerance and resilient dependencies.
Part 5 · Applications
IAM, Kubernetes RBAC and workload identity
Prerequisites: 04-eks-model, 02-objects
05 / Three identity boundaries
Scroll the diagram sideways for readable labels.
The first two lanes summarize documented platform mechanisms. The end-user lane is an illustrative application boundary, independent of cluster administration. [S17][S18][S32][S11]
Learning objective and mechanism verified
An EKS access entry associates an IAM principal with cluster access. Permissions can use EKS access policies or Kubernetes groups connected to RBAC. Kubernetes RBAC controls cluster API operations; it does not authorize an end user to read a particular business record.
EKS Pod Identity maps a namespace and ServiceAccount to an IAM role. A supported AWS SDK uses temporary credentials through the default provider chain and the Pod Identity agent path. Auto Mode supplies the relevant integration. The documented Pod Identity worker scope is Linux EC2; do not assume it works on Fargate.
IRSA uses a projected ServiceAccount OIDC token and AWS STS to obtain temporary role credentials. It remains a relevant option when its supported environment fits. Containers sharing a node are not a strong independent security boundary, and unrestricted instance metadata can expose the node role.
Give a document worker a dedicated ServiceAccount and an IAM role scoped to the required document prefix and actions. Give a deployer only the Kubernetes operations needed for its namespace. Separately enforce tenant and policy access in the application. This is a proposed design; exact IAM conditions depend on the data model.
Avoid attaching all workload permissions to the node role or handing every deployer cluster-admin. Inspect actual credential resolution, test both permitted and denied actions, and restrict instance metadata as appropriate. Verify SDK support and identity-agent access before debugging an S3 denial as a networking problem.
Keep this: Keep engineer-to-cluster, Pod-to-AWS and user-to-product authorization separate.
Check yourself: Can a Pod Identity association grant kubectl permission to a human?
No. Pod Identity controls workload access to AWS APIs; human cluster access uses a separate authentication and authorization path.
Part 6 · Applications
Storage, state and recovery
Prerequisites: 02-objects, 04-eks-model
06 / Ask for storage, then mount it
Scroll the diagram sideways for readable labels.
Dynamic provisioning is shown conceptually. Binding may be delayed until scheduling. Fargate has different storage support; persistent storage is not a backup. [S06][S21][S29]
A disk can pin a Pod to a zone
Scroll the diagram sideways for readable labels.
Simplified single-volume failure example. EBS attachment topology is not cross-zone database replication. EFS is a shared filesystem option with different semantics. [S39][S29][S13]
Learning objective and mechanism verified
A PVC asks for storage; a PV represents allocated storage; a StorageClass describes provisioning policy. A CSI integration connects Kubernetes to the storage system. Pods mount claims. Reclaim policy controls what happens when a claim is released; Pod replacement is a different event from deleting a claim.
Use EBS integration for supported block-volume workloads and EFS integration when shared filesystem semantics fit. EBS is zonal; align scheduling with volume topology. Auto Mode uses the ebs.csi.eks.amazonaws.com provisioner, distinct from the standard EBS CSI provisioner. Fargate cannot mount EBS and supports EFS static rather than dynamic provisioning.
In the insurance scenario, keep transaction state in a managed PostgreSQL service and documents in object storage rather than inside disposable API containers. This proposed boundary reduces the database-operator scope of the EKS platform team, but database networking, credentials, migrations, recovery and cost remain explicit responsibilities.
A persistent volume is not proof that a database can survive losing a zone. Likewise, copying Kubernetes manifests does not restore transaction data. Define which system owns database backups, document versions and workload configuration, then restore them together in a compatible sequence.
Specify an acceptable recovery point (RPO: tolerated data loss) and recovery time (RTO: tolerated outage). Rehearse an isolated restore and validate business records, not just that a Pod becomes Running. These are proposed acceptance criteria, not measured outcomes.
Reasoned design example, not a description of a deployed customer platform. Arrows denote selected dependencies; the queue box represents asynchronous dispatch from the API to workers. [S05][S17][S19][S25]
Objective and assumptions synthesis
This is an illustrative design exercise, not a claim about Tony’s deployed infrastructure. Assume HTTP quote and underwriting APIs, tenant-scoped records, uploaded documents and asynchronous tasks. Demand varies by business hours; the team wants bounded operations and observable recovery.
Clients reach an ALB. Quote and underwriting services run as replicated Deployments in private workload subnets. The API checks product authorization and tenant scope before reading PostgreSQL or issuing document access. Longer tasks use a durable queue and workers with scoped AWS identity.
Separate latency-sensitive API capacity from retryable worker capacity. Keep an appropriate reliable baseline for serving traffic; consider interruption-tolerant compute only for tasks that can safely replay. This inference rests on different failure tolerance, not on a universal claim that workers should use Spot.
Suppose demand increases and HPA adds replicas while the database connection budget stays fixed. More Pods can amplify pool pressure and timeouts. For this exercise, bound pool sizes and task concurrency, apply backpressure and choose scaling signals that reflect useful throughput. Kubernetes cannot infer the database’s safe concurrency from replica count.
Measure completed quotes, error rate, tail latency, oldest task age and database saturation before and during a load test. Require idempotency for retried document or underwriting tasks. A useful next improvement is a failure drill with one unavailable dependency, evaluating whether the API fails quickly and recovers cleanly.
Keep this: Keep the API stateless, bound downstream concurrency and make asynchronous work safe to retry.
Check yourself: Why might adding API replicas increase timeouts?
Each replica can add downstream connections and concurrent work. If the database or another dependency is saturated, extra callers increase contention rather than useful capacity.
Part 8 · Applications
Published cases: healthcare SaaS and ML search ranking
Published architecture elements, simplified for teaching; not the complete topology. [S25]
Booking.com: de-risk the migration
Scroll the diagram sideways for readable labels.
Published migration phases and result; not a transferable latency guarantee. [S24]
Crossuite: reported architecture and result verified
AWS’s Crossuite case describes a healthcare application moving from manually managed Kubernetes to EKS. Its architecture uses ALB, HPA for services, Karpenter for EC2 capacity, RDS PostgreSQL and CloudWatch. The report places the production switch in May 2024 and reports 30% lower overall costs and 99.9% uptime. A headline benefits panel displays a different uptime figure; this atlas uses the narrative’s 99.9% and treats it as a customer report, not an independently verified SLA.
Inference: replica scaling and node provisioning solve different bottlenecks, and external database dependencies must survive the migration. Reuse the two-layer capacity pattern only after measuring startup delay and dependency limits. The reported savings lack a controlled comparison for your workload.
Booking.com: reported architecture and result verified
AWS’s Booking.com case describes a dedicated EKS environment for search ranking and model-serving flexibility. Migration proceeded through inference separation, ranking API and model movement, then hybrid serving. The company benchmarked instance types and the report states approximately 40 ms latency for 99.9% of requests. The same article places some other model workloads on SageMaker; it does not claim every ML workload runs on EKS.
Inference: test the riskiest interface before moving the whole serving system. A phased migration and workload-specific instance tests make the outcome reviewable. The published latency is not an EKS service guarantee, nor evidence that another model or dataset reaches the same result.
For your own migration, define a baseline, representative peak load, acceptable regression and a reversible traffic cutover. Record the old and new compute/network/storage conditions. Avoid attributing revenue growth or every cost change solely to Kubernetes; the case descriptions report a broader modernization.
Keep this: Transfer the decision pattern; do not copy the published performance number as your own forecast.
Check yourself: Which result can be safely copied from these cases?
The architecture and evaluation patterns are useful hypotheses. Their cost and latency figures require your own measurement under your workload and operating conditions.
Probe actions summarize Kubernetes behavior. The database-outage response is a design recommendation; readiness must reflect whether a Pod can provide useful service. [S10]
Availability needs several independent controls
Scroll the diagram sideways for readable labels.
Placement, eviction and application rollout are different controls. A PDB does not constrain the Deployment rollout controller or prevent involuntary loss. [S09][S13][S04]
Learning objective and probe mechanism verified
A startup probe protects slow initialization before normal liveness and readiness checks. Readiness failure removes normal ready participation in Service endpoints; it does not itself restart the container. Liveness failure eventually restarts a container according to its restart policy. These checks answer different questions.
Deployment rolling updates use maxSurge and maxUnavailable. PDBs constrain eligible voluntary evictions, including compliant node drains. PDBs do not prevent node or zone failures, direct Pod deletion, or constrain Deployment rollout logic. Topology spread can distribute replicas across zones or hosts; strict placement rules can also leave Pods Pending.
For an illustrative quote API, start with three replicas, one surge Pod and zero unavailable Pods during a rollout. Use a PDB requiring two available replicas and spread placement across eligible zones and hosts. These are exercise settings, not a claim that three replicas meet every SLA. Reserve surge capacity and test the load remaining replicas can handle.
Do not make every transient database failure kill a healthy process. Design local liveness and useful-service readiness deliberately. During termination, the application must stop accepting new work and finish or release existing work within its grace period; align this with load-balancer draining. This is application design guidance, not a promise of immediate endpoint convergence.
Use the bundled workload example as a review specimen. Tune startup budget, readiness conditions, request duration and drain behavior using a load test. Verify failed releases and replica loss at the user-facing endpoint. A rollout remaining available says little about a backwards-incompatible database migration.
Keep this: Readiness controls participation; liveness restarts; placement and rollout controls govern other failure paths.
Check yourself: Can a PDB guarantee that a Deployment rollout never takes down too many Pods?
No. Deployment rollout availability is controlled by its rollout strategy. PDB governs eligible eviction paths and does not prevent involuntary failures.
Basic HPA ratio example assumes usable metrics and unconstrained bounds. The node capacity path is logical, not a latency guarantee. [S08][S26][S15]
Requests are placement inputs; limits are ceilings
Scroll the diagram sideways for readable labels.
Illustrative single-container arithmetic. Sidecars, defaults and admission policies can affect actual Pod requests and HPA accounting. [S07][S08]
Learning objective and mechanism verified
The scheduler uses resource requests for placement. CPU limits can cause throttling; memory limits can lead to OOM termination. Requests are not exclusive allocation of the specified RAM or CPU. HPA CPU utilization is relative to CPU requests, not the whole node’s CPU capacity.
HPA adjusts workload replica count. A node provisioner such as Karpenter evaluates unschedulable workload requirements and creates suitable capacity. Auto Mode supplies managed compute autoscaling. Placement may still be blocked by topology, taints, IP space, volume constraints or unavailable cloud capacity.
Simplifying the HPA algorithm to usable metrics and unconstrained bounds, three replicas at 90% utilization against a 60% target imply ceil(3 × 90/60) = 5 replicas. A container using 100m CPU against a 250m request reports 40% utilization. Real decisions include missing metrics, readiness, tolerance and stabilization.
For a queue worker, CPU may not reflect an aging backlog; for a Node.js API, downstream waits can dominate latency. Compare useful throughput, queue age and resource saturation before choosing a metric. Avoid two controllers fighting over the same replica field or blindly increasing resources to hide a leak. These are conditional design judgments to validate under load.
For supported retryable workloads, diversify Spot capacity and configure interruption handling. A replacement node is not guaranteed to become ready before interruption. Measure image pull, model warmup, scheduling and ready-to-serve delays; choose baseline capacity and scale-down stabilization using those observations.
Keep this: Measure the bottleneck, then scale the layer that can relieve it.
Check yourself: HPA requests five replicas but only three run. What do you inspect next?
Read the Pending Pod events and node/IP/storage capacity. HPA requests replicas; it does not itself supply EC2 nodes or remove incompatible placement constraints.
Part 11 · Best practices
Tenancy, least privilege and network policy
Prerequisites: 05-identities, 03-networking
11 / Isolation is a stack, not a namespace label
Scroll the diagram sideways for readable labels.
Platform controls summarize Kubernetes guidance. Tenant-scoped business authorization is a reasoned requirement for the worked SaaS scenario. [S11][S12][S33]
Learning objective and isolation model verified
Multi-tenancy requires both control-plane and data-plane isolation. Namespaces scope many API resources but do not create independent kernels. Dedicated nodes reduce co-location but can still share cluster services; stronger isolation may justify sandboxed execution or separate clusters, with extra cost and operations.
Without applicable NetworkPolicies, Pod ingress and egress are allowed by default. Policies require a supporting enforcement implementation and primarily express L4 rules. Default-deny egress also blocks DNS unless allowed. IAM, security groups, NetworkPolicy and business authorization cover different scopes.
In the insurance exercise, use tenant-scoped application queries even if each team has its own namespace. Restrict deployers and service accounts to their required API operations. Add explicit dependency flows before applying default-deny. On a Pod Identity workload, include the identity-agent credential path; on private AWS access, verify required service endpoints.
A tenant allowed to run arbitrary untrusted code changes the threat model. Namespace-only separation is too weak for that assumption; evaluate sandbox or dedicated-cluster designs. A toleration only permits scheduling onto a tainted node and does not, by itself, force exclusive placement. Enforce the intended placement and admission constraints.
Write a permission matrix and a network dependency map, then test denied access as well as allowed access. Verify the real policy implementation on your chosen EKS compute mode. Treat base64 values and read access to Secret objects as sensitive, and keep plaintext secret manifests out of the delivered examples.
Keep this: A namespace is an organizational boundary; build the actual trust boundary explicitly.
Check yourself: Does “one namespace per customer” enforce customer record isolation?
No. The application must enforce record access; Kubernetes namespaces scope API resources and need additional control, network and compute protections.
Reasoned troubleshooting flow based on component responsibilities and observability primitives. A symptom can have more than one root cause. [S27][S07][S05][S19]
Signals serve different consumers
Scroll the diagram sideways for readable labels.
Logical pipelines; collection components depend on compute mode. EKS does not expose every managed control-plane metric endpoint as a self-managed cluster would. [S27][S28]
Learning objective and signals verified
Metrics reveal trends and saturation, logs describe local events and traces connect request steps. The resource Metrics API used by HPA and kubectl top is a separate capability from a full monitoring pipeline. Collecting CPU usage does not provide historical logs or distributed traces.
EKS can export API, audit, authenticator, scheduler and controller-manager logs to CloudWatch. Export is disabled by default and each type is enabled separately. Audit records help investigate who changed cluster resources; retention and ingestion costs require planning.
Imagine the quote API becomes slow after adding replicas. Start with latency and error rate, then inspect traces for database waits, connection counts and per-Pod concurrency. Compare CPU throttling, restarts and queue delay. This hypothetical incident tests whether the database bottleneck diagnosis fits; it is not a claim that every latency spike has that cause.
Decision, pitfall and further improvement synthesis
Choose retention and labels around an actual troubleshooting question. Avoid putting raw tokens or unbounded tenant identifiers into telemetry. Alert on sustained user-facing errors and latency, then include placement, IP capacity and dependencies as diagnostic context. Proposed objective: find a failed release’s cause from retained evidence without needing to reproduce it.
A high-level workflow, not an executable runbook. Some add-ons need preparation before the control-plane step. Rollback has eligibility and compatibility limits and does not revert data or all add-ons. [S22][S31][S38]
Budget the whole platform
Scroll the diagram sideways for readable labels.
Published standard cluster-support rates retrieved on 2026-10-10. Region, partition, features, support tiers and resource usage can change the bill. No compute-price estimate is asserted. [S23]
Learning objective and upgrade mechanism verified
Inventory removed APIs, admission webhooks, CRDs, nodes and add-ons. Rehearse in a representative non-production environment. Follow the EKS procedure for the control plane and compatible components. Standard add-on updates are not automatically completed merely because the control plane was upgraded; Auto Mode changes some operational responsibilities.
The retrieved EKS documentation describes a conditional version rollback: initiate within seven days of an in-place upgrade, return only to the previous minor version and satisfy eligibility and compatibility checks. It preserves data and does not roll back all add-ons or non-Auto-Mode compute for you. Check current supported versions and the detailed prerequisites before relying on it. After that window, a new cluster and migration may be needed.
Published EKS cluster-support charges are $0.10 per cluster-hour for standard support and $0.60 for extended support. For an illustrative 730-hour month: $73 or $438 for this charge alone. EC2 or Fargate compute, Auto Mode charges where applicable, networking, storage, telemetry and optional features are additional. These are not an estimate for a complete production environment.
Use a tested in-place upgrade when its compatibility and recovery constraints fit. Evaluate blue/green clusters when traffic migration and duplicated infrastructure are worth the additional cost. Neither approach automatically reverses a destructive database migration. Keep application and data rollback plans separate from control-plane recovery.
Maintain an upgrade calendar and a cost allocation map. Track useful cost per completed business transaction alongside reliability metrics. Evaluate NAT, data transfer and telemetry retention rather than assuming worker instance price is the entire bill. Revisit region and account-specific pricing before a purchase decision.
Keep this: Budget compatibility work and recovery before the support calendar makes the decision for you.
Check yourself: Does reverting a cluster version revert a changed database schema?
No. Version rollback changes eligible control-plane components, not application data migrations. Database compatibility and recovery need an independent plan.
Part 14 · Further improvements
A practical learning path and improvement experiments
Proposed learning exercises for a disposable environment. Measurement targets are defined by you; no production experiment or measured result is claimed. [S09][S08][S22]
Learning objective and mental model synthesis
Treat the platform as a set of contracts: desired state, traffic delivery, permissions, persistence, capacity and recovery. Improve one weak contract at a time. This is a synthesis of the preceding mechanisms; the right next experiment depends on your workload and operating responsibility.
In a disposable test environment, deploy a small HTTP service, a Service and the supplied manifest specimen after replacing the image and implementing its endpoints. Inspect ownership and labels, replace one Pod and observe how the controller restores the count. Record time until useful traffic resumes. Do not run the specimen unchanged against production.
Increase load gradually, compare resource requests with measured demand, then inspect the HPA metric and node provisioning path. Introduce one dependency failure and one failed rollout in the test environment. A good experiment specifies demand, duration, acceptable latency/error behavior and what evidence would reject the hypothesis.
Choose managed nodes or Auto Mode according to the stated responsibility matrix. Verify private administration access, identity resolution, routing, storage compatibility and subnet headroom. Test an allowed and denied AWS action. Costs accrue for an EKS environment; plan resource cleanup as part of the exercise.
Restore data and workload configuration into an isolated compatible environment; validate records and measure RPO/RTO against your chosen objectives. Then study workload-specific topics: operators for data services, progressive delivery, policy admission, multi-cluster recovery or queue-based scaling. Add those only when the baseline measurements expose a real need.
Examples were structurally reviewed and parsed where supplied as YAML; they were not deployed to an AWS account or live cluster. No load-test, recovery-time or savings result was measured. Official sources were retrieved on 2026-10-10; living docs, region support, feature constraints and pricing need rechecking before implementation.
Keep this: Make every improvement a hypothesis with a measured success condition.
Check yourself: What would make the proposed HPA improvement fail its evaluation?
If more replicas do not improve useful throughput or meet latency/error objectives, or if startup lag and dependency saturation negate the expected benefit, the experiment rejects or narrows the hypothesis.
AWS / Booking.com · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page
Supports: Dedicated ranking cluster, phased migration and reported latency
Read the linked primary source for implementation details and current constraints.
AWS / Crossuite · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page
Supports: HPA, Karpenter, ALB, RDS PostgreSQL, CloudWatch and reported outcomes
Read the linked primary source for implementation details and current constraints.