An illustrated field notebook · Kubernetes + Amazon EKS

KUBERNETES
AWS EKS
ILLUSTRATED ATLAS

Depth 2 · Best practices

From declared state to useful, reliable service.

14 topics · 22 editable diagrams · 39 primary sources · researched 10 October 2026

Depth 2: fundamentals, applications, production decisions and further improvements for a software practitioner. Fourteen topics, twenty-two editable SVG illustrations, two published customer cases and one hypothetical insurance API scenario. Stable API examples; no claim to exhaustive ecosystem or cluster-version coverage. No AWS deployment or measured performance result.

1

Fundamentals

Control loops, objects, traffic paths and EKS ownership.

2

Applications

Identity, persistent data, insurance design and published customer cases.

3

Best practices

Safe releases, capacity, isolation and troubleshooting.

4

Further improvements

Upgrade compatibility, cost and measured learning experiments.

Part 1 · Fundamentals

Kubernetes: APIs, control loops and execution

Prerequisites: Start here / basic technical literacy

01 / The control plane coordinates

The API server stores state in etcd. Controllers and scheduler use the API. Worker kubelet instructs the container runtime to run application containers.

Scroll the diagram sideways for readable labels.

Simplified coordination view. API access also carries status updates; networking and storage plugins are omitted from the arrows. [S01]

Desired state is a promise to keep

A three-replica Deployment loses one Pod, observes only two and creates a replacement that must be scheduled and started.

Scroll the diagram sideways for readable labels.

Illustrative replica-loss sequence. ReplicaSet maintains the count; readiness and application correctness are separate. [S02] [S04]

Learning objective and mental model verified

Trace a declared workload to running containers. The API server exposes and validates the cluster API; etcd holds cluster state. Controllers reconcile objects, the scheduler places unassigned Pods, and node kubelets coordinate with the container runtime. CNI and CSI integrations supply networking and storage interfaces.

[S01]

Mechanism verified

A controller observes resources, compares actual state with the desired specification and requests changes. Different controllers handle separate responsibilities. Reconciliation is repeated, rather than a one-time deployment script.

[S02]

Worked example synthesis

Suppose a quote API should have three replicas and one Pod disappears. Its ReplicaSet creates another Pod. Scheduling, image retrieval, startup and readiness still need to succeed. The application must preserve durable state outside that replaceable process.

[S04] [S02]

Decision, pitfall and improvement synthesis

For an illustrative service, declare the desired count through a Deployment rather than manually creating replacement Pods. Keep spare placement capacity and test recovery delay. A controller cannot repair invalid credentials, an unavailable database or insufficient subnet addresses merely by trying again. This inference follows from the separate scheduling and execution responsibilities; it is not a recovery-time guarantee.

[S01] [S02]
Keep this: Kubernetes continually tries to converge on your declared state; convergence and business correctness are different promises.
Check yourself: Does deleting a Pod necessarily lose customer records?

Only if the records depended on that Pod’s ephemeral state. Replica replacement restores process capacity; durable data and recovery are separate concerns.

Part 2 · Fundamentals

Workloads, configuration and ownership

Prerequisites: 01-control-loops

02 / Choose the owner of the Pods

Four workload rows map Deployment to quote API, StatefulSet to stateful identities, DaemonSet to node collectors and Jobs to exports.

Scroll the diagram sideways for readable labels.

A workload choice map, not a guarantee that a StatefulSet replicates data. Database consistency remains the database or operator’s responsibility. [S03] [S35] [S36]

Learning objective and mental model verified

Choose between Deployment for interchangeable replicas, StatefulSet for stable Pod identities and storage association, DaemonSet for node-local facilities, and Job or CronJob for tasks that complete. A StatefulSet does not implement database replication or application consistency for you.

[S03]

Configuration mechanism verified

A ConfigMap stores non-confidential settings and can supply environment variables or mounted files. A Secret represents sensitive values. Base64 encoding does not make a secret confidential; control API access, storage protection and exposure through logs or manifests.

[S35] [S36]

Worked example synthesis

Use a Deployment for HTTP quote servers, a Job for an export and a DaemonSet for a node log agent on compatible EC2 compute. If running a database inside Kubernetes, choose an operator and storage recovery plan deliberately; stable identities alone do not satisfy the data contract.

[S03]

Decision and common pitfall synthesis

Use labels and selectors consistently so the right controller and Service identify the intended Pods. Updating a Deployment template creates a new ReplicaSet revision; editing a generated Pod is a short-lived change that its owner may replace. Treat runtime configuration changes as releases and verify the application reload behavior.

[S04] [S35]

Further improvement synthesis

Build a small inventory mapping every workload to its owner, dependencies, configuration source and restart behavior. This makes lifecycle ownership reviewable before a rollout or recovery exercise.

[S03] [S04]
Keep this: Choose the workload controller for the lifecycle you need, then separate configuration from the image.
Check yourself: Does StatefulSet mean “highly available database”?

No. It provides workload identity and ordering/storage association primitives; database replication, consistency, failover and backup must still be designed.

Part 3 · Fundamentals

Networking: DNS, Services, routing and VPC IPs

Prerequisites: 01-control-loops, 02-objects

03 / Internal service discovery

Cluster DNS resolves a Service name, and the client sends to a stable Service address which forwards to one of two ready endpoints.

Scroll the diagram sideways for readable labels.

DNS resolution precedes the illustrated request. Service forwarding depends on the cluster dataplane implementation; EndpointSlices describe the backing endpoints. [S05]

ALB control path ≠ request path

A dashed control path leads from routing objects through the AWS controller to ALB configuration. A solid request path runs client to ALB to Pod IP.

Scroll the diagram sideways for readable labels.

Logical routing example with IP targets. In instance-target mode traffic instead reaches a node’s NodePort. Auto Mode manages its load balancing integration separately. [S20] [S30] [S15]

Learning objective and mental model verified

A Service selects backend Pods and offers stable access while Pod endpoints change. EndpointSlices describe the backend endpoints. Cluster DNS provides service discovery. The forwarding implementation may use kube-proxy or another dataplane; a Service object is not itself a running proxy process.

[S05]

AWS integration verified

On standard EKS, the AWS Load Balancer Controller can create an ALB for Ingress and an NLB for a LoadBalancer Service. The diagram uses ALB IP targets, where request traffic reaches Pod IPs. Instance targets take a different path through node ports. Auto Mode has managed load-balancing integration with its own configuration contract.

[S20] [S30] [S15]

VPC mechanics verified

The Amazon VPC CNI assigns VPC addresses to Pods on AWS infrastructure. Plan workload and cluster subnet capacity, including temporary upgrade and scale-out needs. A private API endpoint requires an administration path from the VPC or a connected network; it does not automatically provide internet egress for workloads.

[S34] [S19]

Worked example, decision and pitfall synthesis

For the illustrative quote API, expose only the application route through an ALB and keep worker nodes in private subnets. Check DNS, ready endpoints, ALB target health, security groups, enforced policy and outbound dependencies independently. An empty Service selector or IP shortage can break traffic even while EC2 instances appear healthy.

[S05] [S19] [S20]

Further improvement synthesis

Track available subnet addresses and target-registration delay alongside CPU and memory. Test the complete request path rather than treating “Ingress exists” as proof of reachability.

[S19] [S20]
Keep this: Separate the configuration path from the request path, and treat IP space as a capacity constraint.
Check yourself: Does a client HTTP request pass through the Kubernetes API server?

Normally no. The API configures desired routing; application requests travel through the selected network and load-balancer dataplane.

Part 4 · Fundamentals

EKS responsibility boundaries and compute choices

Prerequisites: 01-control-loops

04 / Managed does not mean ownerless

A responsibility matrix compares standard EKS, managed node groups, Auto Mode and Fargate across AWS and team responsibilities.

Scroll the diagram sideways for readable labels.

Responsibility summary, not a full service contract. Self-managed and Hybrid Nodes are additional options. EKS control plane reliability and application reliability are different layers. [S14] [S15] [S16] [S37]

Learning objective and mental model verified

EKS supplies a managed Kubernetes control plane. AWS documents control-plane operation across three Availability Zones. Workload replicas, dependency resilience and application recovery still require design; a highly available API server does not make a single application replica highly available.

[S14] [S37]

Compute decision verified

Managed node groups simplify EC2 node provisioning and lifecycle operations. Self-managed nodes give more control and more maintenance. Auto Mode extends managed operation into compute autoscaling, networking, load balancing and block storage. Hybrid Nodes cover supported non-AWS compute environments. These are ownership choices, not application availability guarantees.

[S16] [S15]

Fargate boundary verified

Fargate provides per-Pod compute without managing EC2 worker instances. EKS Fargate does not support DaemonSets, privileged containers or GPUs; it uses private subnets and IP load-balancer targets. EBS cannot be mounted to Fargate Pods. EFS support has separate provisioning constraints.

[S30] [S21] [S29]

Worked example and trade-off synthesis

For a typical API with custom node agents, a managed node group may be easier to integrate than Fargate. Auto Mode is a candidate when supported defaults fit and reducing node operations matters. This is a conditional judgment: validate storage classes, IAM, agent support and disruption behavior before choosing.

[S15] [S16] [S30]

Pitfall and further improvement synthesis

“Managed” is not a universal promise of zero maintenance. Write a responsibility matrix for node releases, add-ons, identity, backups, alerts and workload compatibility. Review it whenever compute mode changes.

[S15] [S37]
Keep this: Select how much infrastructure AWS operates, without giving away application ownership.
Check yourself: Will Auto Mode make an application with one replica resilient to every node replacement?

No. It manages infrastructure lifecycle; your workload still needs suitable replication, disruption tolerance and resilient dependencies.

Part 5 · Applications

IAM, Kubernetes RBAC and workload identity

Prerequisites: 04-eks-model, 02-objects

05 / Three identity boundaries

Separate paths show engineer to Kubernetes API, Pod to AWS service APIs and product user to business records.

Scroll the diagram sideways for readable labels.

The first two lanes summarize documented platform mechanisms. The end-user lane is an illustrative application boundary, independent of cluster administration. [S17] [S18] [S32] [S11]

Learning objective and mechanism verified

An EKS access entry associates an IAM principal with cluster access. Permissions can use EKS access policies or Kubernetes groups connected to RBAC. Kubernetes RBAC controls cluster API operations; it does not authorize an end user to read a particular business record.

[S18] [S11]

Pod-to-AWS path verified

EKS Pod Identity maps a namespace and ServiceAccount to an IAM role. A supported AWS SDK uses temporary credentials through the default provider chain and the Pod Identity agent path. Auto Mode supplies the relevant integration. The documented Pod Identity worker scope is Linux EC2; do not assume it works on Fargate.

[S17]

Alternative and limits verified

IRSA uses a projected ServiceAccount OIDC token and AWS STS to obtain temporary role credentials. It remains a relevant option when its supported environment fits. Containers sharing a node are not a strong independent security boundary, and unrestricted instance metadata can expose the node role.

[S32] [S17]

Worked example synthesis

Give a document worker a dedicated ServiceAccount and an IAM role scoped to the required document prefix and actions. Give a deployer only the Kubernetes operations needed for its namespace. Separately enforce tenant and policy access in the application. This is a proposed design; exact IAM conditions depend on the data model.

[S17] [S18] [S11]

Decision, pitfall and improvement synthesis

Avoid attaching all workload permissions to the node role or handing every deployer cluster-admin. Inspect actual credential resolution, test both permitted and denied actions, and restrict instance metadata as appropriate. Verify SDK support and identity-agent access before debugging an S3 denial as a networking problem.

[S17] [S11]
Keep this: Keep engineer-to-cluster, Pod-to-AWS and user-to-product authorization separate.
Check yourself: Can a Pod Identity association grant kubectl permission to a human?

No. Pod Identity controls workload access to AWS APIs; human cluster access uses a separate authentication and authorization path.

Part 6 · Applications

Storage, state and recovery

Prerequisites: 02-objects, 04-eks-model

06 / Ask for storage, then mount it

A Pod references a PVC, which requests a StorageClass and CSI driver to provision a volume. The resulting PV binds to the PVC and is mounted.

Scroll the diagram sideways for readable labels.

Dynamic provisioning is shown conceptually. Binding may be delayed until scheduling. Fargate has different storage support; persistent storage is not a backup. [S06] [S21] [S29]

A disk can pin a Pod to a zone

An EBS volume in AZ A cannot simply attach to a replacement Pod in AZ B. Data recovery must be planned separately.

Scroll the diagram sideways for readable labels.

Simplified single-volume failure example. EBS attachment topology is not cross-zone database replication. EFS is a shared filesystem option with different semantics. [S39] [S29] [S13]

Learning objective and mechanism verified

A PVC asks for storage; a PV represents allocated storage; a StorageClass describes provisioning policy. A CSI integration connects Kubernetes to the storage system. Pods mount claims. Reclaim policy controls what happens when a claim is released; Pod replacement is a different event from deleting a claim.

[S06]

EKS choices verified

Use EBS integration for supported block-volume workloads and EFS integration when shared filesystem semantics fit. EBS is zonal; align scheduling with volume topology. Auto Mode uses the ebs.csi.eks.amazonaws.com provisioner, distinct from the standard EBS CSI provisioner. Fargate cannot mount EBS and supports EFS static rather than dynamic provisioning.

[S21] [S29] [S39]

Worked example and decision synthesis

In the insurance scenario, keep transaction state in a managed PostgreSQL service and documents in object storage rather than inside disposable API containers. This proposed boundary reduces the database-operator scope of the EKS platform team, but database networking, credentials, migrations, recovery and cost remain explicit responsibilities.

[S25] [S06]

Common pitfall synthesis

A persistent volume is not proof that a database can survive losing a zone. Likewise, copying Kubernetes manifests does not restore transaction data. Define which system owns database backups, document versions and workload configuration, then restore them together in a compatible sequence.

[S06] [S21]

Further improvement synthesis

Specify an acceptable recovery point (RPO: tolerated data loss) and recovery time (RTO: tolerated outage). Rehearse an isolated restore and validate business records, not just that a Pod becomes Running. These are proposed acceptance criteria, not measured outcomes.

[S06] [S22]
Keep this: Persistence, replication and backup solve different failure cases.
Check yourself: Is a Pod rescheduled to another AZ enough to recover an EBS-backed database?

Not by itself. The volume’s zone, data replication or restore strategy, database consistency and available capacity must all be considered.

Part 7 · Applications

Worked case: a multi-tenant insurance API on EKS

Prerequisites: 03-networking, 05-identities, 06-data

07 / Insurance API: a worked EKS design

Clients reach an ALB and quote or underwriting API Pods in private subnets. APIs use managed PostgreSQL, S3 and a durable queue for worker tasks.

Scroll the diagram sideways for readable labels.

Reasoned design example, not a description of a deployed customer platform. Arrows denote selected dependencies; the queue box represents asynchronous dispatch from the API to workers. [S05] [S17] [S19] [S25]

Objective and assumptions synthesis

This is an illustrative design exercise, not a claim about Tony’s deployed infrastructure. Assume HTTP quote and underwriting APIs, tenant-scoped records, uploaded documents and asynchronous tasks. Demand varies by business hours; the team wants bounded operations and observable recovery.

[S25] [S19]

Request and task flow synthesis

Clients reach an ALB. Quote and underwriting services run as replicated Deployments in private workload subnets. The API checks product authorization and tenant scope before reading PostgreSQL or issuing document access. Longer tasks use a durable queue and workers with scoped AWS identity.

[S05] [S17] [S25]

Engineering decision and trade-off synthesis

Separate latency-sensitive API capacity from retryable worker capacity. Keep an appropriate reliable baseline for serving traffic; consider interruption-tolerant compute only for tasks that can safely replay. This inference rests on different failure tolerance, not on a universal claim that workers should use Spot.

[S26] [S09]

Failure case synthesis

Suppose demand increases and HPA adds replicas while the database connection budget stays fixed. More Pods can amplify pool pressure and timeouts. For this exercise, bound pool sizes and task concurrency, apply backpressure and choose scaling signals that reflect useful throughput. Kubernetes cannot infer the database’s safe concurrency from replica count.

[S08] [S07]

Validate and improve synthesis

Measure completed quotes, error rate, tail latency, oldest task age and database saturation before and during a load test. Require idempotency for retried document or underwriting tasks. A useful next improvement is a failure drill with one unavailable dependency, evaluating whether the API fails quickly and recovers cleanly.

[S27] [S09]
Keep this: Keep the API stateless, bound downstream concurrency and make asynchronous work safe to retry.
Check yourself: Why might adding API replicas increase timeouts?

Each replica can add downstream connections and concurrent work. If the database or another dependency is saturated, extra callers increase contention rather than useful capacity.

Part 8 · Applications

Published cases: healthcare SaaS and ML search ranking

Prerequisites: 04-eks-model, 03-networking, 06-data

08 / Crossuite: elasticity across two layers

ALB routes to EKS Pods; HPA scales replicas, Karpenter capacity. RDS stores data; CloudWatch observes.

Scroll the diagram sideways for readable labels.

Published architecture elements, simplified for teaching; not the complete topology. [S25]

Booking.com: de-risk the migration

Inference moved first, then ranking API and models, then hybrid serving. Dedicated EKS capacity was benchmarked.

Scroll the diagram sideways for readable labels.

Published migration phases and result; not a transferable latency guarantee. [S24]

Crossuite: reported architecture and result verified

AWS’s Crossuite case describes a healthcare application moving from manually managed Kubernetes to EKS. Its architecture uses ALB, HPA for services, Karpenter for EC2 capacity, RDS PostgreSQL and CloudWatch. The report places the production switch in May 2024 and reports 30% lower overall costs and 99.9% uptime. A headline benefits panel displays a different uptime figure; this atlas uses the narrative’s 99.9% and treats it as a customer report, not an independently verified SLA.

[S25]

Crossuite: engineering lesson synthesis

Inference: replica scaling and node provisioning solve different bottlenecks, and external database dependencies must survive the migration. Reuse the two-layer capacity pattern only after measuring startup delay and dependency limits. The reported savings lack a controlled comparison for your workload.

[S25]

Booking.com: reported architecture and result verified

AWS’s Booking.com case describes a dedicated EKS environment for search ranking and model-serving flexibility. Migration proceeded through inference separation, ranking API and model movement, then hybrid serving. The company benchmarked instance types and the report states approximately 40 ms latency for 99.9% of requests. The same article places some other model workloads on SageMaker; it does not claim every ML workload runs on EKS.

[S24]

Booking.com: engineering lesson synthesis

Inference: test the riskiest interface before moving the whole serving system. A phased migration and workload-specific instance tests make the outcome reviewable. The published latency is not an EKS service guarantee, nor evidence that another model or dataset reaches the same result.

[S24]

Decision, pitfall and next improvement synthesis

For your own migration, define a baseline, representative peak load, acceptable regression and a reversible traffic cutover. Record the old and new compute/network/storage conditions. Avoid attributing revenue growth or every cost change solely to Kubernetes; the case descriptions report a broader modernization.

[S25] [S24]
Keep this: Transfer the decision pattern; do not copy the published performance number as your own forecast.
Check yourself: Which result can be safely copied from these cases?

The architecture and evaluation patterns are useful hypotheses. Their cost and latency figures require your own measurement under your workload and operating conditions.

Part 9 · Best practices

Probes, rollouts, disruption and availability

Prerequisites: 02-objects, 03-networking, 04-eks-model

09 / Three probes, three different decisions

Startup protects initialization, readiness controls normal Service endpoint participation, and liveness restarts unhealthy containers. Database failure should not automatically trigger liveness failure.

Scroll the diagram sideways for readable labels.

Probe actions summarize Kubernetes behavior. The database-outage response is a design recommendation; readiness must reflect whether a Pod can provide useful service. [S10]

Availability needs several independent controls

Three application replicas occupy different zones. Separate cards distinguish topology spread, voluntary eviction budgets and rolling update settings.

Scroll the diagram sideways for readable labels.

Placement, eviction and application rollout are different controls. A PDB does not constrain the Deployment rollout controller or prevent involuntary loss. [S09] [S13] [S04]

Learning objective and probe mechanism verified

A startup probe protects slow initialization before normal liveness and readiness checks. Readiness failure removes normal ready participation in Service endpoints; it does not itself restart the container. Liveness failure eventually restarts a container according to its restart policy. These checks answer different questions.

[S10]

Rollout and eviction mechanics verified

Deployment rolling updates use maxSurge and maxUnavailable. PDBs constrain eligible voluntary evictions, including compliant node drains. PDBs do not prevent node or zone failures, direct Pod deletion, or constrain Deployment rollout logic. Topology spread can distribute replicas across zones or hosts; strict placement rules can also leave Pods Pending.

[S04] [S09] [S13]

Worked baseline synthesis

For an illustrative quote API, start with three replicas, one surge Pod and zero unavailable Pods during a rollout. Use a PDB requiring two available replicas and spread placement across eligible zones and hosts. These are exercise settings, not a claim that three replicas meet every SLA. Reserve surge capacity and test the load remaining replicas can handle.

[S04] [S09] [S13]

Pitfall and shutdown decision synthesis

Do not make every transient database failure kill a healthy process. Design local liveness and useful-service readiness deliberately. During termination, the application must stop accepting new work and finish or release existing work within its grace period; align this with load-balancer draining. This is application design guidance, not a promise of immediate endpoint convergence.

[S10] [S20]

Further improvement synthesis

Use the bundled workload example as a review specimen. Tune startup budget, readiness conditions, request duration and drain behavior using a load test. Verify failed releases and replica loss at the user-facing endpoint. A rollout remaining available says little about a backwards-incompatible database migration.

[S10] [S04]
Keep this: Readiness controls participation; liveness restarts; placement and rollout controls govern other failure paths.
Check yourself: Can a PDB guarantee that a Deployment rollout never takes down too many Pods?

No. Deployment rollout availability is controlled by its rollout strategy. PDB governs eligible eviction paths and does not prevent involuntary failures.

Part 10 · Best practices

Capacity: requests, HPA and node provisioning

Prerequisites: 04-eks-model, 09-release-reliability

10 / Replicas and nodes are separate loops

Metrics lead HPA to request more replicas. Unschedulable Pods lead a node provisioner to create capacity, enabling scheduling.

Scroll the diagram sideways for readable labels.

Basic HPA ratio example assumes usable metrics and unconstrained bounds. The node capacity path is logical, not a latency guarantee. [S08] [S26] [S15]

Requests are placement inputs; limits are ceilings

Resource requests inform scheduler placement, limits affect enforcement, and HPA CPU utilization divides actual CPU usage by the request.

Scroll the diagram sideways for readable labels.

Illustrative single-container arithmetic. Sidecars, defaults and admission policies can affect actual Pod requests and HPA accounting. [S07] [S08]

Learning objective and mechanism verified

The scheduler uses resource requests for placement. CPU limits can cause throttling; memory limits can lead to OOM termination. Requests are not exclusive allocation of the specified RAM or CPU. HPA CPU utilization is relative to CPU requests, not the whole node’s CPU capacity.

[S07] [S08]

Two scaling loops synthesis

HPA adjusts workload replica count. A node provisioner such as Karpenter evaluates unschedulable workload requirements and creates suitable capacity. Auto Mode supplies managed compute autoscaling. Placement may still be blocked by topology, taints, IP space, volume constraints or unavailable cloud capacity.

[S26] [S15] [S19]

Worked arithmetic verified

Simplifying the HPA algorithm to usable metrics and unconstrained bounds, three replicas at 90% utilization against a 60% target imply ceil(3 × 90/60) = 5 replicas. A container using 100m CPU against a 250m request reports 40% utilization. Real decisions include missing metrics, readiness, tolerance and stabilization.

[S08]

Decision and common pitfall synthesis

For a queue worker, CPU may not reflect an aging backlog; for a Node.js API, downstream waits can dominate latency. Compare useful throughput, queue age and resource saturation before choosing a metric. Avoid two controllers fighting over the same replica field or blindly increasing resources to hide a leak. These are conditional design judgments to validate under load.

[S08] [S07]

Interruption and improvement synthesis

For supported retryable workloads, diversify Spot capacity and configure interruption handling. A replacement node is not guaranteed to become ready before interruption. Measure image pull, model warmup, scheduling and ready-to-serve delays; choose baseline capacity and scale-down stabilization using those observations.

[S26] [S08]
Keep this: Measure the bottleneck, then scale the layer that can relieve it.
Check yourself: HPA requests five replicas but only three run. What do you inspect next?

Read the Pending Pod events and node/IP/storage capacity. HPA requests replicas; it does not itself supply EC2 nodes or remove incompatible placement constraints.

Part 11 · Best practices

Tenancy, least privilege and network policy

Prerequisites: 05-identities, 03-networking

11 / Isolation is a stack, not a namespace label

Business data, Kubernetes API, Pod networking and compute each need their own isolation control. A namespace alone does not provide all four.

Scroll the diagram sideways for readable labels.

Platform controls summarize Kubernetes guidance. Tenant-scoped business authorization is a reasoned requirement for the worked SaaS scenario. [S11] [S12] [S33]

Learning objective and isolation model verified

Multi-tenancy requires both control-plane and data-plane isolation. Namespaces scope many API resources but do not create independent kernels. Dedicated nodes reduce co-location but can still share cluster services; stronger isolation may justify sandboxed execution or separate clusters, with extra cost and operations.

[S33] [S11]

Network-policy mechanism verified

Without applicable NetworkPolicies, Pod ingress and egress are allowed by default. Policies require a supporting enforcement implementation and primarily express L4 rules. Default-deny egress also blocks DNS unless allowed. IAM, security groups, NetworkPolicy and business authorization cover different scopes.

[S12]

Worked example synthesis

In the insurance exercise, use tenant-scoped application queries even if each team has its own namespace. Restrict deployers and service accounts to their required API operations. Add explicit dependency flows before applying default-deny. On a Pod Identity workload, include the identity-agent credential path; on private AWS access, verify required service endpoints.

[S11] [S17] [S12]

Decision and pitfall synthesis

A tenant allowed to run arbitrary untrusted code changes the threat model. Namespace-only separation is too weak for that assumption; evaluate sandbox or dedicated-cluster designs. A toleration only permits scheduling onto a tainted node and does not, by itself, force exclusive placement. Enforce the intended placement and admission constraints.

[S33]

Further improvement synthesis

Write a permission matrix and a network dependency map, then test denied access as well as allowed access. Verify the real policy implementation on your chosen EKS compute mode. Treat base64 values and read access to Secret objects as sensitive, and keep plaintext secret manifests out of the delivered examples.

[S36] [S12] [S11]
Keep this: A namespace is an organizational boundary; build the actual trust boundary explicitly.
Check yourself: Does “one namespace per customer” enforce customer record isolation?

No. The application must enforce record access; Kubernetes namespaces scope API resources and need additional control, network and compute protections.

Part 12 · Best practices

Observe and diagnose the right layer

Prerequisites: 03-networking, 09-release-reliability, 10-scaling

12 / From a symptom to the next useful check

Four diagnostic branches map Pending, CrashLoopBackOff, errors and latency to checks and possible causes.

Scroll the diagram sideways for readable labels.

Reasoned troubleshooting flow based on component responsibilities and observability primitives. A symptom can have more than one root cause. [S27] [S07] [S05] [S19]

Signals serve different consumers

One pipeline collects app and node telemetry for dashboards. A separate resource Metrics API provides CPU and memory for HPA and kubectl top.

Scroll the diagram sideways for readable labels.

Logical pipelines; collection components depend on compute mode. EKS does not expose every managed control-plane metric endpoint as a self-managed cluster would. [S27] [S28]

Learning objective and signals verified

Metrics reveal trends and saturation, logs describe local events and traces connect request steps. The resource Metrics API used by HPA and kubectl top is a separate capability from a full monitoring pipeline. Collecting CPU usage does not provide historical logs or distributed traces.

[S27]

EKS control-plane view verified

EKS can export API, audit, authenticator, scheduler and controller-manager logs to CloudWatch. Export is disabled by default and each type is enabled separately. Audit records help investigate who changed cluster resources; retention and ingestion costs require planning.

[S28]

Worked incident synthesis

Imagine the quote API becomes slow after adding replicas. Start with latency and error rate, then inspect traces for database waits, connection counts and per-Pod concurrency. Compare CPU throttling, restarts and queue delay. This hypothetical incident tests whether the database bottleneck diagnosis fits; it is not a claim that every latency spike has that cause.

[S27] [S07]

Diagnostic commands: review-only examples synthesis

kubectl -n insurance get pods -o wide
kubectl -n insurance describe pod POD_NAME
kubectl -n insurance logs POD_NAME --previous
kubectl -n insurance get events --sort-by=.metadata.creationTimestamp
kubectl -n insurance get endpointslices -l kubernetes.io/service-name=quotes
kubectl -n insurance top pods

Replace POD_NAME. Previous logs require a prior container instance; top requires a working Metrics API provider.
[S27] [S05]

Decision, pitfall and further improvement synthesis

Choose retention and labels around an actual troubleshooting question. Avoid putting raw tokens or unbounded tenant identifiers into telemetry. Alert on sustained user-facing errors and latency, then include placement, IP capacity and dependencies as diagnostic context. Proposed objective: find a failed release’s cause from retained evidence without needing to reproduce it.

[S27] [S28]
Keep this: Start from user-visible failure, then follow evidence down the request path.
Check yourself: If kubectl top works, are application traces and audit logs automatically available?

No. Resource metrics, application telemetry and control-plane logging are different pipelines that require their own configuration and retention.

Part 13 · Further improvements

Upgrades, recovery paths and the full bill

Prerequisites: 04-eks-model, 09-release-reliability, 12-observability

13 / Upgrade with compatibility checkpoints

Inventory and rehearse, prepare compatibility, upgrade control plane, roll nodes and add-ons, validate behavior, then choose a documented recovery path if needed.

Scroll the diagram sideways for readable labels.

A high-level workflow, not an executable runbook. Some add-ons need preparation before the control-plane step. Rollback has eligibility and compatibility limits and does not revert data or all add-ons. [S22] [S31] [S38]

Budget the whole platform

Five cost rows cover cluster support, compute, networking, storage and operations. Standard support is 73 dollars for a 730-hour illustrative month, extended support 438 dollars.

Scroll the diagram sideways for readable labels.

Published standard cluster-support rates retrieved on 2026-10-10. Region, partition, features, support tiers and resource usage can change the bill. No compute-price estimate is asserted. [S23]

Learning objective and upgrade mechanism verified

Inventory removed APIs, admission webhooks, CRDs, nodes and add-ons. Rehearse in a representative non-production environment. Follow the EKS procedure for the control plane and compatible components. Standard add-on updates are not automatically completed merely because the control plane was upgraded; Auto Mode changes some operational responsibilities.

[S22] [S31] [S15]

Current recovery boundary verified

The retrieved EKS documentation describes a conditional version rollback: initiate within seven days of an in-place upgrade, return only to the previous minor version and satisfy eligibility and compatibility checks. It preserves data and does not roll back all add-ons or non-Auto-Mode compute for you. Check current supported versions and the detailed prerequisites before relying on it. After that window, a new cluster and migration may be needed.

[S38] [S31]

Worked cost calculation verified

Published EKS cluster-support charges are $0.10 per cluster-hour for standard support and $0.60 for extended support. For an illustrative 730-hour month: $73 or $438 for this charge alone. EC2 or Fargate compute, Auto Mode charges where applicable, networking, storage, telemetry and optional features are additional. These are not an estimate for a complete production environment.

[S23]

Decision and pitfall synthesis

Use a tested in-place upgrade when its compatibility and recovery constraints fit. Evaluate blue/green clusters when traffic migration and duplicated infrastructure are worth the additional cost. Neither approach automatically reverses a destructive database migration. Keep application and data rollback plans separate from control-plane recovery.

[S22] [S38]

Further improvement synthesis

Maintain an upgrade calendar and a cost allocation map. Track useful cost per completed business transaction alongside reliability metrics. Evaluate NAT, data transfer and telemetry retention rather than assuming worker instance price is the entire bill. Revisit region and account-specific pricing before a purchase decision.

[S23] [S22]
Keep this: Budget compatibility work and recovery before the support calendar makes the decision for you.
Check yourself: Does reverting a cluster version revert a changed database schema?

No. Version rollback changes eligible control-plane components, not application data migrations. Database compatibility and recovery need an independent plan.

Part 14 · Further improvements

A practical learning path and improvement experiments

Prerequisites: 13-lifecycle, 11-security, 10-scaling

14 / Improve a baseline with evidence

A baseline is tested with replica loss, traffic growth and isolated recovery drills, each tied to observable success criteria.

Scroll the diagram sideways for readable labels.

Proposed learning exercises for a disposable environment. Measurement targets are defined by you; no production experiment or measured result is claimed. [S09] [S08] [S22]

Learning objective and mental model synthesis

Treat the platform as a set of contracts: desired state, traffic delivery, permissions, persistence, capacity and recovery. Improve one weak contract at a time. This is a synthesis of the preceding mechanisms; the right next experiment depends on your workload and operating responsibility.

[S02] [S09] [S22]

Lab 1: local fundamentals synthesis

In a disposable test environment, deploy a small HTTP service, a Service and the supplied manifest specimen after replacing the image and implementing its endpoints. Inspect ownership and labels, replace one Pod and observe how the controller restores the count. Record time until useful traffic resumes. Do not run the specimen unchanged against production.

[S02] [S04] [S05]

Lab 2: capacity and failures synthesis

Increase load gradually, compare resource requests with measured demand, then inspect the HPA metric and node provisioning path. Introduce one dependency failure and one failed rollout in the test environment. A good experiment specifies demand, duration, acceptable latency/error behavior and what evidence would reject the hypothesis.

[S08] [S26] [S10]

Lab 3: EKS integration synthesis

Choose managed nodes or Auto Mode according to the stated responsibility matrix. Verify private administration access, identity resolution, routing, storage compatibility and subnet headroom. Test an allowed and denied AWS action. Costs accrue for an EKS environment; plan resource cleanup as part of the exercise.

[S19] [S17] [S15] [S23]

Lab 4: recovery and next topics synthesis

Restore data and workload configuration into an isolated compatible environment; validate records and measure RPO/RTO against your chosen objectives. Then study workload-specific topics: operators for data services, progressive delivery, policy admission, multi-cluster recovery or queue-based scaling. Add those only when the baseline measurements expose a real need.

[S06] [S22] [S27]

Evidence limits open-question

Examples were structurally reviewed and parsed where supplied as YAML; they were not deployed to an AWS account or live cluster. No load-test, recovery-time or savings result was measured. Official sources were retrieved on 2026-10-10; living docs, region support, feature constraints and pricing need rechecking before implementation.

Keep this: Make every improvement a hypothesis with a measured success condition.
Check yourself: What would make the proposed HPA improvement fail its evaluation?

If more replicas do not improve useful throughput or meet latency/error objectives, or if startup lag and dependency saturation negate the expected benefit, the experiment rejects or narrows the hypothesis.

Sources & further reading

  1. [S01] Kubernetes cluster architecture

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: API server, etcd, scheduler, controllers, kubelet and runtime responsibilities

    Read the linked primary source for implementation details and current constraints.

  2. [S02] Kubernetes controllers

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Controllers compare desired and observed state and act through APIs

    Read the linked primary source for implementation details and current constraints.

  3. [S03] Workloads

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Deployment, StatefulSet, DaemonSet, Job and CronJob use cases

    Read the linked primary source for implementation details and current constraints.

  4. [S04] Deployments

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: ReplicaSet ownership, rolling updates, surge and unavailable settings

    Read the linked primary source for implementation details and current constraints.

  5. [S05] Service

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service selection, stable access and EndpointSlices

    Read the linked primary source for implementation details and current constraints.

  6. [S06] Persistent volumes

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: PVC, PV, provisioning and reclaim lifecycle

    Read the linked primary source for implementation details and current constraints.

  7. [S07] Resource management

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Requests, limits, scheduling, CPU throttling and memory enforcement

    Read the linked primary source for implementation details and current constraints.

  8. [S08] Horizontal Pod Autoscaling

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metric ratio algorithm, CPU utilization relative to requests, metric pipeline

    Read the linked primary source for implementation details and current constraints.

  9. [S09] Disruptions

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: PDB limits eviction; involuntary failure and rollout limitations

    Read the linked primary source for implementation details and current constraints.

  10. [S10] Liveness, readiness and startup probes

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Probe actions and startup protection

    Read the linked primary source for implementation details and current constraints.

  11. [S11] RBAC good practices

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Least privilege and weak namespace boundaries

    Read the linked primary source for implementation details and current constraints.

  12. [S12] Network policies

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: L4 policy, enforcement dependency, default allowance, DNS egress

    Read the linked primary source for implementation details and current constraints.

  13. [S13] Topology spread constraints

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Zone and host distribution, strict scheduling constraints

    Read the linked primary source for implementation details and current constraints.

  14. [S14] What is Amazon EKS?

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Managed Kubernetes and AWS integration

    Read the linked primary source for implementation details and current constraints.

  15. [S15] EKS Auto Mode

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Managed compute, networking, load balancing and block storage

    Read the linked primary source for implementation details and current constraints.

  16. [S16] EKS compute options

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Auto Mode, managed nodes, self-managed nodes, Fargate, Hybrid Nodes

    Read the linked primary source for implementation details and current constraints.

  17. [S17] EKS Pod Identity

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service account role associations, agent, SDK support and Linux EC2 scope

    Read the linked primary source for implementation details and current constraints.

  18. [S18] EKS access entries

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: IAM principal access with EKS access policies or Kubernetes groups

    Read the linked primary source for implementation details and current constraints.

  19. [S19] VPC and subnet considerations

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Cluster subnets, private endpoint access, upgrade IP headroom

    Read the linked primary source for implementation details and current constraints.

  20. [S20] AWS Load Balancer Controller

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Ingress to ALB and Service LoadBalancer to NLB integration

    Read the linked primary source for implementation details and current constraints.

  21. [S21] EBS CSI integration

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: EBS support limits and different Auto Mode provisioner

    Read the linked primary source for implementation details and current constraints.

  22. [S22] Cluster upgrade best practices

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Compatibility, removed APIs, node upgrades and add-on coordination

    Read the linked primary source for implementation details and current constraints.

  23. [S23] Amazon EKS pricing

    AWS · pricing · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Standard and extended support charges and separate compute costs

    Read the linked primary source for implementation details and current constraints.

  24. [S24] Booking.com search ranking on Amazon EKS

    AWS / Booking.com · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Dedicated ranking cluster, phased migration and reported latency

    Read the linked primary source for implementation details and current constraints.

  25. [S25] Crossuite migration to EKS

    AWS / Crossuite · customer case study · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: HPA, Karpenter, ALB, RDS PostgreSQL, CloudWatch and reported outcomes

    Read the linked primary source for implementation details and current constraints.

  26. [S26] Karpenter best practices

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Pending workload capacity, diversification and interruption handling

    Read the linked primary source for implementation details and current constraints.

  27. [S27] Kubernetes observability

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metrics, logs, traces and resource Metrics API distinction

    Read the linked primary source for implementation details and current constraints.

  28. [S28] EKS control plane logs

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Optional API, audit, authenticator, scheduler and controller logging

    Read the linked primary source for implementation details and current constraints.

  29. [S29] EFS CSI integration

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Shared filesystem integration and Fargate static provisioning constraint

    Read the linked primary source for implementation details and current constraints.

  30. [S30] AWS Fargate for EKS

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: No DaemonSets, privileged containers or GPUs; private subnets and IP targets

    Read the linked primary source for implementation details and current constraints.

  31. [S31] Update an EKS cluster

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Control plane and component upgrade sequencing

    Read the linked primary source for implementation details and current constraints.

  32. [S32] IAM roles for service accounts

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: OIDC service account tokens and STS temporary credentials

    Read the linked primary source for implementation details and current constraints.

  33. [S33] Kubernetes multi-tenancy

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Control and data plane isolation; shared kernels and dedicated clusters

    Read the linked primary source for implementation details and current constraints.

  34. [S34] Amazon VPC CNI

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: VPC private addresses per Pod and Auto Mode networking ownership

    Read the linked primary source for implementation details and current constraints.

  35. [S35] ConfigMaps

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Non-confidential runtime configuration

    Read the linked primary source for implementation details and current constraints.

  36. [S36] Kubernetes Secret good practices

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Base64 is encoding, least privilege and encryption considerations

    Read the linked primary source for implementation details and current constraints.

  37. [S37] EKS reliability

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Control plane across three AZs; shared data plane responsibility

    Read the linked primary source for implementation details and current constraints.

  38. [S38] EKS version rollback

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Current documented seven-day, one-minor-version rollback window and limitations

    Read the linked primary source for implementation details and current constraints.

  39. [S39] Amazon EBS volumes

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Volumes and attached instances must be in the same Availability Zone

    Read the linked primary source for implementation details and current constraints.