EKS responsibility boundaries and compute choices
Prerequisites: 01-control-loops
04 / Managed does not mean ownerless
Scroll the diagram sideways for readable labels.
Learning objective and mental model verified
EKS supplies a managed Kubernetes control plane. AWS documents control-plane operation across three Availability Zones. Workload replicas, dependency resilience and application recovery still require design; a highly available API server does not make a single application replica highly available.
[S14] [S37]Compute decision verified
Managed node groups simplify EC2 node provisioning and lifecycle operations. Self-managed nodes give more control and more maintenance. Auto Mode extends managed operation into compute autoscaling, networking, load balancing and block storage. Hybrid Nodes cover supported non-AWS compute environments. These are ownership choices, not application availability guarantees.
[S16] [S15]Fargate boundary verified
Fargate provides per-Pod compute without managing EC2 worker instances. EKS Fargate does not support DaemonSets, privileged containers or GPUs; it uses private subnets and IP load-balancer targets. EBS cannot be mounted to Fargate Pods. EFS support has separate provisioning constraints.
[S30] [S21] [S29]Worked example and trade-off synthesis
For a typical API with custom node agents, a managed node group may be easier to integrate than Fargate. Auto Mode is a candidate when supported defaults fit and reducing node operations matters. This is a conditional judgment: validate storage classes, IAM, agent support and disruption behavior before choosing.
[S15] [S16] [S30]Pitfall and further improvement synthesis
“Managed” is not a universal promise of zero maintenance. Write a responsibility matrix for node releases, add-ons, identity, backups, alerts and workload compatibility. Review it whenever compute mode changes.
[S15] [S37]Check yourself: Will Auto Mode make an application with one replica resilient to every node replacement?
No. It manages infrastructure lifecycle; your workload still needs suitable replication, disruption tolerance and resilient dependencies.