Capacity: requests, HPA and node provisioning
Prerequisites: 04-eks-model, 09-release-reliability
10 / Replicas and nodes are separate loops
Scroll the diagram sideways for readable labels.
Requests are placement inputs; limits are ceilings
Scroll the diagram sideways for readable labels.
Learning objective and mechanism verified
The scheduler uses resource requests for placement. CPU limits can cause throttling; memory limits can lead to OOM termination. Requests are not exclusive allocation of the specified RAM or CPU. HPA CPU utilization is relative to CPU requests, not the whole node’s CPU capacity.
[S07] [S08]Two scaling loops synthesis
HPA adjusts workload replica count. A node provisioner such as Karpenter evaluates unschedulable workload requirements and creates suitable capacity. Auto Mode supplies managed compute autoscaling. Placement may still be blocked by topology, taints, IP space, volume constraints or unavailable cloud capacity.
[S26] [S15] [S19]Worked arithmetic verified
Simplifying the HPA algorithm to usable metrics and unconstrained bounds, three replicas at 90% utilization against a 60% target imply ceil(3 × 90/60) = 5 replicas. A container using 100m CPU against a 250m request reports 40% utilization. Real decisions include missing metrics, readiness, tolerance and stabilization.
[S08]Decision and common pitfall synthesis
For a queue worker, CPU may not reflect an aging backlog; for a Node.js API, downstream waits can dominate latency. Compare useful throughput, queue age and resource saturation before choosing a metric. Avoid two controllers fighting over the same replica field or blindly increasing resources to hide a leak. These are conditional design judgments to validate under load.
[S08] [S07]Interruption and improvement synthesis
For supported retryable workloads, diversify Spot capacity and configure interruption handling. A replacement node is not guaranteed to become ready before interruption. Measure image pull, model warmup, scheduling and ready-to-serve delays; choose baseline capacity and scale-down stabilization using those observations.
[S26] [S08]Check yourself: HPA requests five replicas but only three run. What do you inspect next?
Read the Pending Pod events and node/IP/storage capacity. HPA requests replicas; it does not itself supply EC2 nodes or remove incompatible placement constraints.