Part 10 · Best practices

Capacity: requests, HPA and node provisioning

Prerequisites: 04-eks-model, 09-release-reliability

10 / Replicas and nodes are separate loops

Metrics lead HPA to request more replicas. Unschedulable Pods lead a node provisioner to create capacity, enabling scheduling.

Scroll the diagram sideways for readable labels.

Basic HPA ratio example assumes usable metrics and unconstrained bounds. The node capacity path is logical, not a latency guarantee. [S08] [S26] [S15]

Requests are placement inputs; limits are ceilings

Resource requests inform scheduler placement, limits affect enforcement, and HPA CPU utilization divides actual CPU usage by the request.

Scroll the diagram sideways for readable labels.

Illustrative single-container arithmetic. Sidecars, defaults and admission policies can affect actual Pod requests and HPA accounting. [S07] [S08]

Learning objective and mechanism verified

The scheduler uses resource requests for placement. CPU limits can cause throttling; memory limits can lead to OOM termination. Requests are not exclusive allocation of the specified RAM or CPU. HPA CPU utilization is relative to CPU requests, not the whole node’s CPU capacity.

[S07] [S08]

Two scaling loops synthesis

HPA adjusts workload replica count. A node provisioner such as Karpenter evaluates unschedulable workload requirements and creates suitable capacity. Auto Mode supplies managed compute autoscaling. Placement may still be blocked by topology, taints, IP space, volume constraints or unavailable cloud capacity.

[S26] [S15] [S19]

Worked arithmetic verified

Simplifying the HPA algorithm to usable metrics and unconstrained bounds, three replicas at 90% utilization against a 60% target imply ceil(3 × 90/60) = 5 replicas. A container using 100m CPU against a 250m request reports 40% utilization. Real decisions include missing metrics, readiness, tolerance and stabilization.

[S08]

Decision and common pitfall synthesis

For a queue worker, CPU may not reflect an aging backlog; for a Node.js API, downstream waits can dominate latency. Compare useful throughput, queue age and resource saturation before choosing a metric. Avoid two controllers fighting over the same replica field or blindly increasing resources to hide a leak. These are conditional design judgments to validate under load.

[S08] [S07]

Interruption and improvement synthesis

For supported retryable workloads, diversify Spot capacity and configure interruption handling. A replacement node is not guaranteed to become ready before interruption. Measure image pull, model warmup, scheduling and ready-to-serve delays; choose baseline capacity and scale-down stabilization using those observations.

[S26] [S08]
Keep this: Measure the bottleneck, then scale the layer that can relieve it.
Check yourself: HPA requests five replicas but only three run. What do you inspect next?

Read the Pending Pod events and node/IP/storage capacity. HPA requests replicas; it does not itself supply EC2 nodes or remove incompatible placement constraints.

Sources & further reading

  1. [S07] Resource management

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Requests, limits, scheduling, CPU throttling and memory enforcement

    Read the linked primary source for implementation details and current constraints.

  2. [S08] Horizontal Pod Autoscaling

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metric ratio algorithm, CPU utilization relative to requests, metric pipeline

    Read the linked primary source for implementation details and current constraints.

  3. [S15] EKS Auto Mode

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Managed compute, networking, load balancing and block storage

    Read the linked primary source for implementation details and current constraints.

  4. [S19] VPC and subnet considerations

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Cluster subnets, private endpoint access, upgrade IP headroom

    Read the linked primary source for implementation details and current constraints.

  5. [S26] Karpenter best practices

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Pending workload capacity, diversification and interruption handling

    Read the linked primary source for implementation details and current constraints.