Part 14 · Further improvements

A practical learning path and improvement experiments

Prerequisites: 13-lifecycle, 11-security, 10-scaling

14 / Improve a baseline with evidence

A baseline is tested with replica loss, traffic growth and isolated recovery drills, each tied to observable success criteria.

Scroll the diagram sideways for readable labels.

Proposed learning exercises for a disposable environment. Measurement targets are defined by you; no production experiment or measured result is claimed. [S09] [S08] [S22]

Learning objective and mental model synthesis

Treat the platform as a set of contracts: desired state, traffic delivery, permissions, persistence, capacity and recovery. Improve one weak contract at a time. This is a synthesis of the preceding mechanisms; the right next experiment depends on your workload and operating responsibility.

[S02] [S09] [S22]

Lab 1: local fundamentals synthesis

In a disposable test environment, deploy a small HTTP service, a Service and the supplied manifest specimen after replacing the image and implementing its endpoints. Inspect ownership and labels, replace one Pod and observe how the controller restores the count. Record time until useful traffic resumes. Do not run the specimen unchanged against production.

[S02] [S04] [S05]

Lab 2: capacity and failures synthesis

Increase load gradually, compare resource requests with measured demand, then inspect the HPA metric and node provisioning path. Introduce one dependency failure and one failed rollout in the test environment. A good experiment specifies demand, duration, acceptable latency/error behavior and what evidence would reject the hypothesis.

[S08] [S26] [S10]

Lab 3: EKS integration synthesis

Choose managed nodes or Auto Mode according to the stated responsibility matrix. Verify private administration access, identity resolution, routing, storage compatibility and subnet headroom. Test an allowed and denied AWS action. Costs accrue for an EKS environment; plan resource cleanup as part of the exercise.

[S19] [S17] [S15] [S23]

Lab 4: recovery and next topics synthesis

Restore data and workload configuration into an isolated compatible environment; validate records and measure RPO/RTO against your chosen objectives. Then study workload-specific topics: operators for data services, progressive delivery, policy admission, multi-cluster recovery or queue-based scaling. Add those only when the baseline measurements expose a real need.

[S06] [S22] [S27]

Evidence limits open-question

Examples were structurally reviewed and parsed where supplied as YAML; they were not deployed to an AWS account or live cluster. No load-test, recovery-time or savings result was measured. Official sources were retrieved on 2026-10-10; living docs, region support, feature constraints and pricing need rechecking before implementation.

Keep this: Make every improvement a hypothesis with a measured success condition.
Check yourself: What would make the proposed HPA improvement fail its evaluation?

If more replicas do not improve useful throughput or meet latency/error objectives, or if startup lag and dependency saturation negate the expected benefit, the experiment rejects or narrows the hypothesis.

Sources & further reading

  1. [S02] Kubernetes controllers

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Controllers compare desired and observed state and act through APIs

    Read the linked primary source for implementation details and current constraints.

  2. [S04] Deployments

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: ReplicaSet ownership, rolling updates, surge and unavailable settings

    Read the linked primary source for implementation details and current constraints.

  3. [S05] Service

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service selection, stable access and EndpointSlices

    Read the linked primary source for implementation details and current constraints.

  4. [S06] Persistent volumes

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: PVC, PV, provisioning and reclaim lifecycle

    Read the linked primary source for implementation details and current constraints.

  5. [S08] Horizontal Pod Autoscaling

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metric ratio algorithm, CPU utilization relative to requests, metric pipeline

    Read the linked primary source for implementation details and current constraints.

  6. [S09] Disruptions

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: PDB limits eviction; involuntary failure and rollout limitations

    Read the linked primary source for implementation details and current constraints.

  7. [S10] Liveness, readiness and startup probes

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Probe actions and startup protection

    Read the linked primary source for implementation details and current constraints.

  8. [S15] EKS Auto Mode

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Managed compute, networking, load balancing and block storage

    Read the linked primary source for implementation details and current constraints.

  9. [S17] EKS Pod Identity

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service account role associations, agent, SDK support and Linux EC2 scope

    Read the linked primary source for implementation details and current constraints.

  10. [S19] VPC and subnet considerations

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Cluster subnets, private endpoint access, upgrade IP headroom

    Read the linked primary source for implementation details and current constraints.

  11. [S22] Cluster upgrade best practices

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Compatibility, removed APIs, node upgrades and add-on coordination

    Read the linked primary source for implementation details and current constraints.

  12. [S23] Amazon EKS pricing

    AWS · pricing · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Standard and extended support charges and separate compute costs

    Read the linked primary source for implementation details and current constraints.

  13. [S26] Karpenter best practices

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Pending workload capacity, diversification and interruption handling

    Read the linked primary source for implementation details and current constraints.

  14. [S27] Kubernetes observability

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metrics, logs, traces and resource Metrics API distinction

    Read the linked primary source for implementation details and current constraints.