Part 9 · Best practices

Probes, rollouts, disruption and availability

Prerequisites: 02-objects, 03-networking, 04-eks-model

09 / Three probes, three different decisions

Startup protects initialization, readiness controls normal Service endpoint participation, and liveness restarts unhealthy containers. Database failure should not automatically trigger liveness failure.

Scroll the diagram sideways for readable labels.

Probe actions summarize Kubernetes behavior. The database-outage response is a design recommendation; readiness must reflect whether a Pod can provide useful service. [S10]

Availability needs several independent controls

Three application replicas occupy different zones. Separate cards distinguish topology spread, voluntary eviction budgets and rolling update settings.

Scroll the diagram sideways for readable labels.

Placement, eviction and application rollout are different controls. A PDB does not constrain the Deployment rollout controller or prevent involuntary loss. [S09] [S13] [S04]

Learning objective and probe mechanism verified

A startup probe protects slow initialization before normal liveness and readiness checks. Readiness failure removes normal ready participation in Service endpoints; it does not itself restart the container. Liveness failure eventually restarts a container according to its restart policy. These checks answer different questions.

[S10]

Rollout and eviction mechanics verified

Deployment rolling updates use maxSurge and maxUnavailable. PDBs constrain eligible voluntary evictions, including compliant node drains. PDBs do not prevent node or zone failures, direct Pod deletion, or constrain Deployment rollout logic. Topology spread can distribute replicas across zones or hosts; strict placement rules can also leave Pods Pending.

[S04] [S09] [S13]

Worked baseline synthesis

For an illustrative quote API, start with three replicas, one surge Pod and zero unavailable Pods during a rollout. Use a PDB requiring two available replicas and spread placement across eligible zones and hosts. These are exercise settings, not a claim that three replicas meet every SLA. Reserve surge capacity and test the load remaining replicas can handle.

[S04] [S09] [S13]

Pitfall and shutdown decision synthesis

Do not make every transient database failure kill a healthy process. Design local liveness and useful-service readiness deliberately. During termination, the application must stop accepting new work and finish or release existing work within its grace period; align this with load-balancer draining. This is application design guidance, not a promise of immediate endpoint convergence.

[S10] [S20]

Further improvement synthesis

Use the bundled workload example as a review specimen. Tune startup budget, readiness conditions, request duration and drain behavior using a load test. Verify failed releases and replica loss at the user-facing endpoint. A rollout remaining available says little about a backwards-incompatible database migration.

[S10] [S04]
Keep this: Readiness controls participation; liveness restarts; placement and rollout controls govern other failure paths.
Check yourself: Can a PDB guarantee that a Deployment rollout never takes down too many Pods?

No. Deployment rollout availability is controlled by its rollout strategy. PDB governs eligible eviction paths and does not prevent involuntary failures.

Sources & further reading

  1. [S04] Deployments

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: ReplicaSet ownership, rolling updates, surge and unavailable settings

    Read the linked primary source for implementation details and current constraints.

  2. [S09] Disruptions

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: PDB limits eviction; involuntary failure and rollout limitations

    Read the linked primary source for implementation details and current constraints.

  3. [S10] Liveness, readiness and startup probes

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Probe actions and startup protection

    Read the linked primary source for implementation details and current constraints.

  4. [S13] Topology spread constraints

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Zone and host distribution, strict scheduling constraints

    Read the linked primary source for implementation details and current constraints.

  5. [S20] AWS Load Balancer Controller

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Ingress to ALB and Service LoadBalancer to NLB integration

    Read the linked primary source for implementation details and current constraints.