Part 12 · Best practices

Observe and diagnose the right layer

Prerequisites: 03-networking, 09-release-reliability, 10-scaling

12 / From a symptom to the next useful check

Four diagnostic branches map Pending, CrashLoopBackOff, errors and latency to checks and possible causes.

Scroll the diagram sideways for readable labels.

Reasoned troubleshooting flow based on component responsibilities and observability primitives. A symptom can have more than one root cause. [S27] [S07] [S05] [S19]

Signals serve different consumers

One pipeline collects app and node telemetry for dashboards. A separate resource Metrics API provides CPU and memory for HPA and kubectl top.

Scroll the diagram sideways for readable labels.

Logical pipelines; collection components depend on compute mode. EKS does not expose every managed control-plane metric endpoint as a self-managed cluster would. [S27] [S28]

Learning objective and signals verified

Metrics reveal trends and saturation, logs describe local events and traces connect request steps. The resource Metrics API used by HPA and kubectl top is a separate capability from a full monitoring pipeline. Collecting CPU usage does not provide historical logs or distributed traces.

[S27]

EKS control-plane view verified

EKS can export API, audit, authenticator, scheduler and controller-manager logs to CloudWatch. Export is disabled by default and each type is enabled separately. Audit records help investigate who changed cluster resources; retention and ingestion costs require planning.

[S28]

Worked incident synthesis

Imagine the quote API becomes slow after adding replicas. Start with latency and error rate, then inspect traces for database waits, connection counts and per-Pod concurrency. Compare CPU throttling, restarts and queue delay. This hypothetical incident tests whether the database bottleneck diagnosis fits; it is not a claim that every latency spike has that cause.

[S27] [S07]

Diagnostic commands: review-only examples synthesis

kubectl -n insurance get pods -o wide
kubectl -n insurance describe pod POD_NAME
kubectl -n insurance logs POD_NAME --previous
kubectl -n insurance get events --sort-by=.metadata.creationTimestamp
kubectl -n insurance get endpointslices -l kubernetes.io/service-name=quotes
kubectl -n insurance top pods

Replace POD_NAME. Previous logs require a prior container instance; top requires a working Metrics API provider.
[S27] [S05]

Decision, pitfall and further improvement synthesis

Choose retention and labels around an actual troubleshooting question. Avoid putting raw tokens or unbounded tenant identifiers into telemetry. Alert on sustained user-facing errors and latency, then include placement, IP capacity and dependencies as diagnostic context. Proposed objective: find a failed release’s cause from retained evidence without needing to reproduce it.

[S27] [S28]
Keep this: Start from user-visible failure, then follow evidence down the request path.
Check yourself: If kubectl top works, are application traces and audit logs automatically available?

No. Resource metrics, application telemetry and control-plane logging are different pipelines that require their own configuration and retention.

Sources & further reading

  1. [S05] Service

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Service selection, stable access and EndpointSlices

    Read the linked primary source for implementation details and current constraints.

  2. [S07] Resource management

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Requests, limits, scheduling, CPU throttling and memory enforcement

    Read the linked primary source for implementation details and current constraints.

  3. [S19] VPC and subnet considerations

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Cluster subnets, private endpoint access, upgrade IP headroom

    Read the linked primary source for implementation details and current constraints.

  4. [S27] Kubernetes observability

    Kubernetes · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Metrics, logs, traces and resource Metrics API distinction

    Read the linked primary source for implementation details and current constraints.

  5. [S28] EKS control plane logs

    AWS · documentation · accessed 2026-10-10 · Living documentation; target-cluster compatibility must be checked · Not stated in retrieved page

    Supports: Optional API, audit, authenticator, scheduler and controller logging

    Read the linked primary source for implementation details and current constraints.