Observe and diagnose the right layer
Prerequisites: 03-networking, 09-release-reliability, 10-scaling
12 / From a symptom to the next useful check
Scroll the diagram sideways for readable labels.
Signals serve different consumers
Scroll the diagram sideways for readable labels.
Learning objective and signals verified
Metrics reveal trends and saturation, logs describe local events and traces connect request steps. The resource Metrics API used by HPA and kubectl top is a separate capability from a full monitoring pipeline. Collecting CPU usage does not provide historical logs or distributed traces.
[S27]EKS control-plane view verified
EKS can export API, audit, authenticator, scheduler and controller-manager logs to CloudWatch. Export is disabled by default and each type is enabled separately. Audit records help investigate who changed cluster resources; retention and ingestion costs require planning.
[S28]Worked incident synthesis
Imagine the quote API becomes slow after adding replicas. Start with latency and error rate, then inspect traces for database waits, connection counts and per-Pod concurrency. Compare CPU throttling, restarts and queue delay. This hypothetical incident tests whether the database bottleneck diagnosis fits; it is not a claim that every latency spike has that cause.
[S27] [S07]Diagnostic commands: review-only examples synthesis
kubectl -n insurance get pods -o wide kubectl -n insurance describe pod POD_NAME kubectl -n insurance logs POD_NAME --previous kubectl -n insurance get events --sort-by=.metadata.creationTimestamp kubectl -n insurance get endpointslices -l kubernetes.io/service-name=quotes kubectl -n insurance top pods Replace POD_NAME. Previous logs require a prior container instance; top requires a working Metrics API provider.[S27] [S05]
Decision, pitfall and further improvement synthesis
Choose retention and labels around an actual troubleshooting question. Avoid putting raw tokens or unbounded tenant identifiers into telemetry. Alert on sustained user-facing errors and latency, then include placement, IP capacity and dependencies as diagnostic context. Proposed objective: find a failed release’s cause from retained evidence without needing to reproduce it.
[S27] [S28]Check yourself: If kubectl top works, are application traces and audit logs automatically available?
No. Resource metrics, application telemetry and control-plane logging are different pipelines that require their own configuration and retention.