Node NotReady at 2AM: A Recovery Walkthrough
A node went NotReady at 2AM. Triage kubelet, disk pressure, and network partitions, then decide when to cordon, drain, or wait.
Read the post1 post about incident-response in the conndeck blog — field notes on local-first Kubernetes operations, GitOps, and production debugging.
A node went NotReady at 2AM. Triage kubelet, disk pressure, and network partitions, then decide when to cordon, drain, or wait.
Read the post