SLOs for Kubernetes Services: Alerting on What Actually Matters
Stop paging on CPU and start paging on burn rate. A practical guide to SLOs, error budgets, and multi-window alerts for Kubernetes services.
Read the post4 posts about sre in the conndeck blog — field notes on local-first Kubernetes operations, GitOps, and production debugging.
Stop paging on CPU and start paging on burn rate. A practical guide to SLOs, error budgets, and multi-window alerts for Kubernetes services.
Read the postKubernetes upgrades don't have to be scary. Version skew rules, drain and PDB discipline, deprecation hunting, and how managed vs self-managed changes the job.
Read the postFrom 2 clusters to 20: kubeconfig hygiene, fleet-wide visibility, and GitOps per cluster. The patterns that keep multi-cluster from becoming multi-chaos.
Read the postA deep dive into kube-scheduler internals: filtering vs scoring, topology spread math, PriorityClass preemption, scheduler profiles, and why the descheduler exists.
Read the post