A workload spread across three Availability Zones (AZs) does not necessarily stay spread. In a measured run on an Amazon Elastic Kubernetes Service (Amazon EKS) cluster, we watched a balanced 1,000-pod fleet concentrate into only two zones after a brief node-availability gap, while the third zone dropped to zero pods. After capacity returned, the imbalance held for over 16 hours with every node healthy and schedulable. Every component behaved as designed.

That is precisely why this pattern deserves attention: the drift is silent, it does not self-correct for stable workloads, and the manifests still describe a distribution that the running state no longer matches. In this post, we explain why drift happens under soft topology spread constraints, measure what it costs, and demonstrate how the Kubernetes descheduler restores the distribution your constraints describe.

The second half is a hands-on walkthrough you can reproduce on Amazon EKS: you scale a workload under randomized load, simulate a disruption, confirm the scheduler does not repair the skew on its own, and then watch the descheduler restore balance. Every number in this post comes from that measured run, and the manifests, dashboards, and scripts to reproduce it are in the accompanying repository sample-eks-descheduler-drift-demo. Kubernetes places pods with the default kube-scheduler, a component of the Kubernetes control plane.

When you give a workload topologySpreadConstraints, the scheduler evaluates them at admission (when the pod is first placed) as part of its scoring algorithm, which ranks candidate nodes by fit. This is the one fact that explains everything else in this post, so we state it once with emphasis: the scheduler evaluates topology spread only at pod creation time. It does not revisit running pods . Once a pod is placed, its position is fixed for the life of that pod. Rebalancing running pods is intentionally not part of core Kubernetes.