Kubernetes runtime: agent pods run at priority 0 and preemption is reported as a plain stop

#2528 · open · 0 comments

View on GitHub ↗

ptone

## Summary On GKE Autopilot, agent pods are created with no priority class (priority 0). During cluster rescaling, the scheduler can preempt an agent pod to place a system-critical pod (observed: kube-dns, priority class system-cluster-critical). The agent then shows as plain `stopped`; no preemption reason is surfaced and the pod is not rescheduled. Observed on a test cluster: 1 of the first round of starts, and 2 of 14 starts in a later round. ## Expected - Agent pods are not the first choice for preemption, or the operator can choose: for example a configurable `priorityClassName` on the Kubernetes runtime (and a documented recommended class with a modest priority and `preemptionPolicy: Never`). - When an agent pod is preempted or evicted, the agent status says so (for example an exit reason of `preempted` or `evicted` taken from the pod status reason or the DisruptionTarget condition), rather than a plain stop. ## Notes - Related: agent pods that disappear are now detected by the hub reconcile work in ptone/scion#2492. - Part of ptone/scion#1956.

Comments