ozzycore
### Summary Since 0.22.0, Keel stops applying updates on clusters with a few hundred tracked workloads. The log is flooded with client-go throttling waits of 45–60s and no `resource updated` events are ever produced. Rolling back to 0.21.1 resolves it. ### Environment - Keel 0.22.3 (`ghcr.io/keel-hq/keel:0.22.3`), chart 1.2.2, Helm3 provider enabled - EKS, in-cluster config, arm64 node - ~210 tracked workloads (≈50% Deployments, ≈50% CronJobs) across 4 image repositories - Every workload annotated with `keel.sh/policy: glob:<tag>`, `keel.sh/trigger: poll`, `keel.sh/pollSchedule: "@every 30s"` ### Observed Over 3.5h after a restart: 1,241 of 1,395 log lines are throttling messages, 0 resources updated. The initial scan took ~2.5 minutes to register just 2 `WatchRepositoryTagsJob`s. `event buffer saturated; applying backpressure` is logged repeatedly. ``` I0929 12:31:15.356389 1 request.go:700] Waited for 46.795361769s due to client-side throttling, not priority and fairness, request: GET:https://<apiserver>/api/v1/namespaces/<ns>/pods?labelSelector=name%3D<workload>%2C... I0929 12:31:55.756264 1 request.go:700] Waited for 52.183120676s due to client-side throttling, not priority and fairness, request: GET:https://<apiserver>/api/v1/namespaces/<ns>/pods?labelSelector=name%3D<workload>%2C... ``` ### Root cause (as far as I can tell) 0.22.0 added two per-workload lookups inside `kubernetes.Provider.TrackedImages()` (`provider/kubernetes/kubernetes.go`, the `p.platforms.Resolve(gr)` / `p.runningDigests.Resolve(gr)` calls): - `k8s.PlatformResolver.Resolve()` → `Pods(namespace, selector)` - `k8s.RunningDigestResolver.Resolve()` → `Pods(namespace, selector)` again Both are live `CoreV1().Pods().List()` calls (`provider/kubernetes/implementer.go`), not informer/lister reads, and `TrackedImages()` iterates **every** tracked workload. It is called: 1. by the poll manager every `POLL_SCAN_INTERVAL` (default 1m), and 2. by **every** `WatchRepositoryTagsJob.Run()` via `computeEvents()` (`trigger/poll/multi_tags_watcher.go`), which resolves all workloads and only afterwards filters them with `getRelatedTrackedImages()`. Meanwhile the clientset is built from `rest.InClusterConfig()` with client-go defaults (QPS 5, burst 10), and there's no way to override them. With N=210 workloads that is 420 LISTs per `TrackedImages()` call, so a single call takes ≥84s at 5 QPS. With 4 repository jobs on `@every 30s` plus the 1m scan, demand is ~35–65 req/s against a 5 req/s budget, so calls pile up indefinitely behind the shared rate limiter. The cost is O(repositories × workloads), so it gets worse as clusters grow. Even the 1m manager scan alone (420/60 ≈ 7 req/s) exceeds the budget at this size. 0.21.1 has neither `internal/k8s/platform.go` nor `internal/k8s/running_digests.go` and works fine on the same cluster. ### Why it can't be worked around with configuration - No env/flag for the Kubernetes client QPS/burst - No toggle to disable platform resolution / running-digest resolution - Lowering poll frequency enough to fit in 5 QPS means ~10m for both `POLL_SCAN_INTERVAL` and `pollSchedule`, and it breaks again as workloads are added ### Suggested fixes (any of these would help; 1 + 2 would fully solve) 1. Serve pod lookups from a shared informer/lister (Keel already watches deployments/statefulsets/daemonsets/cronjobs this way) instead of live LIST calls. 2. In `computeEvents()`, filter to related tracked images **before** r digests, or resolve lazily only for workloads that share the job'simage. 3. Compute platforms/running digests once per scan and cache them, ratImages()` call. 4. Expose the client rate limits, e.g. `KUBE_API_QPS` / `KUBE_API_BURST` env vars applied to the `rest.Config` in `NewKubernetesImplementer`. ### Repro Create ~200 Deployments/CronJobs sharing 2–4 image repositories, annotlicy and `@every 30s` poll schedule, run Keel 0.22.x, push a new digestfor one of the tags, and observe the throttling log lines and that no update is applied. Same setup on 0.21.1 updates normally.