Remote security scanners can exhaust PIDs across an OpenShift worker
Zombie processes leaked inside container namespaces can push a node to EAGAIN, breaking mounts and leaving workloads in ContainerCreating or CrashLoopBackOff.
Red Hat has verified a node-level failure path in which remote security scanners leave defunct processes inside container namespaces until an OpenShift worker exhausts its process IDs. Once that limit is reached, new forks fail with Resource temporarily unavailable (EAGAIN), according to the Customer Portal solution updated Sept. 18.
The failure path
The documented symptom chain begins with thousands of zombie processes—commonly [sh], [rpm] or [ps]—whose parents are standard platform containers such as kube-rbac-proxy, cephcsi, csi-attacher or fluentd. PID exhaustion is node-wide, so the impact can escape the container that a scanner originally inspected.
Red Hat reports that pods may then fail to attach or mount volumes, with messages saying a kubelet pod path “is not a mountpoint.” Applications can suffer broad downtime, while containers remain in ContainerCreating or CrashLoopBackOff. The listed environment is OpenShift Container Platform 4.12 and later, with ODF and CRI-O also named.
Detection signals
The strongest signal is the combination, not any one error: fork failures with EAGAIN, an unusually large zombie-process count, and workload lifecycle or mount failures on the same worker. A scanner run immediately before the increase is useful correlation, but operators should preserve process ancestry because Red Hat specifically notes zombies parented by core platform containers.
Monitoring should therefore include node PID pressure or process-count trends alongside kubelet and CRI-O errors. During an incident, record affected nodes, zombie counts, parent processes, recent scanner jobs and workload events before remediation changes the process table.
Safer scanner controls
Red Hat’s detailed resolution is subscriber-only, so the public record does not specify a universal concurrency number. The conservative mitigation is to bound scanning rather than choose a guessed threshold: limit simultaneous remote inspections per node, stagger scans across the worker pool, and stop increasing concurrency when process counts fail to return to baseline.
A practical validation plan is:
- Baseline process and zombie counts on representative workers.
- Run the scanner against one canary node at the lowest useful concurrency.
- Observe process recovery after the scan completes.
- Increase scope only if counts return to baseline and workload events remain clean.
- Alert on EAGAIN, mount failures and sudden zombie growth.
- Use Red Hat’s supported remediation if a node is already exhausted.
The key architectural point is that scanner isolation is not complete when helper processes can be orphaned under long-lived platform containers. Treat scanner concurrency as node capacity, not merely scan throughput.
sources
comments · 0