Red Hat tests Argo CD Agent across 480 OpenShift clusters
A pull-model GitOps lab run synchronized 3.12 million resources in about 13 minutes, while exposing bottlenecks in the hub API, informers and retry queues.
Red Hat’s latest scale test gives platform teams a concrete look at how far the pull-based Argo CD Agent architecture can stretch—and where it starts to strain. In a lab run described by Red Hat engineers, the system synchronized 3.12 million Kubernetes resources across 480 OpenShift clusters in about 13 minutes (Red Hat Developer).
The result is useful evidence for teams evaluating the agent model, but it is not a blanket capacity promise. The test used a purpose-built environment, a fixed application shape and tuning that included throttling API queries at the largest scale (Red Hat Developer).
What Red Hat tested
The environment combined OpenShift 4.21, Advanced Cluster Management 2.16 and OpenShift GitOps 1.20 with Argo CD 3.3.z. Red Hat says it provisioned one three-node hub and 500 virtualized spoke clusters on 42 physical machines, then drove 65 ApplicationSets that generated one application per target cluster, with 100 resources per application (Red Hat Developer).
At the top measured phase, 480 clusters hosted 31,200 applications and 3.12 million resources. Initial bootstrap completed in 13 minutes and 3 seconds; a later Git commit reached the synchronized state in 13 minutes. Cascading deletion of all 31,200 applications took 4 minutes and 53 seconds (Red Hat Developer).
Why the pull model matters
Traditional centralized Argo CD keeps credentials for remote clusters and reaches their Kubernetes APIs from the hub. In the agent architecture, each workload cluster runs a local Argo CD instance and agent, initiates communication with the control plane, and performs reconciliation locally. Red Hat’s documentation says this reduces cross-cluster credentials and API exposure while retaining a central view of applications (OpenShift GitOps 1.21 documentation).
Red Hat reported up to an 80% reduction in cross-cluster network traffic in its testing. The agent design also keeps existing applications under local management during a control-plane communication interruption, with event streams resuming when connectivity returns (Red Hat Developer; Red Hat blog).
The bottlenecks are the important part
The test did not eliminate central pressure. At the largest phase, Kubernetes API and etcd throughput on the hub became limiting, prompting the team to cap the principal’s API query rate. Red Hat also observed slow informer callbacks, retry events that clogged queues when acknowledgements took longer than one second, and high memory use in the ApplicationSet controller (Red Hat Developer).
For operators, that means the architectural shift distributes reconciliation but does not make the hub irrelevant. Capacity planning still has to account for ApplicationSet generation, control-plane API load and event processing. Red Hat characterizes Argo CD Agent as an emerging configuration with unsupported features, and its current documentation targets advanced users already familiar with OpenShift GitOps (OpenShift GitOps 1.21 documentation).
There is also a licensing prerequisite: the agent requires an OpenShift Platform Plus subscription on every cluster that runs it, although the GitOps control plane remains available with OpenShift Container Platform (Red Hat Developer).
sources
- Scaling GitOps: The "pull" architecture for global fleetsdevelopers.redhat.com
- Argo CD Agent architecture — OpenShift GitOps 1.21docs.redhat.com
- Manage clusters and applications at scale with Argo CD Agentwww.redhat.com
comments · 0