Red Hat confines AI to one stage of CI failure triage—and validates everything it returns
Pipeline Failure Analyzer uses deterministic code for log cleaning, reports and tickets, reserving model reasoning for cross-platform root-cause grouping.
Red Hat’s Pipeline Failure Analyzer offers a useful counterexample to “put an agent on the pipeline.” AI touches one stage of a five-stage workflow; deterministic Python handles log preparation, report assembly, ticket creation and publishing.
The constraint is the architecture. Red Hat’s platform team was routinely facing dozens of failed jobs with traces reaching 50,000 lines, even when the failures collapsed to a few underlying causes. Sending every log through a model produced nondeterministic reports, linear inference costs and context pressure.
Compress before reasoning
Pipeline Failure Analyzer first removes ANSI codes, CI section markers and progress-bar updates. It then anchors on 22 error patterns and retains bounded context around each match. Red Hat reports that this turns a 10,000-line trace into roughly 200–500 lines, a 95–98% reduction.
That preprocessing is deterministic and independently testable. Structured metadata parsed from job names adds architecture, operating-system and hardware context before the model sees anything.
The AI stage then groups semantically related failures that text similarity misses. A missing dependency may produce a direct package error on x86 but surface as a failed compiler subprocess on Arm. The model connects those symptoms, while a second analysis step investigates a root cause against the exact source revision.
Treat model output as untrusted data
Five validation layers stand between the analysis and downstream automation. Findings must contain required fields and valid enums; group artifacts and references must exist; paths must be canonical; the complete summary must pass schema and cross-field checks; and source traces must be large enough and must not be an HTML error page returned by a broken CI API.
If analysis times out, hallucinates or fails validation, the system creates one catch-all group and still notifies the team. It does not silently discard failures.
Ticket deduplication also maps uncertainty to different actions. High confidence adds a comment to an existing issue, medium confidence creates a new issue with a review note, and low confidence assumes a new problem. Human approval remains mandatory for merging generated fixes.
The reported production result
Red Hat says the system has operated since June 2026, diagnosing 110 pipeline failures. Seventy-eight generated fixes were merged, and 68% of tickets created by the analyzer reached resolution. Those are internal operational results rather than an independent benchmark, but the project’s code, tests and AI skill definitions are open source for inspection.
The transferable lesson is to automate structure before buying reasoning. Teams can start with the published log cleaner and a few known error patterns, measure compression, and only then add a model for cross-platform grouping. Any AI-generated conclusion should enter the rest of the pipeline through a validated schema and a safe fallback path.
This is less autonomous than a free-running CI agent—and much easier to audit.
sources
- How we cut CI failure triage from hours to minutes using AIdevelopers.redhat.com
comments · 0