Red Hat’s Krkn adds resiliency scoring and deeper chaos telemetry
The chaos-engineering project now measures weighted SLO failures, endpoint downtime and KubeVirt VM connectivity instead of relying on a binary pass/fail result.
Red Hat has detailed a broader telemetry model for Krkn chaos engineering, combining a beta Resiliency Score with continuous HTTP health checks and KubeVirt virtual-machine checks. The change is meant to show how severely a system degrades during a fault experiment, not merely whether it remains nominally available.
A weighted score for the chaos window
Krkn’s beta Resiliency Score ranges from 0% to 100%. It evaluates Prometheus data collected during the chaos window against configured service-level objectives. A triggered alert marks the associated SLO as failed, and a weighted model determines the overall result.
The default weighting assigns one point to warnings and three points to critical outages, while teams can set custom weights and add SLOs. That gives operators a way to make failures in business-critical services count more heavily than lower-impact alerts. Krkn emits both an overall score and per-scenario reports showing passed and failed SLOs and points lost.
Endpoint and VM checks expose the duration
Continuous HTTP health checks probe configured application endpoints at user-defined intervals, with support for bearer tokens and credential tuples. Their telemetry records status codes plus the start, end and duration of downtime. Teams can therefore measure recovery time instead of stopping at a failed-health-check flag.
For virtualized workloads, Krkn’s KubeVirt checks monitor virtual-machine instances throughout the run. They test SSH connectivity through virtctl, or through worker nodes in disconnected environments, and record loss of connectivity, node placement, IP changes after rescheduling and downtime duration. A final post-experiment check flags VMs that remain unreachable after fault injection ends.
A CI quality gate
Red Hat positions the three layers—weighted SLO scoring, endpoint availability and VM connectivity—as a versioned quality gate. Teams can run the same scenario against successive application releases and compare a single resiliency score with the underlying outage evidence.
That approach should make regressions easier to distinguish from harmless alert noise. It also broadens the audience for chaos results: platform teams get Prometheus and infrastructure detail, while application teams get endpoint and recovery-time measurements tied to a release.
sources
- Beyond pass or fail: The new era of chaos engineeringdevelopers.redhat.com
comments · 0