Komodor's platform scans clusters for a set of out-of-the-box reliability issues: outdated Kubernetes versions, high container restart counts, CPU throttling, and more. Each violation type surfaces very different kinds of data, so a single generic "alert card" wasn't going to cut it.
The first version of this was a flat list of violations. It worked, but it didn't answer the question users actually had first: what is this actually affecting? We restructured the top-level view around impact groups — Node Pressure, Degraded Service, Cluster Upgrades — each summarized by the outcomes it's causing (e.g. "12 noisy neighbors, 456 victims, 7 under-provisioned workloads") before a user ever drills into an individual rule.
Every violation detail view followed the same three-part structure: what happened → why is it important → how to fix this. That consistency means users always know where to look. But the "what happened" section changes shape depending on the relevant data:
Some violations included inline actions — reconfigure resources, edit HPA — letting users act directly from the violation rather than context-switching elsewhere. These leaned on capabilities the platform had already built, so the design layered a fix path onto existing infrastructure rather than inventing new backend work.
The interesting design decision here wasn't any single screen — it was recognizing that a template needs to flex its middle section to match what the data actually needs to say, while keeping the framing (what / why / how) constant so users never have to relearn the page.