← Roi Weinberg

Auto Remediation (Self Healing)

Auto-remediation was a feature we'd wanted to build for a long time — the idea was that Komodor shouldn't just tell you something is wrong, it should be able to fix it. This was an early experiment to test that idea in practice: AI identified what was wrong, and a policy decided whether and how to act — no chat, no back-and-forth. It never went GA, but it shaped a lot of how we thought about trust and control in everything that came after.

A self-healing event drawer showing a step-by-step audit timeline
A dedicated "self-healing" event drawer that explains what happened, what changed, and the final outcome — a full, step-by-step audit timeline.

Action needs an audit trail, or nobody will trust it

The core design problem wasn't the automation itself, it was that automated changes to production are scary if you can't see exactly what happened. So every self-healing action got its own event drawer, structured as a timeline: issue detected → policy triggered → root cause identified → action applied → outcome observed. Nothing happened silently. If Komodor touched a deployment, there was an exact record of what changed and why, sitting right next to the resource it changed.

Guardrails before autonomy

Detection could tell you what was wrong, but it couldn't tell you what was safe to do about it — that judgment had to come from humans. Policies scoped exactly what was allowed — by cluster, by action type, by resource kind. A "Production" policy could be locked down to reruns and restarts only, while "Non-Prod" could allow broader edits. This wasn't a detail bolted on after the fact; it was the thing that made the feature safe to ship at all.

A policy governing what actions are and aren't allowed to run automatically
Policies would "govern" what can and cannot be done automatically.

Making the blast radius explicit at creation time

The policy builder pushed the same idea one step further: instead of a vague "allow automated fixes" toggle, it broke permissions into named actions (edit workload resources, in-place rollback, revert Helm release, rerun jobs), each scoped to specific resource kinds, with inline warnings when one permission implied another (e.g. editing a StatefulSet also permits a restart). The goal was that nobody building a policy could be surprised later by what it let the system do.

The policy builder, showing named actions scoped to specific resource kinds and environments
The policy configuration — explicitly stating what actions are allowed to perform automatically, and on which environments (clusters).

Why this holds up as a pattern

The feature itself didn't survive — detection alone wasn't a strong enough foundation to justify the trust automated action was asking for. But the pattern did: automation earns trust through visibility (a clear audit trail) and constraint (explicit, human-authored boundaries), not through confidence in the system's own judgment. That gap between detecting a problem and safely acting on it is exactly what a reasoning layer like Klaudia was later built to close.