← Roi Weinberg

Designing Klaudia: From Fragmented AI Features to a Persistent AI SRE Assistant

I redesigned Komodor's fragmented AI experience into a single, persistent assistant that "follows" an investigation across an entire cluster instead of being boxed into one resource. Adoption was instant — no onboarding or transition period was needed — and investigations got deeper and more exploratory, validating the thesis that root cause analysis naturally spans more than one resource.

Klaudia, Komodor's AI SRE assistant, shown as a persistent side-by-side panel
Klaudia — Komodor's persistent, side-by-side AI SRE assistant.

Context

Komodor is a Kubernetes troubleshooting, observability, and reliability platform. Klaudia is its AI assistant, built to investigate incidents, explain root causes, and act on fixes across a customer's clusters.

Klaudia's capabilities had grown the way most AI features do inside an existing product: one at a time, each built onto whatever area made sense in the moment. A logs analyzer here. Root cause analysis on issues there, and eventually the ability to "chat" and ask follow-up questions — but boxed into a tight space. As Klaudia grew more capable, she still showed up as a handful of disconnected buttons, each aware only of the single resource it was bolted onto.

The target user is an SRE or developer dealing with an issue — moving between a failed deployment, a related service, a node, and a configuration change, trying to piece together what actually happened.

The problem

Klaudia's AI capabilities were fragmented across surfaces and bound to single resources, so she couldn't follow an investigation as it naturally moved across a cluster. Two things converged into the real motivation for this project: fragmentation — each capability lived on its own surface, with its own entry point and its own memory — and resource-limited context, since real root cause analysis rarely stays inside a single resource.

An early, isolated AI log-analysis feature bolted onto a single resource
The ad-hoc log analysis — an early, isolated AI feature bolted onto a single resource.
A follow-up chat option inside an event drawer
A follow-up chat option inside an event drawer that let users keep discussing a specific issue or failure.

Underneath the UX pain points was a business one: Klaudia was gaining new capabilities (with the rise of MCPs, skills, and similar), but a fragmented, resource-limited presence meant the product's perceived intelligence was lagging behind its actual intelligence, and behind the market. The design challenge became: how might Klaudia follow the investigation instead of being tied to the feature?

Any redesign had to work within Komodor's existing product — Klaudia's role was to augment the Kubernetes platform, not replace it — and it couldn't simply remove the entry points users already knew and relied on.

Outcome


Key stages

Deciding on the solution pattern

We explored three experience models for where Klaudia should live. A floating chat was quickly rejected because it covered the interface the user was trying to investigate. A full-screen assistant was rejected for phase one, since Klaudia's role was to augment the platform, not replace it. We landed on a persistent, side-by-side assistant: Klaudia stays visible alongside the platform while the user moves between resources and continues the same investigation, with a constant "Klaudia AI SRE" button always available to start a session.

Comparison sketch of floating, full-screen, and side-by-side assistant models
Comparison sketch of the three explored experience models: floating, full-screen, and side-by-side.

Design challenges

Choosing side-by-side raised a second problem: once Klaudia could relate to "the resource on stage," users needed to know what the conversation was actually about. We introduced a context chip in the chat input showing the active resource — modeled after how similar context is surfaced in developer IDEs like VS Code — with the ability to pin or disable it while navigating the platform.

A context chip in the chat input showing the active resource
The context chip — modeled after how similar context is shown in developer IDEs like VS Code, with the ability to pin or disable it while navigating the platform.

As Klaudia gained new capabilities — from data collection to making actual changes in a customer's environment — transparency and the "human in the loop" idea became leading values. We surfaced what tools Klaudia was using during a session, so users could see what was actually being done, and, more importantly, approve or reject what she was about to do.

A list of tools Klaudia can invoke, from data collection to environment changes
Some of the "tools" Klaudia can invoke — ranging from basic data collection to performing actual changes in a customer's environment.

Another key constraint was that we couldn't just replace what was already working and familiar — every existing entry point into an AI feature represented a moment users had already learned. Removing them to force a single unified trigger would have traded the fragmentation problem for a new one. Instead, we turned each trigger into a split action: the primary action still starts a focused new chat scoped to that resource, exactly as before, while a second option, "Add to active chat," lets the user fold that same trigger into whatever investigation is already open. No feature was removed — what changed is that each one now feeds the same persistent conversation instead of an ad-hoc monologue.

Validating the approach

Rather than a formal usability testing round, we opened the feature internally once it was stable enough for real use — Komodor's own developers used it against real clusters during their day-to-day work. The response was positive: the side-by-side model and split-action pattern held up under actual investigation use, which gave us the confidence to move toward a broader rollout.

The implemented Klaudia design inside the logs section
The implemented design in the logs section.

The solution & polish

The shipped experience is a persistent, side-by-side AI SRE assistant: always reachable via a constant "Klaudia AI SRE" button, aware of the active resource via the context chip, transparent about what tools it's invoking, and fed by every existing entry point through the split-action pattern rather than a single forced trigger.

The final Klaudia design, a persistent side-by-side assistant for Kubernetes troubleshooting
Final design — a simple, persistent, side-by-side assistant to help developers with their Kubernetes troubleshooting.
Different conversation states: initial suggestions, root cause analysis, and waiting on a user decision
Different conversation states: initial suggestions, root cause analysis in progress, and waiting on a user decision about how to proceed.
File upload support with a guardrail against uploading sensitive data
Supporting file uploads, with a guardrail to make sure no sensitive data gets uploaded.

Impact

The rollout was gradual rather than a single switch-flip, and the response was immediate:

Together, these shifted Komodor's AI capabilities from a collection of isolated features into an assistant that could support an investigation from diagnosis through remediation, built around three principles: continuity (investigations span resources without restarting), visible context (users see what Klaudia is reasoning about), and explicit control (users approve or reject every proposed action).

Closing the loop

The new design solved where Klaudia lived. It didn't solve whether what she said was any good — and now that users were approving real actions she proposed, that question mattered more than ever.

The solution was a 1–6 rating sitting directly under the investigation itself — no modal, no separate flow people would just skip. A score of 4 or below triggers a quick follow-up: a reason chip (wrong root cause, missing detail, not actionable, among others) plus an optional free-text field, so a bad score comes with why.

The inline feedback rating widget below a Klaudia investigation
The feedback widget was added directly above the user input.
The short follow-up feedback form with selectable reasons and a free-text field
The feedback form was short — letting users select a reason for their discontent and write any additional message directly inside the product.

Individually, these ratings are just noise. Rolled up into a dashboard — average score, satisfaction split, and a ranked "top 5 reasons for low scores" — they become a trackable signal the team can actually prioritize against. It's the other half of the same bet the side-by-side model made: continuity and visible context aren't enough if there's no way to know when Klaudia's reasoning fell short.

A feedback dashboard showing average score, satisfaction split, and top reasons for low scores
The feedback dashboard, rolling individual ratings up into a trackable pattern.