Sai Joshitha Kathari stands as a prominent figure in the evolving landscape of Site Reliability Engineering, bringing a wealth of experience in managing high-stakes distributed systems and Kubernetes environments. As a Senior SRE, her work focuses on the delicate balance between high-speed innovation and the rigorous predictability required for global-scale infrastructure. Her technical journey is defined by a transition from the structured world of development to the chaotic, high-pressure reality of production operations, where she has pioneered automation strategies that significantly reduce manual overhead. Beyond her technical contributions to OpenShift and cloud-native architectures, she is a vocal advocate for technical ownership and visibility for women in infrastructure roles. This conversation explores the intricacies of reliable system design, the impact of AI on observability, and the systematic mindset required to navigate complex architectural failures.
The shift from a controlled development environment to the unpredictable nature of production often catches many engineers off guard; what specific complexities did you encounter that redefined your understanding of system reliability?
When I first transitioned into production environments, it felt like moving from a laboratory into a storm. In development, you have the luxury of isolated variables, but production introduces a chaotic symphony of traffic spikes, resource limits, and intricate networking dependencies that are almost impossible to simulate perfectly. I remember the first time I saw an application fail not because of a bug in the code, but because of a subtle downstream service delay that triggered a cascading failure across the cluster. Working with Kubernetes and OpenShift opened my eyes to how resource pressure and autoscaling behavior can create unexpected interactions that look like application errors but are actually infrastructure signals screaming for attention. It was these moments, where the smell of metaphorical smoke came from the networking layer rather than the logic, that pulled me toward the SRE discipline where you have to be part detective and part architect.
You’ve successfully implemented automation that slashed Kubernetes post-deployment validation from 45 minutes to just two minutes; what was the technical architecture behind this shift and how did it change the daily rhythm of your team?
That project was born out of pure necessity because watching an engineer spend 45 minutes manually clicking through health checks and verifying pod states was a massive drain on our collective morale. We built a validation framework that hooked into our CI/CD pipeline, automatically aggregating signals from application logs, infrastructure metrics, and cluster events to confirm a successful rollout. I can still recall the first time we ran the full automated suite; the silence in the room as the process finished in 120 seconds instead of the usual three-quarters of an hour was incredibly rewarding. Beyond the numbers, it eliminated the cognitive load of having to remember a sequence of dozens of manual steps, which meant we could focus on actual engineering rather than repetitive checklists. It turned a high-friction task into a background process, giving the team the confidence to deploy more frequently without the fear of missing a critical validation step.
As we move further into 2026, the volume of logs, metrics, and traces has become overwhelming; how do you see AI-assisted SRE workflows helping engineers find the needle in the haystack without losing the vital element of human judgment?
The sheer volume of operational data we deal with today is staggering, and during a major incident, the challenge isn’t a lack of information, but the overwhelming noise of a thousand alerts firing at once. I see AI as a powerful tool for pattern recognition—it can summarize an incident timeline or group related alerts in seconds, tasks that would take a human ten minutes of frantic scrolling. However, I am very cautious about the “black box” nature of AI in production; if a model suggests a remediation step based on a hallucination or an incomplete understanding of our specific context, the fallout could be disastrous. The engineer must remain the final arbiter, using AI-generated summaries as a starting point rather than a final conclusion. My focus is on creating a feedback loop where the AI highlights unusual patterns in the cluster events, but a human engineer validates the logic before any rollback or scaling action is triggered.
There is a distinct tension between the speed of automation and the risk of automated mistakes; how do you design your reliability frameworks to ensure that a small logic error doesn’t escalate into a site-wide outage at machine speed?
The danger of automation is that it can apply a mistake with 100% efficiency and zero hesitation, which is why I advocate for a “trust but verify” approach to every script we write. We build in multiple layers of validation and safety checks, ensuring that any automated action has a clear path for rollback and is gated by observability signals. I have seen instances where a poorly configured autoscaler could have wiped out a cluster if we didn’t have hard limits and human-in-the-loop triggers in place. It’s about building guardrails that are just as robust as the automation itself, ensuring that if a process detects an anomaly it doesn’t recognize, it defaults to a safe state or alerts a human rather than guessing. We prioritize visibility so that at any given microsecond, an engineer can look at a dashboard and see exactly why an automated decision was made.
Breaking the glass ceiling in infrastructure and SRE roles often requires more than just technical skill; how did you personally move from a place of assuming others knew more to taking full ownership of high-stakes technical decisions?
In the early stages of my career, I often sat in rooms full of infrastructure veterans and felt like I was the only one who didn’t have the entire distributed system mapped out in my head. It took several high-pressure incidents for me to realize that in a system as complex as Kubernetes, nobody actually has the full picture—we are all just testing assumptions and following signals. That realization was liberating; it allowed me to stop worrying about what I didn’t know and start focusing on how to find the answer systematically. I began taking ownership of problems even when the solution wasn’t clear, trusting my ability to troubleshoot and ask the right questions rather than waiting for a “more senior” person to step in. Writing and speaking about these challenges also helped reinforce my confidence, as I realized that the problems I was solving were the same ones keeping other experts up at night.
Troubleshooting a major production incident is often a high-stress “trial by fire” moment; can you describe your systematic approach to staying calm and narrowing down a problem when the stakes are at their highest?
The first time I handled a major outage, my heart was racing and I wanted to fix it so badly that I almost jumped to the wrong conclusion based on a single error log. Now, I’ve developed a mental checklist that starts with a deep breath and a look at what changed in the last 15 minutes: Was there a deployment, a config change, or a spike in external traffic? I systematically move through the layers, checking application health, then pod behavior, and then diving into node pressure or cluster events to see if the infrastructure is choking. By following this structured path, you move away from guessing and toward a process of elimination that eventually reveals the root cause. It’s about slowing down just enough to be fast, ensuring that every action taken on the production cluster is deliberate and backed by a specific metric or signal.
In your experience, what is the most effective way to close the gender gap in specialized technical fields like SRE and systems engineering beyond just general entry-level recruitment?
Closing the gap requires a fundamental shift in how we assign technical ownership; it’s not enough to have women in the room—they need to be the ones designing the systems and leading the incident response calls. When younger engineers see women handling the “hard” infrastructure tasks, like debugging a kernel panic or rearchitecting a global load balancing strategy, it changes their perception of what is possible. Representation is most powerful when it is visible at the highest technical levels, showing that these roles are not just accessible, but are places where women can thrive and lead. We need to encourage more women to speak at technical conferences and participate in deep-level peer reviews, as this builds a culture where technical merit is the primary currency. By providing real opportunities for high-stakes ownership, we create a pipeline of senior leaders who serve as living proof that infrastructure is a field for everyone.
What is your forecast for the evolution of SRE roles as we integrate more autonomous infrastructure management over the next few years?
I believe we are moving toward a reality where the “Reliability” part of SRE will become even more focused on the governance of autonomous systems rather than manual intervention. In the coming years, I expect to see our clusters handling routine self-healing tasks—like rescheduling failed pods or adjusting resource quotas—with almost zero human input, based on years of learned patterns. This doesn’t make the SRE obsolete; instead, it shifts our focus toward higher-level architectural integrity and the ethics of automation. We will spend more of our time designing the “meta-logic” that governs how these autonomous systems interact with one another, ensuring that our infrastructure remains predictable even as it becomes more complex. The role will evolve into a blend of systems architect and data scientist, where understanding the relationship between disparate data signals is the most valuable skill an engineer can possess.
