Who: Anish Agarwal is the Co-founder and CEO of Traversal. A former Columbia University faculty member and MIT researcher specializing in causal machine learning and reinforcement learning, he left academia to build autonomous infrastructure tools for the AI software era.
What: Traversal builds an AI site reliability engineer that automatically troubleshoots, diagnoses, and fixes complex software system outages. By combining causal machine learning with advanced reasoning models, the platform scans logs, metrics, code, and traces to pinpoint root causes, giving engineers verifiable evidence links directly into their observability systems.
Lesson: Emerged from stealth with $48 million in Seed and Series A funding led by Sequoia Capital and Kleiner Perkins. During a six-month deployment with public cloud provider DigitalOcean, serving over 600,000 infrastructure users, Traversal cut the time to resolution for system incidents by 37% while achieving 90% diagnostic accuracy.
Anish Agarwal left theoretical research to found Traversal after recognizing that the explosion of AI-generated code would make complex software systems nearly impossible for human teams to maintain. Backed by $48 million in funding led by Sequoia Capital and Kleiner Perkins, Traversal is building an AI site reliability engineer capable of autonomously diagnosing and fixing large-scale infrastructure failures. By grounding the platform in causal machine learning, Agarwal re-architected Traversal's production models to achieve 90% diagnostic accuracy and cut incident resolution times by 37% for cloud provider DigitalOcean. In this interview, Agarwal reveals how he transitioned from academic research to rapid startup execution, why enterprise AI systems require first-principles reasoning over simple human replication, and how automating software maintenance will preserve human engineering creativity.
Key Takeaways
AI Agents Need Evidence From Observability Systems
Traversal's AI site reliability engineer troubleshoots complex software failures and links its answers to observability data rather than generic web pages. Working with DigitalOcean, the team reduced time to resolution by 37 percent, showing how evidence can make automated incident response more useful.
Research and Entrepreneurship Compress Uncertainty Into Months
Agarwal connects research and entrepreneurship through the challenge of creating structure from uncertainty, but the feedback cycle is much faster in a company. Research may allow years to make an impact, while entrepreneurship can require learning whether an idea works within a month.
Choose Problems at the Intersection of Your Edge
Traversal started without a fixed idea but with a clear preference for problems combining causal machine learning, reinforcement learning, and AI agents. Incident response fit that edge, offered a large market, and carried high technical risk without requiring the founders to know the industry already.
MVP Accuracy Does Not Prove Production Readiness
Traversal's first MVP reached 90 percent accuracy with small companies, then fell to zero percent at larger enterprises such as DigitalOcean. Agarwal uses the reversal to distinguish a promising prototype from a production system that must handle complex, changing environments.
Use Computation to Recover From Zero Accuracy
After accuracy collapsed, Traversal rearchitected its system around what AI models do well, which is inference and computation rather than simply encoding human creativity. Agarwal compares incident response to a detective story, where many evidence points must be connected to find one answer.
Build With People Who Sustain Deep Conviction
Agarwal says difficult companies require a deep desire to solve the problem, quick recovery from intense stress, and trusted people around you. The right advisers, investors, customers, and teammates create guidance through uncertainty, while the work itself must matter enough to sustain commitment.
Watch the full interview now on EO's YouTube channel! Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.
Introducing Anish Agarwal, CEO of Traversal
My name is Anish Agarwal. I'm the CEO and co-founder of Traversal. We've just announced our Series A and come out of stealth with $48 million in seed and Series A funding, led by Sequoia and Kleiner Perkins.
We're building an AI site reliability engineer. When you have large, complex software systems and they break, we troubleshoot to figure out why, and then help fix it automatically. It's a bit like Perplexity. When you ask Perplexity a question, it gives you evidence, citations for how it got to the answer. In our world, the citations aren't web links. The citations are links into your observability system.
One customer we can now talk about publicly, which we've worked with very closely for the last six months, is DigitalOcean. They're a large public cloud provider, I think the third largest by number of people using them, with over 600,000 people running their core infrastructure on it. You can imagine that when they break, every one of those people feels the pain, because that's the main thing powering their systems. In those six months, we've dropped their time to resolution by 37%, which is incredible, because every minute of downtime is worth thousands, if not millions, of dollars.
The ChatGPT Moment Changed Everything
I like sports and I like competing, and I always gravitated naturally to math and science. I liked how abstract and clean they were, and they felt universal in what you could do. I came to MIT because I thought machine learning and AI were really important and I wanted to understand them deeply. This was 2016, eight or nine years ago. One of the biggest moments of my research career was NeurIPS 2017, where the keynote was given by Google's AlphaGo team. I found it incredible that a system could learn something creative by itself. The complaint you always heard before was that it's just copying people, but this was learning creativity through self-play. I thought we should apply that architecture everywhere. I was fortunate to join Columbia's faculty. I like thinking about theoretical problems and mathematical abstractions, and a university is an amazing place to do that.
The thing that changed was ChatGPT. It felt like something incredible had happened in the world, a once-in-a-lifetime thing where the world had fundamentally changed and people didn't even realize it. It felt almost like a religious experience.
I really like uncertainty. I like creating something from zero to one. What research and entrepreneurship have in common is that you have no idea what's happening most of the time, and you have to find ways of creating structure from nothing. The difference is that here the time spans are compressed. In research, you get five or ten years to make an impact. Here you get one month. The feedback cycle is very quick, and in this age, when AI is changing so quickly, being in that quick feedback cycle is very important.
We've entered the industrial age of artificial intelligence. I saw some of the smartest people around me either at OpenAI, Anthropic, or Meta, or creating companies, and that's what made me excited. Starting a company felt like a great expression of that.
Begin with Your Edge
The best AI companies are always going to be at the edge of where the models are. That's how you differentiate yourself. At the edge, sometimes it works and sometimes it doesn't, and you need to know quickly which it is and correct for it.
When we started the company in January 2024, we started without an idea, but with a clear taste for the type of problem we wanted to take on. We wanted to work at the intersection of our research, causal machine learning and reinforcement learning, and AI agents. Causal machine learning is the study of cause and effect. How do you get AI systems to pick up cause-and-effect relationships from data? An A/B test is one way of learning cause and effect, and a clinical trial is another. We were trying to see how that intersected with AI agents, something we had followed for almost two and a half years. We went through a few different ideas.
Then our fourth co-founder, Ammon, joined us and pitched the problem of dealing with incidents. As we looked into it, it felt like a perfect problem, finding a needle in a haystack with many fake needles everywhere. It fit our research in causal machine learning and reinforcement learning. It fit LLMs, because the haystack is made of logs, metrics, traces, code, and configuration files. It fit AI agents, because you have to automate a complex workflow, querying different pieces of the software, reasoning over them, and writing more queries, in a sequential, adaptive flow. And it's a big market, because everyone cares about software not going down, and it's only getting bigger, because so much more code is being written by tools like Cursor, Windsurf, or Copilot. It's a Cambrian explosion of code, and no one really understands it. When that software breaks, it's going to be really difficult to troubleshoot. That's what gave us confidence that this was the problem to take on.
Honestly, the first VC I ever met was Sequoia, and they had been looking for a team to solve this problem. They had reached the same thesis, which is that the AI risk and technical risk are really high, but the market risk is low, because if you can solve it, there's a big market. People like us, who don't come from this world but are strong on the AI side, are the right people to solve it. That validation also gave us confidence.
The First Principle Saved Us From 0% Accuracy
A lot of what AI agent companies do is try to replicate what humans have done, but there is so much more that can be done.Thinking from first principles about what AI systems are good at, and exploiting that rather than replicating a human, is going to be very important for reaching the next level of innovation.
Creating an MVP is so easy with all the tools out there, and you can iterate very quickly and put a product in people's hands. We built our first MVP about three months in, in June or July of last year. With small companies it worked great, because the scale of the data was small, we could look at their historical incidents, see what the playbook was, and put that into an AI agent. Our accuracy was around 90%, which is amazing, so we felt confident it was going to work. But just because it works once doesn't mean it always will, because the world is constantly changing.
Then we hit some of the larger enterprises, including DigitalOcean, and our accuracy went to 0%. That was a tough week. Creating an MVP is easy, but creating a production system that works in complex environments is really hard. Don't confuse an MVP with a production AI system. Those are two very, very different things.
So we re-architected a lot. We asked ourselves, how do we stop trying to get our own creativity into an agent, and instead use what these AI systems are good at, which is computation, inference? That's what unlocked us, and suddenly our accuracy went back up to 90%.
How do you make sure your system gets better as the reasoning models get better? We saw a lot of competing companies that didn't improve at all when the reasoning models came out. So how do you exploit what reasoning models are good at? The way I put it is that reasoning models are very good at detective stories. It's a mystery novel. You're trying to figure out who committed the crime, and there are all these pieces of evidence, and you connect the dots to find the person who did it. That workflow, with a clear answer at the end and lots of moving pieces to connect, is what these models are very good at. It felt perfect for us, because we have all these symptoms happening at the same time, and we need to connect the dots to find the specific right answer, which is who did it. Making sure we exploited that to the maximum was crucial.
How to Survive the AI Coding Era
AI is going to write so much more code, and no one really understands all of it. If you wrote it all, you have in your head how it fits together. But as systems get bigger, and we work with some of the largest Fortune 100 companies, no one has full context. No team has full context, because the system is so complex. And that's happening not just at large companies but at small ones too, because the code is no longer being written by an engineer but by an AI system. That lack of context means that when an incident happens, it's so much harder to debug, because you don't have the context you need. It's already happening. Most people now are not developing code, they're validating, QA-ing, or troubleshooting it, and that doesn't scale. That's the big problem.
Here is the second big problem. With all the developments in AI software engineering, our belief is that engineers should get to do the really creative, fun work, architecting and system design. But over time, all engineers will be doing is troubleshooting, which would be sad. To make the creative work a reality, you need systems not just for developing and building software but for maintaining it. All of software maintenance needs to be reinvented.
You just have to persevere. If something goes wrong, that's okay. In some ways it's a good thing, because if it were easy, everyone could do it. This is one of those problems where the statement is very easy to make but actually solving it is really hard. A lot of companies are trying, and very few are succeeding. Staying resilient, having grit when something doesn't work, and sticking with the problem is a big part of what differentiates us as a company.
Think about what's going to be important ten years from now, you have no idea what's happening, and you have to find ways of creating structure from nothing. In typical hard jobs, in finance or technology, you have a point A and a point B, and it's very hard to get from A to B. In research, and in entrepreneurship, you don't know where point A is. You don't know where you are, and you don't know where point B is, and if you did, it would still be hard. You're constantly in this game of deciding where you are and where you're trying to go. What's interesting about this job is that the stresses can be very high and the lows very low, so getting used to very high highs and very low lows matters. But the one thing I've learned is that my recovery time is very fast. Even when I'm super stressed and super tired, if I take one or two days off, I'm back to full force. If you really enjoy what you're doing, even through moments of stress beyond anything you'd face otherwise, if you really love the problem and are attracted by it, that's what attracts other people too. We all love the problem, and we feel a deep desire to solve it.
Probably the most important thing of all is to surround yourself with people you care about, people you like, people you want to be like. If you do that, life will be okay. In my research life, I found the right PhD advisors, who molded and guided me. In this world, I found the right investors who guided me at the start, and the right customers who helped define the product. You live and die by the people you surround yourself with, and they guide you through the different parts. Building a taste for the right people to mentor you and partner with you is the most important thing, and the rest will figure itself out. But I wouldn't start a company just for the sake of it. You should feel something deep inside you, because it's not easy, and you need that deep belief or motivation to sustain you over time.