Mar 17, 2024

How our AI Startup Grew 60x After a Pivot

Interview with Raza Habib & Jordan Burgess, Co-founders of Humanloop

Founder Focused

💡
At a Glance
  • Who: Raza Habib and Jordan Burgess, co-founders of Humanloop. Habib studied physics before completing a PhD in machine learning and working at Google AI; Burgess built his first website at 15, studied at MIT, and pivoted into machine learning.
  • What: Humanloop is a development platform for teams building AI products with large language models, providing prompt engineering, versioning, collaboration, and evaluation tools. The current version launched in October 2022.
  • Traction: Humanloop has grown about 60x in usage since launch, doubled its revenue in the last quarter, and raised about $7 million in venture capital from investors including Index Ventures and Y Combinator.
In this interview, Raza Habib and Jordan Burgess share how a two-day sales experiment validated their pivot into LLM tooling, why evaluation is the core challenge for AI product teams, and why it's better to build for the slight future than be limited by today's capabilities.

Key Takeaways

A two-day sales experiment validated the pivot to LLM tooling
Habib and Burgess gave themselves two weeks to find ten paying customers for their new LLM-focused product direction. They hit that target in two days, a level of market pull they had never experienced before. The strength of demand made the decision to pivot straightforward.
AI product development requires new forms of evaluation and versioning
Unlike deterministic software, LLM-based products are stochastic and subjective, making traditional testing inadequate. Humanloop provides product teams with prompt engineering, collaboration, versioning, and evaluation tools that bring software engineering rigor to AI development.
Smaller teams find product-market fit faster than larger ones
Habib argues that hiring too many people before product-market fit is one of the most common startup mistakes. A team of three can experiment, iterate, and learn from customers far more effectively than a team of 20. Staying small until confident is the better approach.
Domain experts, not just engineers, should drive AI prompt development
At companies like Duolingo, the people best placed to do prompt engineering are domain experts such as linguists. Humanloop enables those experts to collaborate with engineers on prompt development and performance measurement without being blocked by engineering gatekeepers.
Build for the capabilities of near-future models, not today's limitations
Burgess warns that startups building only for current model capabilities risk becoming victims of AI progress rather than beneficiaries. While building reliable agents remains difficult today, founders should target slightly ahead of where models are now, since they will improve rapidly.
Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.

Introducing Raza and Jordan, Co-Founders of Humanloop

Raza Habib (co-founder and CEO of Humanloop): Hi, my name is Raza Habib. I'm one of the co-founders and CEO of Humanloop. Humanloop helps companies that want to build AI products with large language models to both develop and then make those products reliable.
Jordan Burgess (co-founder and Chief Product Officer of Humanloop): Unlike software engineering, when you're dealing with deterministic code and you can predict how it will behave, using AI is non-deterministic. It will produce different outputs depending on very minor changes or other artifacts. That's why evaluation is core.
Raza: Typically, the people using Humanloop are product teams who are building an AI-focused application, collaborating with an engineering team to get it into production. We launched the current version of the product in October 2022. We've grown about a factor of 60 in usage since then. We doubled our revenue in the last quarter, and we've raised about $7 million of venture capital from some of the leading investors in the world, including Index Ventures and Y Combinator.

Falling in Love with AI 10 Years Ago

Raza: As an undergrad, I studied physics. I was just really curious about it. I've always been a little bit nerdy, and my first job after university, I joined a small venture-backed startup called Justpark. It was a startup that was doing peer-to-peer car parking rental, and I joined that company when they were only five or six people. By the time I left, it was 45 people. It had grown to a million users.
That was a really exciting experience to scale a company from a very small stage. But it wasn't as deeply technical. Coming from a physics background, I wanted to work on something on the frontier of what was possible, and I also became increasingly convinced that AI was likely to be one of the biggest, if not the biggest, technological achievements that humans ever did.
This was probably around 2015. I decided to go back to school, and I did a master's and then a PhD in machine learning. Back in 2015, deep learning had sort of begun to take off. Today, we use ChatGPT. You have a conversation with it, it's incredibly fluent. It feels like you're talking to a person. In 2015, people were getting excited if you could produce a language model that could open and close a bracket, and things that now people take completely for granted, even two years ago, felt like science fiction.
As soon as you start learning about machine learning, it's hard not to be fascinated by it. It touches so many different things simultaneously. How can I build a machine that can learn in the same way that the human mind learns? So there's an interesting technological question. There's a question about how human intelligence works. You just want to understand how the brain might work. So there's a biological question that's fascinating.
And then there's philosophical questions as well. Is it possible to build a machine in principle that can do all the things that a mind does? Will they be conscious? Will they be sentient? So it's a fascinating field in and of itself. As someone in physics, it feels like the most exciting time in physics was the start of the 20th century when quantum mechanics was being developed. But I feel like the people who were doing that then, if they were alive today, would likely be working on AI. I feel like it's the most interesting intellectual problem of our time.
When I finished the PhD, I spent a little bit of time working at Google AI, but I wanted to be working in an environment where the connection to product was a lot tighter and be part of a very small team with tight deadlines, pushing towards something that feels like you can achieve something amazing, rather than the more relaxed environment of a bigger company.
I had really concluded that I wanted to start a company. I had a list of the ideas I was most excited about and the people I was most excited about. I was trying to go through the smartest people I knew. Jordan was very high up in my list. Jordan has an incredible eye for detail and taste in product. He's the kind of person who notices what makes the difference between a really good product and not.
Jordan: My name is Jordan Burgess. I'm the co-founder and Chief Product Officer at Humanloop. I've always been interested in technology, and I had a lucky break as a 15-year-old that someone bought one of my websites from me. It was Wikipedia, but for guitar tabs and lyrics. I had my first exit for four figures of pounds.
It's sort of a Steve Jobs moment of, the world is created by people no smarter than you. You've created something and people have created value from it and paid you for it. That's a transformative moment in your life. I was very fortunate that I could go on the exchange to MIT for my third year, and I just saw this ecosystem of startups, of people building stuff, and got very enthused by that as a path for me.
When I left university, I did start something in the medical tourism space. If you don't have passion for that area that can last ten years, then you're not going to really sustain the efforts needed to go and create that company. My journey then was to figure out, what are the most impactful technologies that I should be working on and what can I make the investment in.
Feeling around, understanding the space, and realizing that artificial intelligence is going to be the most impactful technology in our lifetimes. I decided to study machine learning with the intention that ultimately this would be where I could happily spend a career working in.

First MVP and Pivot

Raza: I remember going to visit Jordan in Cambridge. Then we started to get into rooms together. There's a room at UCL where all four walls are whiteboards, floor to ceiling. I remember we covered every wall with different iterations of an idea. People talk about the moment you have an idea in a startup, but I much prefer the analogy of an idea maze.
Because once you say, we're going to build tools for NLP, what does that look like? There are 1,001 decisions you have to make about who exactly your customer will be, what product you'll build, and what the product will look like between the initial idea and something you can actually take to market. It was really weeks of time spent in that room and speaking to customers before we got to the first version of Humanloop.
The first version of Humanloop looked quite different from the product today, but the first MVP that was good enough that we felt we could put in front of customers, I think took us about a month to build. The way the first version of the product worked was at a time when people still had to annotate data to train a natural language model. It looked like a UI for annotation.
In 2022, we realized that as the large language models improved, you wouldn't even need to have almost any annotated data anymore. The paradigm was going to shift away from hand-labeling data and fine-tuning towards prompt engineering, and maybe also some fine-tuning.
Jordan: There was a certain transition point when we built something ourselves, just as an internal hackathon, and realized just how capable these models have genuinely become. Speaking to a handful of customers to understand what their needs were, what we heard was the same consistent pain point. How do I know how this system is behaving when it is non-deterministic? How do I go and improve it? Because I genuinely don't know how to evaluate it beyond eyeballing examples.
Raza: We had this idea that there was an opportunity to help the people who were trying to build applications with this, but we weren't sure how big the market was or how pressing the need was. So we gave ourselves the sales experiment. We had an idea for what the product would look like. We sort of had a mockup. If we can get ten paying customers in two weeks, then we know that there's a really strong pull in this direction and the timing is right.
We kicked off that sales experiment, and after two days we had ten paying customers. It didn't take two weeks. We had never felt pull like that before. There was pull from our existing product, but not the kind of pull where you have the first conversation with someone and they basically say, take my money. The fact that the market pull was so strong, and we were so confident about the overall technology direction at that point, made the decision to pivot reasonably straightforward. It just felt like this was obviously a better trajectory to be on.
There are really two pillars to the product that we're solving for them. One is everything related to prompt engineering, versioning, and management. The other is evaluation. The product team uses Humanloop as their development environment for developing their prompts. They're in that UI, collaborating together, iterating on things. Their workflow before Humanloop is usually split between multiple tools. Often it's the OpenAI playground and Excel sheets and other things. There's no collaboration, there's no versioning, there's no history.
The aha moment for the product team is the realization that they as product leaders can drive the process of the development of these AI applications without being as dependent on engineers as gatekeepers. For the engineering teams, it's bringing some of the rigor of normal software development to developing LLM features.
These larger companies have to have confidence that the models are going to perform in the way that they expect in production. They're not going to say something embarrassing. They're not going to create any liability. They're going to do what they're supposed to do. That's hard to do with LLMs because it's subjective and because it's stochastic. In traditional software engineering, we're used to having unit tests and integration tests, and there's a whole process around it. There's good versioning. We give those teams something equivalent for LLMs and the ability to measure performance.
Companies like Duolingo who are building LLM applications, Duolingo has so many use cases. They have this Duolingo Max product, which is an AI chatbot that people can practice speaking with, but also they have to produce a ton of content internally for their applications that could potentially be done with AI assistants. Those are the opportunities and applications. But when they come to build these things, there's a lot of prompt engineering involved.
The people who are best placed to do that prompt engineering are the domain experts, people like the linguists. So you need some way to allow the linguists to collaborate with the engineers and to develop the prompts and then measure performance. You have to have confidence that it's going to work the way that you expect when you deploy it.

Advice for AI Startups

Raza: I fundamentally think that AI startups aren't a different breed of company from normal companies. They're still building a product that should solve a problem for end users. Fundamentally, your end user doesn't care how you're delivering the product to them. They just want to know, are you solving an important problem?
For me, one of the mistakes that most startups make is they hire too many people too early and they hire people before product-market fit. If you have a team of 20 people and you don't have product-market fit, it's really hard to experiment and iterate and learn from customers and change things. It's actually much harder to find PMF with a team of 20 people than it is to find PMF with a team of three people. You want to stay a small team until you're confident.
Another one is that a lot of startups right now are very excited about agents, AI systems that can take actions in the world and act on behalf of the customer. I think that is certainly the future. That's definitely the direction of travel, and very soon we will be there. But today, it's still hard to make those reliable. A lot of people are trying to build something that maybe is slightly overambitious for the current moment, but correct for the near future.
I don't think they're doing it wrong, but it's not surprising to me that it's hard to get it to work right now. I think they're correct to be building for the capabilities of the models a little bit further on from where we are now, because they are going to improve very rapidly. It's better to be building for that slight future point than be limited by the capabilities today.
Jordan: If you're not smart about expecting where the future will be, you will be not a beneficiary of these trends, but a victim of them. You don't want to be building something where if a smarter, better model comes out next week, your company is screwed. You want to make sure that you're riding that wave of improvements.
The question was, how does it feel surfing on this giant wave of AI? Things are moving incredibly quickly. It's exciting, it's tough, it's challenging. But why would you want to do anything else? If you go out there surfing on a little small wave and you're seeing this big AI wave over there, you're going to be a bit jealous.

Join the 1.5M+ founders inbox
to get the latest updates.