Who: Lin Qiao is the CEO featured in this interview, speaking about the work behind Fireworks AI. The discussion connects Lin Qiao's background to company-building choices, focusing on the problem being solved, the people served, and the evidence guiding decisions as the organization develops.
What: The interview examines Fireworks AI and the problem behind Fireworks AI's Mission to Build AI. It follows how the subject tested an idea, learned from users, and made choices about product, technology, team, and growth. Founders can use these details to evaluate ideas before committing resources.
Traction: The transcript records concrete company progress that makes a huge difference. It also follows uncertainty and customer learning. The practical lesson is to connect growth with evidence, measure the response, and adapt the plan as new information changes the opportunity.
In this interview, Lin Qiao discusses work associated with Fireworks AI and the lessons that emerged while building it. The conversation covers Fireworks AI's Mission to Build AI, with concrete examples of how the team found a problem, tested a product, responded to setbacks, and chose what to do next. Readers will learn a practical framework for turning insight into execution while staying close to customers, changing evidence, and the tradeoffs involved.
Key Takeaways
Fireworks Builds Infrastructure for the AI Transition
Lin Qiao says Fireworks AI was built from Meta experience to help companies navigate the AI first transition without a hundred person machine learning and infrastructure team. Its traction, including 150 billion daily tokens, reflects a playbook of turning hard infrastructure needs into shared capability.
Cutting Latency and GPU Costs Enables Scale
Fireworks designed its software stack to reduce GPU requirements, because slow responses make products unappealing and expensive hardware can erase a customer's economics. Lin Qiao connects infrastructure efficiency to viable businesses that can handle ten, one hundred, or one thousand times more customers.
Startups Must Pursue Tenfold Improvement and Speed
Lin Qiao argues that startups should avoid incremental projects and target tenfold gains in adoption, latency, or scale. She recommends asking whether each project will visibly change business metrics, then saying no when it does not, even if that decision disappoints people.
Compound AI Systems Route Questions to Specialists
Lin Qiao sees text-only language models expanding into image and audio systems that resemble real-world communication. Fireworks combines specialist models, function calling, and external APIs so an application can route each question toward the right source instead of forcing one model to answer everything.
Aptitude Matters More Than Experience in Hiring
Lin Qiao hires for hunger, learning speed, and problem-solving aptitude because fast-moving technology quickly makes experience obsolete. Her improvement habit is to compare results with expectations, ask why they differ, and work with the team on small and large process changes.
Watch the full interview now on EO's YouTube channel! Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.
Introduction
Hi I'm Lin. I'm CEO and co-founder of fireworks AI. I started Fireworks AI with my co-founders late 2022. We have a long time experience working at Meta. Building AI infrastructure from ground up. And we work with many companies in the industry and experience their pain going through this AI first transition.
Without teams like us, without no house, without proper hardware, without the right software and tools. Today we're processing more than 150 billion tokens per day and generating more than 1 million images per day. From the funding point of view, we have raised the first round from benchmark 25 million, and then recently we just closed around led by Sequoia with 52 million investment and post-money 552 as the valuation.
So our valuation grew by four times.
Fireworks AI's Mission to Build AI Infrastructure
Throughout my career, I've been through waves of technology enabled business transformation, waves of it, and that particular intersection of technology and business deeply interests me. And then we go through mobile first transition. That's a huge tectonic shift. I want to join a company where they are on the forefront of this transition. That's Facebook. When I joined, Facebook just finished the transition from desktop to mobile first. So when I initially joined meta, I joined the data infrastructure team.
And when I'm running the team, I look at the stats of data growth, and the data growth is mind blowing. It's just so fast it doesn't make sense I asked myself, I need to figure out what's going on here. And I did a drill down and figure out, oh, majority of data growth is driven by AI. That is clear to me. That's the future. And there's a lot of unsolved problem because there's a new emerging area.
So I moved to the AI space, and that's why I started to build a team from five people to 300 people over the course of five years. So PyTorch is very widely adopted. For example, many of us use Netflix or other streaming media. The content on Netflix is heavily personalized. Many of those recommendation models are PyTorch models. PyTorch has been heavily used in robotics. Self-Driving cars.
Almost all self-driving car companies use PyTorch, including Tesla. Me and my co-founder has spent years at Meta, and we have hundreds of people building all the AI infrastructure, supporting models, AI first transition.
While I also have many friends working other big companies in the industry, they are way behind Meta and when they go through the first transition, they don't have hundreds of people, machine learning team or AI infra team to help them. So our mission of starting fireworks AI is to enable new businesses to flourish. Building on top of this innovative AI technology without 100 people, machine learning, engineering team and infrastructure team.
Generally AI is very big, the model is very large and it's just very slow to run these models. So not having very low latency and fast response makes that product not appealing at all. So we know this is a big pain point. Another big pain point is the general AI model is so big and they have to run on GPU and they have to acquire a lot of GPU. And GPU is so expensive.
If they're lucky to get the GPU, it's they have never seen such a big bill before. So we designed our software stack to minimize your GPU needs to solve your problem. In that way, it's significant control the cost operation cost for the company. So then our customer has will have a viable business when they scale quickly to ten x 100 x 1000 x more customers.
Within the past half a year, our traffic grow by 100 times. Today, we're processing more than 150 billion tokens per day and generating more than 1 million images per day.
If there are a few key things about startups, one is to not work on incremental work. Do not pick on incremental things. It's human nature to work on projects that we know we can deliver. But it's not for startups. Startups are pushing for ten x ten x faster customer adoption, ten x faster infrastructure latency ten x higher scale. Startups are only here for ten x. It's for a huge leap.
The second is startups are not for slow moving pace. It has to kind of outrun and outpace competition. There should have already been incumbents occupying this, occupying the space. And the speed is advantage of startups because we don't have the burden of coronation. We don't have the burden of Slow decision making. Be able to say no is essential for moving fast, because over time it's you can justify hey, why we should add this, why we should add that.
And then the same person got time slices into multiple different top priority things. So I think laser focus on prioritization and really ask hard question does this project really move the bottom line of our business? Can we visually see that our metrics will change? I think that have the principle to drive that conversation and be able to say no, it's very important. Again, the goal is not to make people happy.
That's not the goal. And the goal is to make a solid product decision or strategy decision and have everyone laser focus on delivering that. That's the key point. Starting this company were constantly asking ourselves, are we working on the right problems? Can we deliver this today and can we move faster?
I think that mentality is deeply ingrained into us, and that's also essential to, I think, to build a successful startup is kind of have a huge sense of urgency to keep asking, why not today? Why not yesterday? Why not faster?
How To Survive AI Era As A Software Engineer
2023 has been a year where we laser focus on large language model with texting text out of this year and beyond. That will not be sufficient because a lot of business tasks will emulate what's happening in real life. And in real life, we communicate way beyond text. We speak and we use visual to collect signals.
And so that's where we have to expand beyond large language model into models who understand images, who can generate images, who understand audio, who can generate audio. And I firmly believe the direction for the industry to move forward to is to into Multi-modality. And today on our platform, we already serve more than 100 models across all these modalities, from large language models to audio models to image generation models.
We have a very broad variety of modality for our customer to pick and choose from. With that said, it's still not enough because every single model has limited knowledge, and the fundamental reason is those models are like child in the school. They learn from textbook and those gen AI models learn from training data, and training data is finite. It's not infinite. If you ask the question, it will have to give you an answer in a probabilistic way.
When it goes out of its knowledge, it's going to hallucinate. So that hallucination is a big problem for application developers. And the way to address that is to let each model focus on its specialty. Only answer questions when it's good at, and we are building a power routing layer. It's called function calling. So we have our proprietary knowledge to to figure out which specialty model to routing to, to give the best answer.
It does not just routing to different models, different modality. It can also routing to APIs. APIs could be search, could be weather, could be stock price could be many other things. So then this model is able to pull together the totality of knowledge to give the best answer to to our users. This is the form and shape of compound AI system, and fireworks is growing into and building into our next generation compounding systems.
Very simple, aptitude. It's not about experience. I actually prefer aptitude over experience. I need to see the fire in the belly. Super hungry, super motivated. That trumps anything else because this is a fast moving technology. Everything is new. How fast you pick up, how fast your learner. How determined your problem solver. That makes a huge difference. In the early stage of my career I have this deep imposter syndrome.
So I think a lot it's actually a way of self-reflection, and I constantly think about what I can do better, how I can do things differently. But later on it becomes a habit. It just when we take some action, we do things. I will observe, hey, is this going as expected? And if it's different, then why it's different and how we can be more efficient?
And what are the small changes or big changes we can introduce to the process, to the way we do things to to be better, to move faster. So and then I share my thoughts with my team. Sometimes I will ask them, hey, what do you think? How this is going, what we can do differently. So just kind of build this muscle over time.
We just think together as an as an, as a team to, to kind of always there's always room to improve.
Join the 1.5M+ founders inbox to get the latest updates.