Dec 17, 2024

This Insane AI Video Search Technology Selected by NVIDIA and Snowflake

Interview with Jae Lee, Co-founder of Twelve Labs

Founder Focused

💡
At a Glance
  • Who: Jae Lee is a co-founder and CEO of Twelve Labs, who started the company with two fellow soldiers while serving in South Korea's Cyber Command.
  • What: Twelve Labs builds video foundation models that let developers and enterprises search, classify, and summarize video content through APIs.
  • Traction: Twelve Labs now serves more than 20,000 developers, with adoption from the world's largest creators, media and entertainment companies, sports organizations, and law enforcement, alongside a product partnership with NVIDIA and backing from investors including Snowflake and Databricks.
Today's story is about Jae Lee, the CEO of Twelve Labs. Twelve Labs is developing a multimodal video understanding LLM that comprehends various forms of information. For example, if you search for a specific scene in a video, Twelve Labs' technology can pinpoint and retrieve just that scene. Their technology has caught the attention of NVIDIA, top global creators, and even the professional sports industry. Recently, they also secured an additional $30 million in funding from Snowflake, Databricks, and others. An interesting fact is that Jae started the company while serving in the military. In this video, we have expanded on parts we missed in the previous one to fully capture the founding story of Twelve Labs.

Key Takeaways:

Why Twelve Labs Bet on Video Instead of Text or Image AI
When the transformer paper landed in 2017, capital alone could win in text and image foundation models. Video was still wide open, so the founders chose the frontier where being passionate and a little reckless mattered more than fundraising size.
Getting Noticed Matters as Much as What You've Built
If an idea is truly impactful, someone else is likely already pursuing something similar. The determining factor isn't originality alone, it's whether you find a way to put your work in front of the right people.
Never Get Comfortable With the Scale You've Already Reached
When a VC asked how long it would take to index a billion hours of video, Twelve Labs was still thinking in terms of a million hours and ten years. No matter how impressive your current scale feels, someone else is already imagining an order of magnitude beyond it.
Why the Best Customer Is the One Who Says No
Twelve Labs once fought hard to convert a reluctant early customer into a paying one, only to find the customer never actually used the product. The lesson was to optimize for hearing no quickly instead of forcing a yes, since that's what finds true innovators faster.
The Mentor Advice Test: Would It Turn You Into an Underwear Company
Jae compares taking every piece of mentor advice to ending up building something unrecognizable from your original vision, like an underwear company instead of an AI company. Learning to say thank you, but no thank you, to well-meaning advice was how he grew most as a founder.
How an Oracle Event Turned Into an NVIDIA Product Partnership
A brief conversation with Jensen Huang at an Oracle Cloud partner event led to Twelve Labs being featured at NVIDIA's 2023 GTC. That exposure brought NVIDIA's venture team back with an offer that went beyond capital, into an actual product partnership around video understanding.
The Real Twelve Labs Bet: Mapping Language Onto Video
Twelve Labs' core thesis is that mapping precise human language onto video content unlocks emergent capabilities like search, classification, and summarization, without hand-built rules for each one. That single mapping problem has scaled the company from a $2,000 bagel-shop prototype into a platform serving over 20,000 developers.
Watch the full interview now on EO's YouTube channel! Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.

Introducing Jae Lee, CEO of Twelve Labs

Hi, my name is Jae. I'm one of the co-founders and CEO of Twelve Labs. Twelve Labs is an AI research and product company based here in San Francisco and Seoul. We are building video foundation models for developers and enterprises building video-centric products. We build humongous AI models that can understand videos like humans, and we serve them to developers via APIs, for teams looking to build really powerful semantic search, classification, or summarization into their products. We currently have a little over 20,000 developers actively using our search API. We have the largest creators in the world adopting Twelve Labs, as well as media, entertainment, large sports organizations, and law enforcement.

Chapter 1. A Startup Founded during Military Service

I was born in Seoul. I spent about ten years there, then had a chance to move to the States. I moved when I was eleven, so I went through elementary and middle school in Knoxville, Tennessee. I was able to pick up a lot of the culture, and I was very interested in expanding my perspective and exploring the new world.
My first experience with software engineering, or at least coding, was Matlab. My uncle was getting his PhD at the University of Tennessee, and I'd see him plotting distribution graphs and things like that, which made me curious what he was doing. That's how I got into playing around with small data sets and doing the same thing he was doing, because I wanted to be relevant and I wanted to talk to him about a bunch of things. He probably thought what I was doing was pretty cute, which sparked my interest in learning more about how we capture all this data and create a system that can really understand the distributions of the things of the world. It just kind of felt like if you had that understanding, it gives you the power to predict anything.
I went to Berkeley for college and studied computer science, so I really geeked out on AI and software engineering. I spent about 15 years in the United States, so half of my life was in Seoul and in the States. I was drafted into an organization called Korean Cyber Command, where like-minded people were already there, armed with incredible knowledge in software engineering and AI. I joke about it: I served the country with keyboards rather than a rifle.
I was really fortunate to have met my chief architect, SJ, and then Aiden joined in too. I still clearly remember Aiden with his buzzcut, coming in from boot camp. From day one, we knew we had this common interest in AI, and in what we could do as young scientists to really push the frontier of AI development. We spent a lot of time reading papers, discussing, arguing.
It turned out there were two clear paths. One was that after the military, we'd go pursue a career in academia and become professors. The other was starting something of our own. Looking back, what we realized is that we were spending so much time together, and the military is this really special setting where you're basically jamming together 50 to 100 twenty-year-olds, right, with testosterone. What we thought was, if we're having this much fun in the military, imagine what we can do when we get out. So it was pretty clear to us that we were going to start something.
I think we spent about a year and a half really thinking about what the next frontier for AI was, and how we could contribute to pushing that boundary. There's a seminal paper called "Attention Is All You Need," which some people call the transformer paper, and it's making a lot of impact now. When it was published in 2017, at least for text- and image-based foundation models, capital was probably going to be a moat: whoever raised the most money. What we realized was there was still a lot of unexplored research for multimodal video understanding, and there, rather than capital, it was probably going to be really passionate, smart, but also slightly dumb people, dumb enough to start, who would have a really good chance of succeeding. We also realized that with the explosion of video and other complex multimedia data, it was going to become the infrastructural data for the internet, and that developers and enterprises needed something better than object detection or transcription to make sense of all this video data that humanity is creating. So it was a no-brainer to start building for video understanding.
The model that Twelve Labs is building is basically trying to map human language onto whatever is happening in video content. If you can map precise human language to what's happening within video content, that gives you emergent capabilities, like being able to search for things really well, or being able to classify or summarize them.
We didn't all join Korean Cyber Command at the same time. SJ was already about six months ahead of me, and Aiden was six months behind. So we decided we were going to start this company, but SJ was leaving the next year, I was leaving the year after, and Aiden was leaving six months after that. So how do we do it? It was genuinely very scary.
So we had our ideas. I remember SJ was discharged on a Thursday, and he came back to the military base that Saturday with our laptops. He took us out in front of the base, to a bagel shop called Last Bagel, and that became our office. SJ would bring all of our laptops, and we'd do our research and a little bit of prototyping there. We did that for six months. Then I got discharged, and I did the same thing, carrying laptops and taking out Aiden. We did that for about a year, until everyone was out.
We had a bunch of friends working in AI, crypto, and blockchain, and Web3 was just booming. We had a mutual friend, the founders had a mutual friend, who had a really nice office in Seoul, and he told us we could come in and use the space. So that's what we did. After about three weeks, that company went bankrupt. It got scary. People started coming in while we still had our desktops and GPUs all set up there, so we got really scared and brought everything back out. We found a really tiny office, about the size of a dressing room, and that's where all five of us spent the next six months before we raised our proper seed round.
Looking back, if we were to do it again, I don't know if we'd be able to. But some people say ignorance is bliss, and I think we were just really naive and really excited about building this company. Not knowing what was ahead allowed us to do what we did.

Chapter 2. What a startup with only $2000 can achieve

When we first started the company and hired our first employees, they had a hard time explaining what Twelve Labs does to their parents. What is a foundation model? What is video understanding? It was a new concept that was hard to understand for people not in this space. Nowadays people talk about the foundation layer, the tooling layer, and the application layer, and everyone's very familiar with it. When we started Twelve Labs, I think technologically it made total sense, we knew it was going to happen, but what was uncertain was whether the market would accept it. We were at the verge of a breakthrough in building an AI that could get to a certain level of human understanding of video, so we were betting on the market's acceptance of foundation models.
The founders were pretty much broke. We'd spent two years in the military, and we had $2,000 to start with, barely enough to do anything. So we had to figure out what would be impactful given our current resources, something that would put us on the map, or at least let the world know that what we were doing was relevant. Our tactic was to talk to a bunch of customers. There were early believers in Twelve Labs who took our APIs and built awesome things with us, but we needed more exposure.
As a team, we decided to participate in ICCV, the International Conference on Computer Vision. They were running an awesome competition for video understanding. We talked to Aiden: we had nothing to lose and only to gain. The team was extremely supportive of Aiden spearheading that effort. All I could do to support him was offer some ideas and directional feedback, but we needed compute, and we needed the determination to put some serious cash behind it. Back then, $200,000 in compute was a lot of money for Twelve Labs. Thinking that we were going to blow through $200,000 in ten days of compute was really scary, but the team was able to use that precious capital and build something incredible, and it helped us win the competition.
I think the important thing is, if you're building something really impactful, something you think is going to significantly change the industry you're in, there will always be someone with a very similar thesis. It's just a matter of how you get yourself out there, how you let people know that you exist. For us, that was the competition.
After winning the competition, companies like Index Ventures and Radical Ventures, who had a very strong thesis around multimodal AI and were asking what's next, reached out to us. Video happens to be the most relatable multimodal data. These amazing companies came in inbound, and we started jamming. The conversations turned into the next conversation, and we talked about the technology, and it just happened very serendipitously.
For our first pitch, I was in Seoul, and my first call with Index Ventures was at like 3:30 a.m. Seoul time. We didn't have a pitch deck, we knew nothing about fundraising then, and we didn't even know this was going to be a friendly introduction. I just felt the need to build one, so I remember staying up till about 3 a.m. building it. The storyline was quite simple, we didn't have too much to show for it. The idea was: the problem we're solving is massive. Eighty percent of the world's data is in video, and there's no adequate solution out there for developers and enterprises to make sense of it all. That's the market we're tackling. We want to index all of that, eighty percent of the world's data. And this is the underlying research work we've done. That was it.
VCs asked a lot of hard questions. The most memorable one: if I bring you TikTok as a customer, and they want to index a billion hours of content, how long does it take? We were thinking maybe a million hours, and that it would take about ten years. That's when we realized we should never be comfortable with what we've built, there's this whole new, incredibly large world out there. Maybe some people are impressed that our system can index a million hours, but there are others thinking about a billion, ten billion, a hundred billion hours. That was a really challenging question, because what we'd said during our pitch was that we wanted to index all of the world's videos. That question kind of stunned me, and thinking about all the technical issues we had at the time, he probably found it funny. I was trying to give my best answer as a founder. The company ended up raising about $30 million in seed funding.

Chapter 3. Lessons from Acquiring Early Customers

My name is Soyoung Lee. I'm one of the co-founders of Twelve Labs, and I currently lead our business development and go-to-market.
We had a customer who was paying for our product but wasn't actually using it. We'd gone through a lot of work to get them as a customer, a lot of sales and relationship building. They were extremely early, and we had almost pushed the sale to happen. What we learned from that experience was that we have to optimize, even early on. It might have been better for us not to make that sale, because the customer probably wasn't ready. They didn't have the passion or the innovative drive that our other customers had. We were optimizing for hearing the yes. We tried so hard to turn that no into a yes, and we succeeded, but in the end, I think we probably should have kept it a no, and focused on the other customers where the yes was more clear, and who had a very clear vision of how they could build new experiences with the technology. Especially for earlier products and earlier technologies, where resources are limited, you want to build for your best customers and the innovators in every field. You should probably start optimizing for hearing the no rather than the yes, because that will help you find the right direction faster.
I think we learned through trial and error that the early customers we need to find and work with are true innovators in their field, whether they come from content creation, law enforcement, e-learning, and so on. We've had instances of trying to oversell. We went through sales 101 books and sales methodologies, and we'd pitch to the customer: here's the use case you could build out with our technology, we'll improve your ROI by x percent. We did make some sales from that, but I don't think it was something we should have spent so much time on, because if you find the right customer who's innovative, you don't need to explain anything to them. You show them a use-case-based demo, for us that meant indexing some videos that resemble the customer's, and then showing them how you can search or generate text very easily, just like a person would.
Had they been watching the video, they can draw out the full map of what they want to provide to their customers or users. This is a technology you have, and they can fill that gap in pretty easily. They already know what the return would be, or what the opportunity of that experience would be, even for the kind of really large, high-profile customers we have now. It was all the same process. We actually learned all of this from the customers, because seeing our demo and the early hints of the technology, they were able to teach us how they could utilize it. Even today, it's not just a single use case they want to power, they come to us with four or five different ideas of how the technology can impact different business units and optimize different workflows or build new experiences.
(Back to Jae)
I think we make mistakes probably every day. The one I regret most is that we had this conviction around building a foundation model, but we didn't have any data point on how to build that company. I blindly believed the startup mantra: identify a narrow problem and build a narrow solution for it. We spent a lot of early days thinking about what we'd do with this really powerful AI that could understand video. We knew from the get-go that we wanted to serve it to developers and enterprises, but we'd had mentors and other founders tell us we should build TikTok 2.0, or YouTube 2.0. We spent a lot of time thinking maybe TikTok 2.0 made sense, or maybe Gong 2.0, like sales call analysis, but that didn't really excite us, because we knew we were good at building infrastructure and helping developers build the next thing. That's probably the stupidest thing we've done: spending time thinking about things we weren't excited about.
As a founder and CEO, it's really hard to get distracted. If you are a first-time founder or a young founder, your mentor's advice means a lot to you. But having your own grounding, and relying a little bit on your own gut feeling, is very important, because what we've realized is that if you try to accommodate all of the advice you get from your mentors, your company will most likely become like an underwear company, totally different from what you wanted to build. So my key takeaway is having some fundamental foundation for yourself and for the company, and being able to say thank you for your advice, but no thank you. Having the gut to say no to someone you respect as a founder, I think I grew a lot.
Twelve Labs has a multi-year compute partnership with Oracle Cloud Infrastructure, where we get all of the state-of-the-art Nvidia chips. Oracle put together a small event for their key partners building foundation models. Aiden and I had a chance to meet with Jensen, because Jensen was at that event, and we had about 5 to 10 minutes to talk about Twelve Labs. It seems like he has a special place in his heart for computer vision and video understanding, that was one of the first use cases Nvidia chips powered. So we got to meet with Nvidia folks from that event, and then Twelve Labs was featured in Nvidia's 2023 GTC. I think that sparked other people from Nvidia to be interested in Twelve Labs, and Nvidia's venture team reached out to us. It was quite casual, we were talking about Twelve Labs and the future we're drawing, and the future of multimodal video understanding. I think the venture team also had an idea of how Nvidia and Twelve Labs could partner up as more than just a financial investment, but also think about a really robust product partnership. From then on, what Twelve Labs is doing and what Nvidia wants in vision and video understanding was just a perfect match. It happened quite naturally, from conversing about our technology and our roadmap, and Nvidia's future in producing really powerful chips for edge devices and smart cities. There's that natural fit of two companies' products really creating synergy.

Chapter 4. The Best Engineer is not the Best Coder

Nowadays I am focusing mostly on hiring. I think Twelve Labs is a group of great people, and I spend a lot of time meeting great people. I want to be able to recognize greatness when he or she comes in, good engineers, or even just good people in general, have this core values. I would go near their places and get together at a cafe, and we would speak for three or four hours. My way of deciding whether this person is a good fit for Twelve Labs is whether I'm able to learn from their core values.
Everyone's really good at coding nowadays, but great engineers can apply their core values and their skill sets, and are able to talk about the company they're excited about and how they want to impact it. How do you see the product evolving? How do you see our interfaces evolving? Some of the best engineers are not the best coder, but having that perspective, a really strong perspective and groundedness, is very important, and I try to look for that.
Twelve Labs' vision in the next two years is really becoming horizontal, video understanding infrastructure for all of the businesses and developers working with video data. We want to enter into streaming as well, real-time video data, and really become a visual cortex for modern video applications.

Join the 1.5M+ founders inbox
to get the latest updates.