What: Luma builds multimodal general intelligence, starting with video models like Dream Machine.
Traction: Luma has raised $200 million from Andreessen Horowitz, Amplify Partners, and Matrix Partners. Nvidia, AMD, and Amazon are also investors.
In this conversation, Luma AI co-founder and CEO Amit Jain discusses building multimodal general intelligence, starting with the world's best video models. Amit believes the future isn't just about generating images or videos—it's about building entire worlds through AI that understands multiple modalities. Watch to discover his journey and insights on building revolutionary technology!
Key Takeaways:
Reverse Pricing: Charging More to Find Out Who Actually Values Your Product
Luma priced Dream Machine at $30, then $100, then $500 a month. There was no research behind the numbers, only a bet that whoever paid the most was worth a phone call.
The Worst Response to a New Product Isn't Criticism. It's Silence.
Apathy is the outcome founders should fear most, not complaints. A thousand complaints from paying users in a busy Discord means people care enough to keep showing up.
Why Luma Didn't Scale Its Video Model Until OpenAI Proved Scaling Would Work
At twenty people, Luma had nowhere near OpenAI's compute and couldn't justify scaling on faith alone. Once Sora showed that scaling actually worked, Luma poured in resources and had its first Dream Machine model three months later.
The Hollywood Test: An Unpolished Model, a $500 Customer, and a Request for Movie Rights
One of Dream Machine's earliest paying users turned out to be a well-known Hollywood art director. They used the unfinished, not-good-enough-for-film model to recreate a scene from a famous movie in thirty minutes, then asked Luma to release the rights.
Luma's Real Product Team Lives in a Discord With Its Angriest Users
Luma keeps 2,000 to 3,000 of its most engaged users in one Discord alongside its research, product, and engineering teams. With large models, there's no way to test every capability in advance, so watching what users actually do with it is the only real feedback loop.
How Do You Know a Problem Is Worth Building a Company Around?
Being interested in a problem and being unable to stop thinking about it are different things, and only one of them survives contact with real effort. The test is whether going deeper into a problem makes it more exciting or less, since most people quit the moment it gets hard.
Watch the full interview now on EO's YouTube channel! Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.
Introducing Amit Jain, Co-Founder and CEO of Luma AI
Hey, my name is Amit. I'm one of the co-founders and CEO of Luma AI. At Luma, we are building multimodal general intelligence, and that starts with building the world's best video models.
We have so far raised $200 million from Andreessen Horowitz, Amplify Partners, and Matrix Partners. Nvidia, AMD, and Amazon have invested as well. Luma is training models that are able to learn and generate video, audio, and text all together. Our premise is that by building models that train like the human brain, we will be able to not only generate worlds but actually solve the limitations of LLMs.
How I Found a $200M AI Product Idea
I grew up in India. As a child, I was deeply interested in physics from very early on. I spent most of my time learning about advanced physics and trying to understand how the world works. I studied math and physics in college, and I was about to go to graduate school for a PhD in physics.
Around the same time, I also started building iOS apps. My first one, for solving differential equations, got relatively popular in the App Store. That was interesting: once you make something other people find useful, whether it's a large number of people or a small number, it changes you.
Then a couple of my friends started a company, and I left to join them instead. As time went on, their company was acquired by Apple. At Apple, I started working on an agent system called Shortcuts, or Workflows. When we shipped that, I learned there was a part of Apple where something new was being built. Some of my friends had already gone there, and they went completely silent about what they were doing. It was this insane secret: if you don't want me to know anything, don't tell me it's a secret.
I wanted to figure out what was going on, and it turned out they were building this new thing called Vision Pro. So I joined the team working on this ambitious project for 3D capturing the world with Vision Pro. The idea was that if you're wearing this thing, we can take you anywhere, because we have full control over what your eyes see.
This was really close to my heart, because anything to do with simulating reality draws me in. I worked on that project for about three years. This was also around the time when huge things were happening in language models.
In 2020, two things came out that hit me. There was the DALL·E paper from OpenAI, and a paper on 3D reconstruction called NeRF, or Neural Radiance Fields. That was interesting to me. We were doing all of this procedurally: writing handwritten code for rendering, for capturing, and for reconstruction. If these two things worked, you could learn to represent the world instead. You wouldn't need to write any rendering code or graphics algorithms.
What DALL·E proved was that you could generate images from nothing. That was huge for me. NeRF gave you static 3D worlds, and DALL·E gave you images. But our world is extremely dynamic. There's a lot happening: we walk around, cars move, clouds move, all these kinds of things.
The question was: how can we actually simulate that world? So I started experimenting. All I did was try these methods, write all kinds of networks, and train them. Within about three months, I was convinced this is how most of the video things humans see will be made.
At work, I was still doing all these traditional techniques. So my choices were: stay at Apple and try to convince this giant organization that this is the future, that I need $100 million to actually do this, or go find fifteen or twenty of the most brilliant people on the planet who could do it for us.
I tried this at Apple first. Yeah, this wasn't going to happen there. I left, and we started Luma in 2022. In 2023, we built a 3D generative model called Genie. At the time, large-scale infrastructure like training systems and encoders didn't exist. So we had to build and invent all of it ourselves.
We built a lot of large-scale data collection systems, large-scale training systems, all those kinds of things. It took about a year, year and a half. Then we started work on the first video model Luma released in 2024, called Dream Machine. Around that time, our chief scientist Jiaming had joined our team. Jiaming had been leading image and video generation research at Nvidia.
At the time, these new training chips were coming out from Nvidia, the H100s. When we looked at that for the first time, we thought: the chips are capable enough. We think we can gather enough data, and the algorithms are there from our previous work on 3D generative models. We can actually do this.
So we started that process, and it took about four and a half months to go from having some of the infrastructure to training a proper model we could release to the world. But that came after a year and a half of work on encoders, learning models, and building infrastructure. It was also on the heels of OpenAI announcing Sora in February, which was really interesting.
Before Sora, our video efforts were smaller. We were a very small company at the time, barely twenty people. We had a good amount of compute, but nowhere near OpenAI's level. We couldn't have scaled without knowing that scaling would work.
Once that evidence was in front of us, we scaled our efforts significantly. Three months later, we had the first Dream Machine model. It was funny; it was a very early model. Today you wouldn't consider it very good. But at the time, it was huge, and it got incredibly popular.
It was on Good Morning America, it was on CNN, and people were absolutely astonished and mesmerized by it: you can generate video. That made Luma into a well-known household name, and it gave us all the resources we needed to continue our research and our work.
That was the first Dream Machine. Whenever you're developing a new capability, a new technology, it never works the first try. It never works the tenth try. I'd say: if you've found a good market and you want to find what fits into it, iterate like hell.
Anything that slows down your iteration velocity, avoid. If you have an engineering mindset like mine, you want to build a very stable, extensible system with all the right technologies. But sometimes those systems slow you down, because you end up building a monolith, and it's good. It's very scalable, it's all the things. But the problem is, to change one thing, you now need to change five modules. Compared to that, a barebones thing you built in Python, you put it up there. You have no allegiance to it, so you don't care. You just iterate. You change it all day, all night. I think that's important.
Why Hollywood Paid $500 for Our Imperfect Product
Honestly, we thought not many people would try it, because it's a new thing. We knew eventually the demand would be huge, but initially we thought not many people would try it. But the number of people who tried it was insane.
So we put some pricing on it, to see how much people were willing to pay. I had this extremely unscientific way of doing it: make it expensive. We'd give a small number of videos for free. Then it would be $30, then $100, then $500. The reason was solely to discover who was willing to pay for it. If someone is willing to pay $30, they're getting something out of it. If someone is willing to pay $100, they're clearly making money from it. If someone is willing to pay $500, I need to go talk to them, because I need to understand why they're paying so much per month to use this product. So we took the approach of putting something out in a very unpolished state. If someone found it valuable, we would learn a lot, and we did.
So this was one of the people paying $100 or $500 a month. I was curious what they were doing with this thing. I got on the call, and this person was grinning ear to ear. They said, "I'm really glad to meet you." I was like, wait, is that you?
They were well known, especially among art directors and people in the world who make movies. What I learned was that they were trying to create a scene for one of the more famous movies. They weren't able to do that with their traditional techniques. They tried Dream Machine and got a great scene out in about 30 minutes. Now this person was talking to me because they wanted me to release the rights so they could use it in the movie. I was blown away.
The first version of the model really wasn't good enough to be used in movies generally. But here was someone who found even that useful in the early days. Put it out. Talk to your users. See what they're doing, what they're not doing. Talk to them: "How are you using our stuff? If you're not using our stuff, how are you not using it? Do you know about our stuff? How did you find out? How did you not find out?" We have a group of about 2,000 to 3,000 of our most engaged users on Discord. Myself and our research, product, and engineering teams are all in that chat all day. When these are your most dedicated users, the paying users who spend a lot of time on your platform, they have a thousand things to complain about. That's very good.
The worst thing that can happen to someone making anything in the world is apathy. You put something out, and nobody cares. That's the worst scenario. The second-worst scenario is you put something out and everybody is very happy with it, because that means there's nothing else left to do. Generally, if everyone's just telling you all good things about it, that probably means they either want to interview you or they're lying to you.
Unlike traditional products, you know exactly how feature A interacts with features B and C. You know the whole state very well. With large models, people can do anything with it. People can generate anime. People can generate videos of a tomato rolling down a hill. With large models, you don't know what your users are going to do, and there's no physical scenario in which you could test all the capabilities of the model. So you really have to see what people are doing with it, where it's succeeding, where it's failing.
We're Not a Video Company: A Path to AGI
We are not a video model company. We are not building video models or image models. Our goal is very simply to solve multimodal general intelligence. What does that mean? Our belief is that people don't want image generation models or video generation models. What people want are worldbuilders.
Every video, every movie is a world, a universe someone created. Whether you're talking about high fantasy like Lord of the Rings, that's a whole universe with its own laws of physics and characters. Or if you think about a YouTube video or TikTok, they create a personality, a persona of their own. We need models that let people create worlds and then hit play. For that, you need to build a very different kind of intelligence. Text alone is not enough to build it.
Think about the way humans learn a concept. We see it with our eyes, in video. We hear it with our ears. We reason about it logically in text. Everything humans learn, day in and day out, doesn't just happen in text. It happens across all these modalities. So if you want to build intelligence that can collaborate with humans digitally and physically, that can understand us and entertain us, you need to build intelligence trained on all the data the human brain is trained on. Video is a big part of that. So we believe that multimodal intelligence, multimodal data, is on the critical path to AGI.
When you're working on a problem that's worth solving, that really motivates you. I find a lot of things interesting. But finding something interesting, versus finding something you're so mad about that you want to do it, is very different. Passion could be a moment in time; it can come and go. "Oh, this problem, it would be so great if you could solve that," and you imagine spending your whole life doing it. Generally, that's not the case. My suggestion would be to try a lot of things and go deep into them. What you want to find is not whether you were interested in the depth of the problem, but whether that depth gets you more excited or less excited. When you go deep into something, you have to put in effort.
Sometimes you put in effort and think, this is as boring as it gets. Don't do that. If you've found something where the harder it gets, the more excited you get, that's a unique thing. Honestly, I can guarantee you that most people around you will quit. Try a lot of things. Go deep into them. See if you can stay excited, not artificially, even after you're done thinking about the problem for the day. You can't stop thinking about it. It comes to you at night. It comes to you like, "But how do I do that?" If you can find that, and somehow that also happens to be an opportunity a company can be built around, that's it.
Join the 1.5M+ founders inbox to get the latest updates.