Who: Catie Cuan is the founder and CEO of ART Lab (AI Robot Technology), a Stanford-based roboticist, choreographer and researcher who was previously an artist-in-residence at Google before turning her human-robot interaction research into a company.
What: ART Lab is building a Vision Language Interaction (VLI) model, a robot-learning approach that scores a robot's actions by whether a human's emotional response improves, not just whether the task got done.
In this interview, Catie Cuan explains why she thinks robotics is having its own "ChatGPT moment," why making robots look human only raises expectations they can't meet, and how a Google art project called Music Mode changed how she thinks about machines.
She also opens up about leaving a finance job to become a professional dancer, watching her father hooked up to hospital machines, and why she now teaches Stanford engineers to ask not what they can build, but why they're building it.
Key Takeaways
The Real Robotics Challenge Isn't Intelligence, It's Human Interaction
Cuan argues that as billions of robots enter offices, hospitals and homes, the bottleneck isn't giving them more capability, it's making humans and robots legible to one another. Her VLI model tries to solve this by judging robot behavior on whether it improves a human's emotional response, not just whether the task got done.
Making Robots Look Human Only Raises Expectations They Can't Meet
Citing roboticist Rodney Brooks, Cuan says the more a robot resembles a person, the higher people's expectations climb, whether or not the robot can actually deliver. She points to her own family, including her 77-year-old father, as a check on whether a humanoid is really what people want.
A Google Art Project Taught Cuan What People Actually Want From Robots
While working among what a later peer-reviewed paper puts at more than 180 robots at Google, Cuan and composer Peter Van Straten built Music Mode, mapping each robot's movements to musical samples. Employees who had been indifferent to the robots started emailing her in tears, a reaction she calls proof that people need to feel something from the technology, not just be served by it.
Robots Don't Have to Look Like Us to Matter
Cuan points to the Paro therapeutic seal robot and Roombas as evidence that useful, beloved robots don't need humanoid form. She says the field has locked itself into a narrow, humanoid-first vision of the future and wants that aperture broadened.
The Question Isn't What You Can Build, It's Why
In her Stanford class CS334 Robots and Arts, Cuan pushes students past the question of what they can build with AI and robots, since almost anything is now possible, toward the harder question of why they're building it. She frames time as the one resource nobody can scale or reclaim, which makes the reason behind a project matter more than the project itself.
Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.
Introducing Catie, the Founder of ART Lab
Catie Cuan, founder of Stanford's ART Lab
My name is Catie Cuan. I am the founder and CEO of ART Lab. It stands for AI Robot Technology. I'm a robot choreographer, a roboticist, a researcher, an entrepreneur and an artist. My mission is to make nascent technologies like AI and robots more humanistic, whimsical and widely accessible. And I do that through business, research and art.
My definition of humanistic is human centered, so thinking about how we can drive interactions and experiences with these machines that help humans feel more empowered and safe, but also help them live healthier, better lives.
A Unitree humanoid robot
We're at the Stanford Robotics Center on the Stanford campus in Palo Alto, California. The idea was to bring together a variety of disciplines and engineering backgrounds to work on flagship projects in robotics like domestic robotics, industrial robotics, medical, underwater and exploration. Or in my case, education, art, culture, movement and dance.
Billions of Robots Are Coming, Are We Ready?
We're going to have billions of robots in our lifetimes. It only took 13 years for there to be a billion iPhones on earth. We have so many people working in robotics right now. And so we're going to see billions upon billions of these robots showing up in our everyday spaces. Offices, hospitals, hotels, homes, caretaking facilities.
When you take all of these robots out of these very niche, narrow vertical applications and you give them a lot more intelligence and put them in situ with people, you have this huge opportunity and also challenge, which is how do you make humans and robots legible to one another? And that's what my work is interested in.
Human interaction is a grand challenge for robots. It's been the case for the last several decades. Humans are really unpredictable. We tend to do and say things and move in ways that you might not always expect. The problems that I'm interested in are how can we make robots more legible to people, but also more animating, more inspiring, to delight and bring awe into our lives.
And so the thing that I have always been curious about throughout my whole career is what kinds of interactions can we facilitate that both inspire and surprise us, but also open the door into what technology can be, where it's not exclusively a tool, but it also is something that you could socialize with or learn from.
We are building a kind of model that we call a VLI, a Vision Language Interaction Model. This is the first kind of model that's existed which takes the impetus of what it should be doing from humans in the environment. LLMs are acting on human information and then map that again to a series of robot actions, but measure the success or the failure of those robot actions on the basis of a human response.
So we can tell, is the human increasing in positive affect or decreasing? Was that what they wanted or not? And so we think this is the first start for robots that can really naturally and intuitively socialize and interact with people. I think it's so exciting and it's so necessary because we're going to have billions of these autonomous machines. They're going to be interacting with billions of people. That's trillions of interactions. And we want those interactions to be safe, legible and clear.
An interactive exhibit at the Exploratorium
I grew up in the Bay Area from the time that I was really young. I always loved numbers and I always loved to dance. I loved math and I loved science, too. My dad's Cuban, so my whole family grew up social dancing. And I would say, if you don't know how to dance, you're probably not Cuban. I always grew up with those two very strong interests.
And then at some point when I was getting towards the end of college, I had interned here at a big technology company in the Bay Area. And I knew I really wanted to try what it was like to be a professional dancer, and that you only have a very small window in your early adulthood when you can go for it.
I was a little scared. I didn't know anyone who had been a professional dancer in my immediate sphere: my friends, my family. So when I graduated from college, I wound up taking a job doing private equity consulting in New York. I realized really quickly that wasn't for me, mainly because I had this childhood dream.
And I figured, I'm only going to be 21 once, and if I don't go for this now, I'll regret it for the rest of my life. And so I actually wound up leaving that job and then became a professional dancer in New York City.
I had my own dance company for a long time. I danced for the Metropolitan Opera Ballet at Lincoln Center in New York, which was a crazy experience. I'll never forget what it's like to bow on that stage.
Around that time, my dad also got sick. He spent a while in the hospital. I think my dad, seeing him be sick in the hospital, was really scary, because here's somebody who you care about and you love so much, and you want them to be healthy and safe.
I think the fact that he was surrounded by all these big machines that were supporting him, they were meant to help him literally survive. They were beeping and booping and had these proportions and colors and this sort of overall presence that I think was oppressive to him, and it was also opaque.
He didn't understand what they were doing, and there were very few opportunities where someone could explain it to him. And seeing him be that vulnerable, it occurred to me that a lot of other people are in that same position all the time.
One of the projects I worked on when I was at Google was called Music Mode. When I first got there, we had 200 robots, and they were all driving around Google buildings doing useful work. They were wiping tables, sorting trash, resetting conference rooms. They were doing valuable things.
And I was there as an artist in residence. And I remember going around and asking a few people who were software engineers, what do you think about the robot? And they were like, it's fine. It's a robot. It was kind of cute or kind of cool initially, and now it very slowly does these different tasks.
And I'm thinking, what a missed opportunity. Here's this amazing new flagship vanguard technology showing up in your lives, and you feel a little bit ambivalent towards it. I think for some people, that would be a success. For me, it was like, huh, we could be doing a lot more.
So my first idea was to get the robots to play music. I thought, why don't they sort some trash and wipe some tables, and then they can go over here and play the cello. We worked on that for a few weeks, and then my friend and colleague Tom came up to me one day and he said, you know, Catie, the robot generates a lot of data.
Catie Cuan during the interview
Instead of the robot playing music, what if the robot was the music? What if this data became the music? That was such a clever breakthrough that he came up with, because we then started working with a composer named Peter Van Straten on this piece of software that we called Music Mode.
And what we did was that every time a different part of the robot moved, we would map it to a different sample of music. For example, me as the robot, I open and close my gripper, it goes cling, cling, cling. Or if I turn the torso of the robot, it would make this beautiful bass noise like boom, boom, boom.
And Peter, who was our composer, such an amazing artist to work with, because he understood that you could move all these joints of the robot in any combination. So it was a real combinatorics problem. And so we worked on this for a couple months, and then we shipped it on the robot, this software called Music Mode.
And we would just be able to say, hey, robot, turn on music mode. And then as it's wiping the table, you just hear this symphony. Then I started getting emails from people at Google who I did not know that were like, this is the most incredible thing I have ever seen.
I'm sitting at my desk crying because I never thought that a robot could be so beautiful. I never thought that a machine could make me feel this way. So we shipped Music Mode on every single one of the robots. And it was crazy.
We would have nine robots that were sorting trash at the same time, and you'd go around and say, hey, robot, turn on music mode. And just this beautiful symphony arose. That for me was such a proof point that we need these experiences with our technologies.
Break Out of the Humanoid Imagination
We want to feel something, we want to feel engaged. We have been so influenced by fiction and media. There are dozens, I think, at this point, of humanoid companies that are making robots that look like humans. And often the argument that you hear is, well, they're made for human spaces, so we need them to look like people.
Rodney Brooks says it really well. He's a titan in our field. And he says the closer that something looks to a person, the higher the expectations people have for it. And you can't help yourself, your mirror neurons go off. You start having all these expectations that because it's bilaterally symmetric and it's roughly in the same proportions as a human, it should be able to do what a human does.
The expectations that you set are so high, and I think also for people, if I ask whether it's my 77 year old dad or whether it's one of my nieces and nephews, and I'm like, hey, is that the robot that you want? I think there isn't enough expansive thinking into what robots can be.
We've sort of locked ourselves, we meaning the broader robotics field, into a slightly myopic vision of what our future with robots can look like. And I would love for that aperture to be broadened.
There's this phrase called dirty, dull and dangerous. So for a long time, robots were useful for tasks that were too dirty, dull, or dangerous for people to do. And then I think there was appetite for robots to do other things that were not exclusively dirty, dull, and dangerous jobs.
It's just that I think we've still gotten locked into this paradigm of robots are for utility. They need to alleviate certain parts of physical labor in our lives. When I think robots can do a lot more than that.
The Paro robot was a very soft seal robot that was popular, I think, in the mid 2010s, that would just provide companionship for people who were ailing or lonely. There have been robots used for helping children with autism learn social skills.
A Unitree humanoid robot mid-flip
So I think we have a lot of opportunities for robots to show up in our lives and be beneficial for humans that are not exclusively humanoids that do our dishes and our laundry.
I've worked with a lot of different robots in my career. I've gotten to use Roombas in a dance project. Small humanoid robots, drones, big industrial robot arms, these mobile manipulators from Google. I've used quadrupeds from Boston Dynamics.
The very first experience I ever had dancing with a robot in Amy's lab in 2017 was transcendent. It was the first time that I got to imbue myself inside of something and then also interact with it. So I created this choreography for the machine, and then it performed the choreography, and then I got to dance with it, and I was like, whoa.
I felt myself on this very long continuum of humans and our tools, everything from early stone tools and fire to the wheel and the printing press and the modern computer and these robots. And I thought, wow, this is definitionally part of what it means to be a human, is to be in relationship with tools.
Aristotle used to say we shouldn't be called Homo sapiens, wise man, intelligent man. There are a lot of species in the world that are intelligent. We should be called Homo technae. The thing that makes us unique is our relationship to our tools and our ability to make tools.
What I've learned since then, oh, my gosh, so much frustration, so much debugging, so many limitations. How my relationship has changed is I've become much more acutely aware of how amazing human beings are relative to all of the robots that I work with. They don't have the expressive potential that people have.
Humans need to give ourselves a lot of credit. We can do a wide diversity of things. If you think about what was required to come into this building today, you had to open a door. It was a door that you had never seen before with a handle that you couldn't have predicted.
Catie Cuan during the interview
There are infinitely many different kinds of door handles with different friction and torque required to get the door open. You had to walk down stairs that you'd never seen before in a space with a proportionate move, a large object around, whether that's a backpack or something else that you're carrying.
You needed to then socialize and show up and make conversation and be acutely aware of timing, and to build the type of dexterity and acuity and specificity that we have took not only your lifetime to evolve, but also the many generations before you that encoded all of that stuff in your DNA. We are really remarkable animals.
I read a beautiful article yesterday about writing as a craft and how it's a hard earned craft in many cases, that anything that's worth doing, anything that is a craft in the richest sense of the word, takes effort.
I feel that way about many different things, not only writing, that anything that's really worth doing or worth earning is hard, otherwise everybody would do it. Otherwise it wouldn't be considered worthy.
I'm reminded again of this professor here at Stanford named Mark Skylar-Scott, who said, my goal in my research is to 3D print a human heart, and I think it's going to take me 25 years to do it, but it is a challenge that's worthy of your life's attention, because the impact at the end is so life altering, literally, that it will have been worthy of your life's attention.
I think we need things that are still challenging. Part of what I learned in my dance training is that to really be at the top, to be at the highest levels, you need to work hard. It is not easy, otherwise everyone else would do it.
When you show up to an audition and there's 400 people there and they're going to hire four people, you better be good. Otherwise everybody else in that room is good too, and you should go home and get better. And a thing I certainly don't see myself delegating in the future is this pursuit of high impact, challenging problems with solutions that I want to build with people who inspire me.
You Can Build Anything. The Question is Why
I teach a class at Stanford called CS334 Robots and Arts. It is the first class that's cross listed between computer science and theater arts and performance studies. And the invitation was that they could use AI and robots to make a creative statement with a political, philosophical or social underpinning.
You couldn't just have a robot do something. You actually needed to make an assertion with that project: I believe the world should be like this, or this is something that inspires me about the world, or here's something that I don't like and I would like to change it.
The very first day of the class, I was thinking about, here are all of these amazing students at Stanford, a nexus of technology, culture, art. But we're in this moment that is so special when it seems that everything is possible.
Researchers here at Stanford are 3D printing organs. You can take an autonomous car all around San Francisco or Austin or Phoenix, depending on where you are. We're also using large language models to potentially make sequences of DNA so we could build new life forms. So we're in this absolutely unbelievable moment right now.
Well, ostensibly this class is you're going to make a project, but really what this class is about is: when you are in a place of infinite opportunity to build so quickly, easily and seamlessly, the question that's interesting is not what I should build, but the question is, why am I building it? Do I want it to say or to do?
Catie Cuan during the interview
When you're in this space of infinite choice, you need to understand why you're making decisions, where your values are. We all have a finite amount of resources that we can dedicate to anything. We have our energy, our brilliance, our network, our relationships, our technical skills.
And we have this really precious thing which is our time. That is the only resource that you cannot scale, you cannot get back, that is equal for everyone. It does not matter who you are, where you are, your time is tremendously valuable.
The dedication of that time, it's one of the biggest choices that you make in your life. And we need to know why we make choices, because it's a reflection of the spending of this really precious resource, which is our time.
And knowing why you're building and making what you're doing, what the statement is that you want it to make. So I think the why for my students and for everyone else, when it's authentic, when it's real, it will carry you through a lot.
If you know why you're doing it, it can create less resilience in these challenging circumstances where you need to be able to overcome and do better, be better, work harder, learn faster, go farther than you thought. The why carries you through a lot of that.
Join the 1.5M+ founders inbox to get the latest updates.
Explore more
"ChatGPT Moment" for Robotics Is Coming. The Real Problem Isn't Intelligence