NVIDIA's New AI Broke My Brain

NVIDIA's New AI Broke My Brain

Two Minute Papers

0:00 Let’s see what is going on here.

0:03 This is me around 9am.

0:05 A bit wobbly, steps are unsure, yup, that checks out.

0:10 Now then, give me my fake badge.

0:13 Thank you sir.

0:14 Hehehe, no one noticed.

0:16 Now let’s proceed to the next step of my mastermind plans.

0:22 Let’s eat all their food.

0:25 Wait, they noticed.

0:27 Proceed to the next step.

0:30 What was that?

0:32 Oh yes, run!

0:34 Now, jokes aside, look at that.

0:35 Sign up for this one baby.

0:37 Oh yes, please mow my lawn.

0:40 That is excellent.

0:41 Rake the leaves!

0:43 Perfect.

0:44 Hey, don’t slack off, that’s my job!

0:48 Okay, so what is going on here.

0:51 Let’s start with the good news,

0:54 this is a new teleoperated robot controller and more.

0:58 They call it Sonic.

1:00 Now the work here is not the robot, but the software controlling it.

1:07 At least in this footage, watch until the end and you might get surprised.

1:13 This means there is a human performing these movements,

1:17 and the robot is able to understand these motions,

1:20 and then translate them to a bunch of joint positions in 3D space.

1:26 It’s kind of insane that this is possible.

1:30 But it will just get better and better as we continue the video.

1:34 So, before you ask, yes it can do kung fu.

1:39 Provided that you can do kung fu.

1:42 It understands whole body movement,

1:44 so you can get it to crawl into some space you don’t want to go to.

1:49 And that is super useful, people are already using robots for that.

1:55 Why?

1:56 Well, chiefly, for exploring under explored and dangerous areas.

2:00 This means tons of useful applications, for instance,

2:04 a variant of this could help save humans stuck under rubble,

2:09 or perhaps later, even explore other planets without putting humans at risk.

2:15 But that’s still nothing.

2:17 Because this is a multimodal system.

2:20 Meaning that the input can be almost anything.

2:23 So, you say that I don’t have to pretend

2:27 to mow the lawn to actually mow the lawn, because where is the fun in that?

2:34 Well, just tell it to do that.

2:38 Can you?

2:39 Well, currently, for simpler tasks,

2:41 like moving around or behaving like a monkey, yes you can!

2:47 Absolutely incredible.

2:48 And I love how expressive it is.

2:50 You can ask it to walk happily, stealthily, or like an injured person.

2:56 And you know, just the fact that it is stable and does not fall is remarkable.

3:03 Previously, even in simple characters in simulated worlds,

3:06 you needed thousands and thousands of tries to teach

3:10 them to just be able to walk without falling.

3:15 And now, this, is a huge leap forward.

3:19 Wow.

3:20 But it gets better, we said multimodal.

3:23 Yup, that means that the input can also be music.

3:28 I’ll show you the dancing, but not the music because of Youtube reasons,

3:38 but I put a link in the description where you can check it out.

3:41 And we haven’t even talked about the most insane part of the whole thing.

3:48 Now hold on to your papers Fellow Scholars,

3:52 because this runs with about 42 million parameters.

3:57 That is a neural network so simple,

4:01 it can run so easily on your phone it barely notices it.

4:05 It may even run on your toaster these days.

4:10 That size is absolutely nothing.

4:14 This is an incredible achivement.

4:19 Okay, but how?

4:20 How is that even possible?

4:22 Dear Fellow Scholars, this is Two Minute Papers with Dr.

4:26 Károly Zsolnai-Fehér.

4:27 Well, first, it looked at 100 million frames of human

4:32 motion to understand what we do and how we do it.

4:36 The incredible thing is that this system

4:38 does not require human-made action labels,

4:40 so we don’t have to explain our movements.

4:43 It just watches the raw motions and figures out

4:47 how to transition between tasks without any unnatural pauses!

4:52 So then, your multi-modal input goes in, a video of you,

4:55 your voice, music, or just text.

4:57 A motion generator turns these into human motion,

5:00 and the human encoder processes it into a latent space,

5:05 and then a quantizer converts it to universal tokens.

5:10 Once again, universal tokens, that is key, you’ll see a bit later.

5:16 Then, the decoder translates these tokens into motor commands.

5:20 But there is a big problem.

5:23 Learning to convert one to the other is super hard.

5:29 First of all, robots do not work like humans,

5:33 that is one of the fundamental challenges.

5:35 So if the user commands you to turn around, it should be turning around.

5:42 Okay, sure.

5:43 But how fast exactly?

5:44 You don’t want to try to turn 180 degrees too quickly,

5:49 because you would fall apart.

5:51 To solve this, in their research paper,

5:54 they propose what they call a root trajectory spring model.

5:58 This dampens sudden, quick user commands so the robot does not get injured.

6:04 Yes, robots can get injured too, which is kind of hilarious.

6:10 Now there is an exponential term as a function of time.

6:15 What is that?

6:16 That is a physical brake.

6:18 As time increases, this term rapidly shrinks to 0,

6:23 which forces the whole mathematical expression to decay smoothly.

6:27 This serves two goals: one, the robot does not injure itself and two,

6:34 it will settle at a target position without oscillating back and forth forever.

6:40 Nice.

6:41 Now, do the dampening too much, and of course,

6:44 you’ll get a little slug that can’t get anything done,

6:48 so it’s really tough to do well.

6:50 Well done folks.

6:51 Now, all this took 128 GPUs and 3 days to train.

6:57 That is expensive.

6:59 But here’s the key, after the training is done,

7:03 the final product is so lightweight,

7:05 we don’t need this kind of hardware to run it at all.

7:09 In fact, all of the models showcased in these videos

7:12 will be given to all of us for free, forever.

7:17 They run on your phone, easy-peasy.

7:20 That is incredible.

7:21 Open research for the benefit of humanity.

7:24 Love it, thank you so much.

7:27 This project is led by professor Zhu and Jim Fan, who I love dearly.

7:33 Jim started the humanoid robots lab at NVIDIA just 2 years ago,

7:38 and they are raining research papers on us, breakthrough after breakthrough.

7:44 Insanity.

7:45 And to compress all this human movement knowledge down into a tiny little AI

7:51 controller that can be used by any of us is simply a stunning achievement.

7:56 It turns out, training a good AI requires coding good thinking into a machine.

8:03 But, surprisingly, we ourselves can also learn a lot

8:07 of good life advice from this kind of thinking too.

8:10 For instance, the model compresses a messy,

8:13 diverse soup of inputs into a kind of pure, abstract token.

8:17 You know, in life, when asking other people for advice,

8:22 you will inevitably hear everything, and its opposite too.

8:26 That is also a big soup of inputs.

8:29 But try to look at all of them, side by side,

8:33 and you’ll find that they often share an underlying truth.

8:36 This works, as is showcased by this incredible project too.

8:41 And note that this work is not the end of anything, this is just a start.

8:47 An early work at a nascent area.

8:50 Two more papers down the line, and I really hope this is going

8:54 to start folding my laundry and cooking my lunch.

8:58 That would be amazing.

8:59 What a time to be alive!

9:02 And this is not some proprietary nonsense,

9:05 this is open knowledge and open just dropped.

9:27 If you are interested in hearing more hopefully soon,

9:30 subscribe and hit the bell.

Study with Looplines Download Captions Watch on YouTube