How DeepMind’s New AI Predicts What It Cannot See

How DeepMind’s New AI Predicts What It Cannot See

Two Minute Papers

0:00 This absolutely incredible paper from the

0:02 Google DeepMind lab promises something that

0:05 sounds like science fiction.

0:07 Full 4 dimensional reconstruction of scenes.

0:10 Hmm.

0:11 Does this mean that

0:13 things disappear into another spatial dimension

0:16 like in this game called Miegakure?

0:19 No.

0:20 No,

0:21 because this game is in the works and it has been for more than 11 years now.

0:26 Wow.

0:26 Okay, I won’t say anything because I also worked

0:30 on this research paper called Gaussian Material

0:32 Synthesis that took me 3,000 work hours to finish.

0:35 And while I was working on it,

0:38 no papers appeared and people thought I was dead.

0:42 God, I haven’t even started the episode and we’ve gone off the rails already.

0:47 Okay, Károly, focus.

0:48 Okay, so what is this 4D thing?

0:51 Well, 3 spatial dimensions, and 1 dimension that

0:54 is time.

0:55 It’s not crazy wormholes, it’s worse!

0:58 It’s like building IKEA furniture,

1:00 but as you start tightening the screws, the cabinet is running away.

1:05 Okay, so what the heck is this crazy person talking about.

1:09 So in goes a video of a scene

1:12 of your choice.

1:13 And out comes a virtual version of it in the form of a point cloud.

1:18 However, the catch is that things are allowed to move around as they please.

1:23 And this is fantastic, I mean look at these highly dynamic judo scenes and all

1:28 kinds of craziness,

1:30 and it understands how these points are moving around over time.

1:34 I am always fascinated by the fact that an AI can look at a 2D photograph,

1:40 and understand the underlying spatial reality.

1:42 This is just a bunch of numbers for them,

1:46 yet they understand what is close and what is far away.

1:50 Crazy.

1:51 We humans are good at that,

1:53 but we have a brain that evolved for that for millions of years.

1:57 And this is just a bunch

1:59 of sand that learned to think.

2:01 So that is already amazing.

2:03 But it gets better.

2:05 DeepMind says it could have unlimited applications, yes,

2:10 unlimited power!

2:12 Woo-hoo!

2:12 Károly.

2:13 Ok, ok.

2:15 Now performing this is really tough.

2:18 Previous techniques could do this kind

2:20 of 4D reconstruction, but you needed a bunch of specialized models for it.

2:25 You’d have one AI for depth, another for motion, and a third for camera angles.

2:31 And then you have to glue

2:33 all of these together into an abomination.

2:36 Using the abomination requires a technique

2:39 called test-time optimization.

2:41 Yes.

2:42 Here, your computer sits there sweating for minutes,

2:46 trying to make the different models agree with

2:49 each other so the geometry doesn't fall apart.

2:52 Now this new technique doesn’t do that.

2:55 This is called D4RT, if you want to sound cool,

2:59 pronounce it as dart.

3:01 Now this one uses one AI technique.

3:04 Just one transformer.

3:05 Everything that you see here in the middle is just part of one thing.

3:10 And this one thing can handle depth,

3:12 motion,

3:12 and camera pose simultaneously without needing them to talk to each other.

3:18 But it gets better.

3:20 A lot better.

3:21 It can even track through occlusion.

3:23 It is able to guess

3:25 where these points are, even if it doesn’t see them.

3:29 How on Earth is that possible?

3:31 Well, these points we have seen before, and will see again, so it is able

3:36 to make an educated guess as to where they are, even if it doesn’t see them.

3:42 Crazy.

3:43 And it can reconstruct massive scenes by

3:46 just briefly looking through them.

3:49 Absolutely incredible.

3:50 Now hold on to your papers Fellow Scholars,

3:53 because as a result, it is incredibly fast.

3:56 I mean, wow.

3:57 Look at how it compares to previous techniques.

4:00 Depending on what you compare to,

4:02 it is up to 300 times faster.

4:05 That is mind blowing.

4:07 I’ll tell you in a moment how it works.

4:11 Now, wait wait wait.

4:12 Hold the phone.

4:14 we can represent scenes in other ways too,

4:17 not just with point clouds.

4:19 Most games and animation movies use 3D mesh geometry,

4:22 and Gaussian Splats are also the new rage.

4:26 How does this relate to those?

4:28 It is better and also worse in 3 ways.

4:32 First, it excels at handling motion.

4:34 While meshes and splats often struggle with ghosting,

4:38 leaving behind artifacts as objects move, D4RT treats movement as a core part of

4:44 the math.

4:44 Second, it is up to 300x faster than previous methods.

4:49 It skips the slow,

4:51 iterative optimization loops that Gaussian splats usually require.

4:55 Third, the model recovers depth, tracks, and camera parameters simultaneously.

5:00 These are incredibly appealing.

5:03 However, let’s not overstate things here.

5:05 Now come the bad news.

5:07 3 things it is not so good

5:09 at.

5:10 Because it outputs a point cloud, the data is let’s say unintelligent.

5:15 It’s just a bunch

5:16 of dots.

5:17 You can't 3D print it or use it

5:20 for physics collisions without an extra meshing step.

5:23 It is also not meant to look pretty.

5:26 Meshes and Gaussian Splats remain the

5:29 kings of photorealistic reflections while D4RT

5:32 focuses strictly on geometric accuracy.

5:35 Finally, it is worse for editing,

5:38 because without the structured faces of a mesh,

5:41 you can't exactly hop into Blender and sculpt it like digital clay.

5:46 Okay, so how is all this incredible work possible?

5:49 How do we assemble that cabinet that wants to

5:52 run away?

5:53 Dear Fellow Scholars, this is Two Minute Papers with Dr.

5:57 Károly Zsolnai-Fehér.

5:57 First, the encoder.

5:59 This is a master carpenter.

6:01 This looks at the scene and

6:03 tries to understand the past and the present of the furniture.

6:08 Understand what it’s about.

6:10 This they call a global scene representation.

6:13 Then, we get the decoder.

6:15 These are the magic elves.

6:17 Now let’s build.

6:18 Here comes the genius

6:20 part.

6:20 Instead of trying to build the whole cabinet at once, which is heavy and slow.

6:27 Yes, we all know that from building IKEA furniture.

6:31 How the heck can this box have 100 screws?

6:34 No one knows.

6:36 Okay, so the carpenter just points to a spot and yells at a tiny

6:41 elf: “Hey YOU!

6:42 Yes, you!

6:43 Where is this specific screw at timestamp 10?"

6:46 The elf, which is the query grabs the info and zaps the screw into existence.

6:52 Now here comes the genius part.

6:55 Elves don’t need to talk to each other.

6:58 Oh yes, finally!

6:59 So because of that, you can have 10 elves or 1,000,000 elves doesn’t matter.

7:06 Yes, the technique is completely parallelizable!

7:08 That is the other reason why it is so bloody fast.

7:13 And here is the kicker.

7:15 The decoder, so the elves see in a way that is a bit blurry.

7:19 They have terrible eyesight,

7:21 so the objects they are working on become a bit blurry.

7:26 So scientists say, let’s

7:27 give them magic glasses.

7:29 How?

7:30 Well, by feeding the technique the original, high-resolution video

7:34 pixels back into the decoder.

7:36 So this is what they saw before, and this is what they see now.

7:42 That is insane,

7:43 because now it can reconstruct details finer than the AI's own internal brain!

7:49 But I haven’t explained the part where the cabinet wants to run away.

7:53 How do we handle that?

7:55 Well, in a normal 3D scan, if the camera can't see the leg of the cabinet,

8:00 the computer just gives up.

8:02 Incomplete information and moving things cannot be

8:05 handled well.

8:06 They just leave a giant hole in your geometry.

8:10 Total disaster.

8:10 But remember, our master carpenter is not looking at just one photo.

8:15 He has watched the entire video tape from start to finish.

8:19 He has seen the past, and the present.

8:22 So when the cabinet leg disappears behind the sofa, the elf cries out, "Master!

8:28 The screw is gone!

8:29 I cannot build what I cannot see!" I do not know why an elf has this voice.

8:35 Now, the wise carpenter smiles and says: "Relax.

8:38 I saw that screw five seconds ago,

8:41 and I see it pop out the other side five seconds later.

8:45 Based on that, right now, it is hiding...

8:49 exactly here!” And boom!

8:50 The elf is now suddenly able to assemble

8:54 the cabinet.

8:55 In other words,

8:56 this is how it tracks through occlusion and disappearing information.

9:01 Now surprisingly, there is more to learn here.

9:04 Listen.

9:04 The elves build the scene

9:06 300x faster because they do not talk to each other.

9:11 That is excellent life

9:12 advice.

9:13 Sometimes collaboration has a tax.

9:15 Sometimes instead you need to create a few

9:19 hours of zero-communication deep work blocks where you are unreachable.

9:24 Whenever I do that,

9:26 I am often surprised by how much I can get done in little time.

9:30 This is a collaboration between the wizards at Google DeepMind,

9:34 University College London, and University of Oxford.

9:37 These are the people inventing the power tools of the

9:41 future and giving it away for all of us for free.

9:45 Thank you so much!

9:46 What a time to be alive!

9:48 So, here you go.

9:50 A glimpse of the future and how digital worlds could be created soon.

9:54 A really advanced paper described in simple words anyone can understand.

9:59 If you appreciate that,

10:01 make sure to subscribe, hit the bell and leave a kind comment.

10:05 So you’ll get more

10:06 videos like this.

10:07 Don’t worry about it, we are all paper addicts here.

Study with Looplines Download Captions Watch on YouTube