How DeepMind’s New AI Predicts What It Cannot See
Two Minute Papers
0:00 This absolutely incredible paper from the
0:02 Google DeepMind lab promises something that
0:05 sounds like science fiction.
0:07 Full 4 dimensional reconstruction of scenes.
0:10 Hmm.
0:11 Does this mean that
0:13 things disappear into another spatial dimension
0:16 like in this game called Miegakure?
0:19 No.
0:20 No,
0:21 because this game is in the works and it has been for more than 11 years now.
0:26 Wow.
0:26 Okay, I won’t say anything because I also worked
0:30 on this research paper called Gaussian Material
0:32 Synthesis that took me 3,000 work hours to finish.
0:35 And while I was working on it,
0:38 no papers appeared and people thought I was dead.
0:42 God, I haven’t even started the episode and we’ve gone off the rails already.
0:47 Okay, Károly, focus.
0:48 Okay, so what is this 4D thing?
0:51 Well, 3 spatial dimensions, and 1 dimension that
0:54 is time.
0:55 It’s not crazy wormholes, it’s worse!
0:58 It’s like building IKEA furniture,
1:00 but as you start tightening the screws, the cabinet is running away.
1:05 Okay, so what the heck is this crazy person talking about.
1:09 So in goes a video of a scene
1:12 of your choice.
1:13 And out comes a virtual version of it in the form of a point cloud.
1:18 However, the catch is that things are allowed to move around as they please.
1:23 And this is fantastic, I mean look at these highly dynamic judo scenes and all
1:28 kinds of craziness,
1:30 and it understands how these points are moving around over time.
1:34 I am always fascinated by the fact that an AI can look at a 2D photograph,
1:40 and understand the underlying spatial reality.
1:42 This is just a bunch of numbers for them,
1:46 yet they understand what is close and what is far away.
1:50 Crazy.
1:51 We humans are good at that,
1:53 but we have a brain that evolved for that for millions of years.
1:57 And this is just a bunch
1:59 of sand that learned to think.
2:01 So that is already amazing.
2:03 But it gets better.
2:05 DeepMind says it could have unlimited applications, yes,
2:10 unlimited power!
2:12 Woo-hoo!
2:12 Károly.
2:13 Ok, ok.
2:15 Now performing this is really tough.
2:18 Previous techniques could do this kind
2:20 of 4D reconstruction, but you needed a bunch of specialized models for it.
2:25 You’d have one AI for depth, another for motion, and a third for camera angles.
2:31 And then you have to glue
2:33 all of these together into an abomination.
2:36 Using the abomination requires a technique
2:39 called test-time optimization.
2:41 Yes.
2:42 Here, your computer sits there sweating for minutes,
2:46 trying to make the different models agree with
2:49 each other so the geometry doesn't fall apart.
2:52 Now this new technique doesn’t do that.
2:55 This is called D4RT, if you want to sound cool,
2:59 pronounce it as dart.
3:01 Now this one uses one AI technique.
3:04 Just one transformer.
3:05 Everything that you see here in the middle is just part of one thing.
3:10 And this one thing can handle depth,
3:12 motion,
3:12 and camera pose simultaneously without needing them to talk to each other.
3:18 But it gets better.
3:20 A lot better.
3:21 It can even track through occlusion.
3:23 It is able to guess
3:25 where these points are, even if it doesn’t see them.
3:29 How on Earth is that possible?
3:31 Well, these points we have seen before, and will see again, so it is able
3:36 to make an educated guess as to where they are, even if it doesn’t see them.
3:42 Crazy.
3:43 And it can reconstruct massive scenes by
3:46 just briefly looking through them.
3:49 Absolutely incredible.
3:50 Now hold on to your papers Fellow Scholars,
3:53 because as a result, it is incredibly fast.
3:56 I mean, wow.
3:57 Look at how it compares to previous techniques.
4:00 Depending on what you compare to,
4:02 it is up to 300 times faster.
4:05 That is mind blowing.
4:07 I’ll tell you in a moment how it works.
4:11 Now, wait wait wait.
4:12 Hold the phone.
4:14 we can represent scenes in other ways too,
4:17 not just with point clouds.
4:19 Most games and animation movies use 3D mesh geometry,
4:22 and Gaussian Splats are also the new rage.
4:26 How does this relate to those?
4:28 It is better and also worse in 3 ways.
4:32 First, it excels at handling motion.
4:34 While meshes and splats often struggle with ghosting,
4:38 leaving behind artifacts as objects move, D4RT treats movement as a core part of
4:44 the math.
4:44 Second, it is up to 300x faster than previous methods.
4:49 It skips the slow,
4:51 iterative optimization loops that Gaussian splats usually require.
4:55 Third, the model recovers depth, tracks, and camera parameters simultaneously.
5:00 These are incredibly appealing.
5:03 However, let’s not overstate things here.
5:05 Now come the bad news.
5:07 3 things it is not so good
5:09 at.
5:10 Because it outputs a point cloud, the data is let’s say unintelligent.
5:15 It’s just a bunch
5:16 of dots.
5:17 You can't 3D print it or use it
5:20 for physics collisions without an extra meshing step.
5:23 It is also not meant to look pretty.
5:26 Meshes and Gaussian Splats remain the
5:29 kings of photorealistic reflections while D4RT
5:32 focuses strictly on geometric accuracy.
5:35 Finally, it is worse for editing,
5:38 because without the structured faces of a mesh,
5:41 you can't exactly hop into Blender and sculpt it like digital clay.
5:46 Okay, so how is all this incredible work possible?
5:49 How do we assemble that cabinet that wants to
5:52 run away?
5:53 Dear Fellow Scholars, this is Two Minute Papers with Dr.
5:57 Károly Zsolnai-Fehér.
5:57 First, the encoder.
5:59 This is a master carpenter.
6:01 This looks at the scene and
6:03 tries to understand the past and the present of the furniture.
6:08 Understand what it’s about.
6:10 This they call a global scene representation.
6:13 Then, we get the decoder.
6:15 These are the magic elves.
6:17 Now let’s build.
6:18 Here comes the genius
6:20 part.
6:20 Instead of trying to build the whole cabinet at once, which is heavy and slow.
6:27 Yes, we all know that from building IKEA furniture.
6:31 How the heck can this box have 100 screws?
6:34 No one knows.
6:36 Okay, so the carpenter just points to a spot and yells at a tiny
6:41 elf: “Hey YOU!
6:42 Yes, you!
6:43 Where is this specific screw at timestamp 10?"
6:46 The elf, which is the query grabs the info and zaps the screw into existence.
6:52 Now here comes the genius part.
6:55 Elves don’t need to talk to each other.
6:58 Oh yes, finally!
6:59 So because of that, you can have 10 elves or 1,000,000 elves doesn’t matter.
7:06 Yes, the technique is completely parallelizable!
7:08 That is the other reason why it is so bloody fast.
7:13 And here is the kicker.
7:15 The decoder, so the elves see in a way that is a bit blurry.
7:19 They have terrible eyesight,
7:21 so the objects they are working on become a bit blurry.
7:26 So scientists say, let’s
7:27 give them magic glasses.
7:29 How?
7:30 Well, by feeding the technique the original, high-resolution video
7:34 pixels back into the decoder.
7:36 So this is what they saw before, and this is what they see now.
7:42 That is insane,
7:43 because now it can reconstruct details finer than the AI's own internal brain!
7:49 But I haven’t explained the part where the cabinet wants to run away.
7:53 How do we handle that?
7:55 Well, in a normal 3D scan, if the camera can't see the leg of the cabinet,
8:00 the computer just gives up.
8:02 Incomplete information and moving things cannot be
8:05 handled well.
8:06 They just leave a giant hole in your geometry.
8:10 Total disaster.
8:10 But remember, our master carpenter is not looking at just one photo.
8:15 He has watched the entire video tape from start to finish.
8:19 He has seen the past, and the present.
8:22 So when the cabinet leg disappears behind the sofa, the elf cries out, "Master!
8:28 The screw is gone!
8:29 I cannot build what I cannot see!" I do not know why an elf has this voice.
8:35 Now, the wise carpenter smiles and says: "Relax.
8:38 I saw that screw five seconds ago,
8:41 and I see it pop out the other side five seconds later.
8:45 Based on that, right now, it is hiding...
8:49 exactly here!” And boom!
8:50 The elf is now suddenly able to assemble
8:54 the cabinet.
8:55 In other words,
8:56 this is how it tracks through occlusion and disappearing information.
9:01 Now surprisingly, there is more to learn here.
9:04 Listen.
9:04 The elves build the scene
9:06 300x faster because they do not talk to each other.
9:11 That is excellent life
9:12 advice.
9:13 Sometimes collaboration has a tax.
9:15 Sometimes instead you need to create a few
9:19 hours of zero-communication deep work blocks where you are unreachable.
9:24 Whenever I do that,
9:26 I am often surprised by how much I can get done in little time.
9:30 This is a collaboration between the wizards at Google DeepMind,
9:34 University College London, and University of Oxford.
9:37 These are the people inventing the power tools of the
9:41 future and giving it away for all of us for free.
9:45 Thank you so much!
9:46 What a time to be alive!
9:48 So, here you go.
9:50 A glimpse of the future and how digital worlds could be created soon.
9:54 A really advanced paper described in simple words anyone can understand.
9:59 If you appreciate that,
10:01 make sure to subscribe, hit the bell and leave a kind comment.
10:05 So you’ll get more
10:06 videos like this.
10:07 Don’t worry about it, we are all paper addicts here.