DeepMind’s New AI Just Changed Science Forever

DeepMind’s New AI Just Changed Science Forever

Two Minute Papers

0:00 I appeared on camera for an interview not so long ago.

0:03 And I was really surprised by how many of you

0:06 Fellow Scholars said that you would like to see more.

0:09 So first of all, thank you so much to all of you for the kind words.

0:13 Second, I thought let's try this and hope that you will enjoy it.

0:17 Dear Fellow Scholars, this is Two Minute Papers with Dr.

0:22 Károly Zsolnai-Fehér.

0:23 Look, it only took 1,000 episodes.

0:26 Now, I have an amazing paper for you

0:29 because scientists at DeepMind did something pretty insane.

0:33 Our question today is can an AI invent

0:37 something that is fundamentally new and pushes humanity forward?

0:42 Well, they said that their new AI agent

0:45 can actually do research and even write research papers.

0:49 Most of the core content anyway.

0:52 Is that insane?

0:54 Well…it’s not.

0:55 A lot of other people have tried it and the only

0:59 insane thing about it was how many poor papers they wrote.

1:03 But it turns out… there is levels to this game.

1:06 You see, I visited the research group that is behind this work last year.

1:11 I flew to Mountain View into this crazy lab,

1:14 and a grumpy guard didn’t even want to let me in first.

1:18 Crazy town.

1:19 So I was very surprised that they are

1:21 guarding these secrets and they take them very seriously.

1:25 What is even more surprising is that now they give

1:29 some of those secrets away to all of us for free.

1:33 Now that is insane!

1:34 More on that in a moment.

1:36 So I talked to these scientists, this was the research group of Quoc Le.

1:41 They are brilliant.

1:43 They wrote an AI that was able to do

1:46 a gold medal worthy performance on the mathematical olympiad.

1:49 This is serious business.

1:51 Then they released this technique, anyone who is made out of money bags

1:56 and pays for the Gemini Advanced can use it, it is called Deep Think.

2:01 And now, this AI is even better than that.

2:06 They call it Aletheia.

2:08 Now that, once again is insane.

2:11 Okay, so what does it do?

2:13 Well, it promises that it does research.

2:16 It solves novel problems.

2:18 This is something that could push humanity forward.

2:22 Now that is so much harder than the mathematical olympiad.

2:26 Why is that?

2:27 Well, in these contests,

2:29 you have a not that huge piece of core knowledge you are supposed to have,

2:34 and every problem can be guaranteed to be solved by those small set of tools.

2:40 Every problem is nice, shiny, and polished.

2:45 Tough, but polished.

2:47 You know what is not polished at all?

2:49 Real life problems.

2:50 With these open problems, we don’t even know if they are solvable at all.

2:55 Maybe they are impossible, or maybe possible, but not with our current tools.

3:00 That’s the point: no one knows.

3:03 When this technique is given a problem, the generator starts working on it,

3:07 creates a candidate solution,

3:09 and now here is one of the important parts of the paper.

3:14 The verifier.

3:14 This takes a look, and says, okay bro this is junk.

3:20 Start again.

3:21 This is essentially a filter.

3:23 You know, that’s actually good life advice.

3:26 Sometimes it’s good to have a filter,

3:29 so you don’t just shoot those hot takes out there into the ether.

3:33 Now every now and then, the solution looks pretty good,

3:37 and could maybe pass with a few modifications.

3:40 Then, it gets polished for another round of reviews, and so it goes.

3:47 Sounds simple…maybe even trivial right?

3:50 So what is so scientific about this?

3:53 Why doesn’t every system do that?

3:55 Well, that’s easier said than done.

3:57 In fact, it is almost impossible to pull off.

4:01 Why?

4:02 One, when the AI is doing something fundamentally new,

4:07 unfortunately, hallucinations still happen.

4:09 Yup.

4:09 It just makes stuff up.

4:12 Fake papers, fictitious authors, you name it.

4:15 All kinds of junk comes out.

4:17 Two, when you want to compute 1+1 or other simple things,

4:22 you have tons of training data about it out there.

4:26 You can verify that easily.

4:28 But if you want to do frontier research?

4:31 There is no training data on what we don't even know yet.

4:36 Of course there isn’t!

4:37 You are trying to invent things no one understands yet.

4:41 These two factors make it extremely difficult to get

4:44 an AI to do something fundamentally new and useful.

4:47 So how did they pull it off?

4:51 With three key steps.

4:53 First, Alethia does not use this formal

4:56 rigid math language to check its own proofs.

4:59 It uses natural English language.

5:01 That is notoriously hard, because when the AI checks its own writing,

5:07 it just blindly agrees with it.

5:10 We humans do that too!

5:12 Now here, the researchers found a way

5:14 to separate the thinking part from the answer part.

5:18 So the messy train of thought is hidden from the verifier,

5:23 it cannot trick itself into just blindly agreeing with itself.

5:27 Brilliant.

5:28 Our brains would need something like that too.

5:32 Then, two they let the computer think longer.

5:37 That’s not new.

5:39 However, they added some optimizations to this, so much so that the model

5:43 they have now is just as smart as the one from 6 months ago.

5:48 But hold on to your papers Fellow Scholars, because yes,

5:52 same smarts, but it uses a 100 times less compute.

5:58 What!

5:59 Crazy.

5:59 They trained a much stronger base model

6:02 which made it more efficient at reasoning.

6:05 So this one, even without internet access,

6:09 beats the mathematical olympiad gold AI easily.

6:13 About 65% was improved to 95%.

6:16 Wow.

6:17 It went from a bit better than a coinfip to destroying

6:23 the tasks made for some of the best human minds.

6:25 All this in just a few months.

6:26 I am out of words.

6:27 Now three, they gave the AI the ability to search for stuff.

6:31 We are talking about Google after all.

6:34 Once again, that is easy.

6:36 However, getting the AI to read and combine techniques from dozens

6:41 and dozens of cutting-edge research papers without losing its mind.

6:45 Now that is hard.

6:47 You saw it earlier, this really happens!

6:50 They heavily trained this AI to be able to use

6:53 these tools and research works that are out there.

6:56 That was what finally stopped it from making up junk.

7:00 Okay, so how good is it?

7:02 First I saw that it solved a few of these Erdős problems.

7:07 It autonomously found the answer to 4 open

7:10 math puzzles left behind by a legendary Hungarian mathematician.

7:14 Is that insane?

7:16 I asked a mathematician friend.

7:18 He told me yeah, that’s pretty good,

7:21 but there are so many of these problems out there,

7:24 and not a ton of people work on them.

7:27 In other words, they are fairly easy,

7:29 they were just ignored by experts for years.

7:32 So not nearly as good as I thought.

7:34 But then, it stepped up its game

7:37 and wrote the core contents of a research paper.

7:40 On something new.

7:41 Note that the final paper is written up by a human scientist.

7:47 They had one paper on calculating constants in arithmetic geometry.

7:51 And then it helped human scientists write 4 other papers,

7:55 like finding new limits for interacting particles.

7:59 So how good are these research works?

8:02 Well, they are submitted for peer review and that’s going to take quite a while.

8:07 So, in the meantime, they had a bunch of math experts look at it,

8:11 many of them independent scientists.

8:13 They checked it for correctness and novelty, and it checks out man.

8:18 I think for the first time ever,

8:21 an AI created core parts of a research work that is new,

8:26 it has impact, it is useful.

8:29 That is…wow.

8:30 What a time to be alive!

8:33 So I told you there is levels to this game.

8:36 So where are we now?

8:38 Level 0 is negligible novelty work, it can do that.

8:42 Level 1 is somewhat novel work, it can do that too.

8:48 But now, it can help a person create publishable-level research.

8:52 That is incredible.

8:52 But wait, it can also do that autonomously.

8:54 An absolute game changer.

8:54 Levels 3 and 4, those are groundbreaking works, these are out of reach,

9:00 but I ask you Fellow Scholars, given the pace of progress, for how long?

9:05 For 6 more months?

9:07 And I think that is something that needs to be talked about more.

9:11 Research helping the people live a better life.

9:14 Love it.

9:15 And thank you so much to all of you

9:18 Fellow Scholars for watching us over the years.

9:20 We can only exist because of you Fellow Scholars.

9:24 I really hope that you enjoyed this.

9:26 It allows me to talk about papers where there is not a lot of visual content,

9:31 and I really wanted to share this with you.

9:33 Let me know in the comments if we should do more.

Study with Looplines Download Captions Watch on YouTube