DeepMind’s New AI Just Changed Science Forever
Two Minute Papers
0:00 I appeared on camera for an interview not so long ago.
0:03 And I was really surprised by how many of you
0:06 Fellow Scholars said that you would like to see more.
0:09 So first of all, thank you so much to all of you for the kind words.
0:13 Second, I thought let's try this and hope that you will enjoy it.
0:17 Dear Fellow Scholars, this is Two Minute Papers with Dr.
0:22 Károly Zsolnai-Fehér.
0:23 Look, it only took 1,000 episodes.
0:26 Now, I have an amazing paper for you
0:29 because scientists at DeepMind did something pretty insane.
0:33 Our question today is can an AI invent
0:37 something that is fundamentally new and pushes humanity forward?
0:42 Well, they said that their new AI agent
0:45 can actually do research and even write research papers.
0:49 Most of the core content anyway.
0:52 Is that insane?
0:54 Well…it’s not.
0:55 A lot of other people have tried it and the only
0:59 insane thing about it was how many poor papers they wrote.
1:03 But it turns out… there is levels to this game.
1:06 You see, I visited the research group that is behind this work last year.
1:11 I flew to Mountain View into this crazy lab,
1:14 and a grumpy guard didn’t even want to let me in first.
1:18 Crazy town.
1:19 So I was very surprised that they are
1:21 guarding these secrets and they take them very seriously.
1:25 What is even more surprising is that now they give
1:29 some of those secrets away to all of us for free.
1:33 Now that is insane!
1:34 More on that in a moment.
1:36 So I talked to these scientists, this was the research group of Quoc Le.
1:41 They are brilliant.
1:43 They wrote an AI that was able to do
1:46 a gold medal worthy performance on the mathematical olympiad.
1:49 This is serious business.
1:51 Then they released this technique, anyone who is made out of money bags
1:56 and pays for the Gemini Advanced can use it, it is called Deep Think.
2:01 And now, this AI is even better than that.
2:06 They call it Aletheia.
2:08 Now that, once again is insane.
2:11 Okay, so what does it do?
2:13 Well, it promises that it does research.
2:16 It solves novel problems.
2:18 This is something that could push humanity forward.
2:22 Now that is so much harder than the mathematical olympiad.
2:26 Why is that?
2:27 Well, in these contests,
2:29 you have a not that huge piece of core knowledge you are supposed to have,
2:34 and every problem can be guaranteed to be solved by those small set of tools.
2:40 Every problem is nice, shiny, and polished.
2:45 Tough, but polished.
2:47 You know what is not polished at all?
2:49 Real life problems.
2:50 With these open problems, we don’t even know if they are solvable at all.
2:55 Maybe they are impossible, or maybe possible, but not with our current tools.
3:00 That’s the point: no one knows.
3:03 When this technique is given a problem, the generator starts working on it,
3:07 creates a candidate solution,
3:09 and now here is one of the important parts of the paper.
3:14 The verifier.
3:14 This takes a look, and says, okay bro this is junk.
3:20 Start again.
3:21 This is essentially a filter.
3:23 You know, that’s actually good life advice.
3:26 Sometimes it’s good to have a filter,
3:29 so you don’t just shoot those hot takes out there into the ether.
3:33 Now every now and then, the solution looks pretty good,
3:37 and could maybe pass with a few modifications.
3:40 Then, it gets polished for another round of reviews, and so it goes.
3:47 Sounds simple…maybe even trivial right?
3:50 So what is so scientific about this?
3:53 Why doesn’t every system do that?
3:55 Well, that’s easier said than done.
3:57 In fact, it is almost impossible to pull off.
4:01 Why?
4:02 One, when the AI is doing something fundamentally new,
4:07 unfortunately, hallucinations still happen.
4:09 Yup.
4:09 It just makes stuff up.
4:12 Fake papers, fictitious authors, you name it.
4:15 All kinds of junk comes out.
4:17 Two, when you want to compute 1+1 or other simple things,
4:22 you have tons of training data about it out there.
4:26 You can verify that easily.
4:28 But if you want to do frontier research?
4:31 There is no training data on what we don't even know yet.
4:36 Of course there isn’t!
4:37 You are trying to invent things no one understands yet.
4:41 These two factors make it extremely difficult to get
4:44 an AI to do something fundamentally new and useful.
4:47 So how did they pull it off?
4:51 With three key steps.
4:53 First, Alethia does not use this formal
4:56 rigid math language to check its own proofs.
4:59 It uses natural English language.
5:01 That is notoriously hard, because when the AI checks its own writing,
5:07 it just blindly agrees with it.
5:10 We humans do that too!
5:12 Now here, the researchers found a way
5:14 to separate the thinking part from the answer part.
5:18 So the messy train of thought is hidden from the verifier,
5:23 it cannot trick itself into just blindly agreeing with itself.
5:27 Brilliant.
5:28 Our brains would need something like that too.
5:32 Then, two they let the computer think longer.
5:37 That’s not new.
5:39 However, they added some optimizations to this, so much so that the model
5:43 they have now is just as smart as the one from 6 months ago.
5:48 But hold on to your papers Fellow Scholars, because yes,
5:52 same smarts, but it uses a 100 times less compute.
5:58 What!
5:59 Crazy.
5:59 They trained a much stronger base model
6:02 which made it more efficient at reasoning.
6:05 So this one, even without internet access,
6:09 beats the mathematical olympiad gold AI easily.
6:13 About 65% was improved to 95%.
6:16 Wow.
6:17 It went from a bit better than a coinfip to destroying
6:23 the tasks made for some of the best human minds.
6:25 All this in just a few months.
6:26 I am out of words.
6:27 Now three, they gave the AI the ability to search for stuff.
6:31 We are talking about Google after all.
6:34 Once again, that is easy.
6:36 However, getting the AI to read and combine techniques from dozens
6:41 and dozens of cutting-edge research papers without losing its mind.
6:45 Now that is hard.
6:47 You saw it earlier, this really happens!
6:50 They heavily trained this AI to be able to use
6:53 these tools and research works that are out there.
6:56 That was what finally stopped it from making up junk.
7:00 Okay, so how good is it?
7:02 First I saw that it solved a few of these Erdős problems.
7:07 It autonomously found the answer to 4 open
7:10 math puzzles left behind by a legendary Hungarian mathematician.
7:14 Is that insane?
7:16 I asked a mathematician friend.
7:18 He told me yeah, that’s pretty good,
7:21 but there are so many of these problems out there,
7:24 and not a ton of people work on them.
7:27 In other words, they are fairly easy,
7:29 they were just ignored by experts for years.
7:32 So not nearly as good as I thought.
7:34 But then, it stepped up its game
7:37 and wrote the core contents of a research paper.
7:40 On something new.
7:41 Note that the final paper is written up by a human scientist.
7:47 They had one paper on calculating constants in arithmetic geometry.
7:51 And then it helped human scientists write 4 other papers,
7:55 like finding new limits for interacting particles.
7:59 So how good are these research works?
8:02 Well, they are submitted for peer review and that’s going to take quite a while.
8:07 So, in the meantime, they had a bunch of math experts look at it,
8:11 many of them independent scientists.
8:13 They checked it for correctness and novelty, and it checks out man.
8:18 I think for the first time ever,
8:21 an AI created core parts of a research work that is new,
8:26 it has impact, it is useful.
8:29 That is…wow.
8:30 What a time to be alive!
8:33 So I told you there is levels to this game.
8:36 So where are we now?
8:38 Level 0 is negligible novelty work, it can do that.
8:42 Level 1 is somewhat novel work, it can do that too.
8:48 But now, it can help a person create publishable-level research.
8:52 That is incredible.
8:52 But wait, it can also do that autonomously.
8:54 An absolute game changer.
8:54 Levels 3 and 4, those are groundbreaking works, these are out of reach,
9:00 but I ask you Fellow Scholars, given the pace of progress, for how long?
9:05 For 6 more months?
9:07 And I think that is something that needs to be talked about more.
9:11 Research helping the people live a better life.
9:14 Love it.
9:15 And thank you so much to all of you
9:18 Fellow Scholars for watching us over the years.
9:20 We can only exist because of you Fellow Scholars.
9:24 I really hope that you enjoyed this.
9:26 It allows me to talk about papers where there is not a lot of visual content,
9:31 and I really wanted to share this with you.
9:33 Let me know in the comments if we should do more.