Gemini Exponential, Demis Hassabis' ‘Proto-AGI’ coming, but …

Gemini Exponential, Demis Hassabis' ‘Proto-AGI’ coming, but …

AI Explained

0:00 In the last 48 hours, we have had two major model releases and about

0:04 10 hours worth of interviews from top leaders about them.

0:09 The insights of which I will try to condense into just 15 minutes or so.

0:14 Because Gemini 3 flash is Google's attempt at finally convincing you

0:20 to switch from chatbt or claude and the results look incredible.

0:26 I'll go through them in a moment,

0:27 but we have two co-founders of Google DeepMind.

0:29 Both seeing the LLM paradigm continuing on this exponential until

0:34 a sketched out protoi model arrives in not too long.

0:38 However, there are some problems with that vision and one

0:42 result in particular I don't want you to miss.

0:44 So, let's get started.

0:46 Here are some of the raw numbers and bear in mind

0:48 that the flash version of Gemini is the quick version,

0:52 the one that can answer almost instantly.

0:54 You guys will know that all companies have a pro

0:57 version of their models that typically take much much longer,

1:00 minutes often to answer a question.

1:01 I want you to notice the comparison with the model released 2 days ago,

1:04 Gemini 3 Flash, with the state-of-the-art model

1:08 as of June of this year, Gemini 2.5 Pro.

1:11 Whether we're talking about academic reasoning, visual reasoning,

1:14 scientific knowledge, coding, mathematics, the results aren't even that close.

1:18 And this is for the dramatically quicker model.

1:20 For example, even without access to tools,

1:23 the new Gemini 3 Flash roughly halves the error

1:26 rate in one very difficult mathematics benchmark, AIM.

1:29 Again, this is comparing Summer's Gemini 2.5 Pro at 88%

1:33 to two days ago's Gemini 3 Flash at 95.2%.

1:38 In fact, in almost any domain you can

1:40 point to, from table and chart analysis, video analysis,

1:44 or going off and being an agent,

1:45 Gemini 3 Flash exceeds the previous huge model performance from the summer.

1:50 You can, of course, optimize models for one particular set of benchmarks.

1:54 And we learned just this morning that Google did indeed apply

1:57 a special type of post-training

1:59 to optimize performance for software engineering.

2:01 For those who code, you may be somewhat incredulous

2:04 to see Gemini 3 Flash outperforming Gemini 3 Pro,

2:08 the heavier model released just a few weeks ago.

2:11 It would be very easy to get carried away with those results

2:14 and say Chat GBT is doomed for consumers and Gemini,

2:18 as Jim Kramer points out, is growing much faster.

2:20 Given that Jim Kramer is wrong about so much,

2:23 the head of applied research at OpenAI took this as a great sign for chatbt.

2:28 But the reality is always more complex than the headlines make

2:31 it seem because Gemini 3 Flash is indeed a great model,

2:34 but it does have a key weakness.

2:36 And if ChhatPut was dying, that's definitely news to investors who keep

2:40 valuing OpenAI higher and higher and higher.

2:43 Now before we get to the proto AGI sketched

2:45 out by Deis Sarvis and another co-founder of DeepMind,

2:49 I just want to spend a moment more on Gemini

2:52 3 Flash because there is a secret about AI

2:54 model releases that I want all of you guys

2:56 to be aware of when you see a new model announced.

2:59 The secret is that models are rarely punished for incorrect answers.

3:04 They are not incentivized to say I don't know.

3:07 So companies like OpenAI and Google Deep Mind

3:09 and Anthropic are heavily incentivized to instruct their models.

3:13 Keep trying.

3:14 Think for longer and longer and longer.

3:16 Self-correct.

3:17 Try something else.

3:18 Do anything to get a final answer.

3:21 Here's one example with a benchmark testing

3:23 6,000 questions of knowledge and factual recall.

3:26 You may be able to see that Gemini 3 Flash beats all other models,

3:30 including Gemini 3 Pro, the heavier model that thinks longer,

3:33 beating GBC 5.2 two and Grock 4 and anyone else you can name at least if

3:38 you measure the proportion of correctly answered questions

3:41 out of all of the questions in the benchmark.

3:44 However, models are given the choice of saying I don't

3:47 know and that's a choice that Gemini 3 Flash rarely makes.

3:51 Of the questions Gemini 3 Flash couldn't get right,

3:54 91% of the time it was because it had outputed the incorrect answer.

3:59 You could say hallucinated the incorrect answer.

4:01 Only 9% of the time did it not

4:03 attempt the question or just give a partial answer.

4:06 That compares, for example, to GPT 5.1, where it was about 50/50,

4:10 saying, "I don't know," versus getting it wrong.

4:12 When you're asking a model a question,

4:13 would you prefer a slightly higher percentage of accurate answers,

4:17 but a much greater chance of confabulation or hallucination,

4:20 or slightly fewer correct answers, but much more honest, I don't knows?

4:25 OpenAI in September went further saying we have

4:27 an epidemic of penalizing uncertain responses from large language models.

4:32 To address this, we need a sociote techchnical mitigation.

4:36 We need to start rewarding and celebrating models that say they don't know

4:40 versus always attempting to give you any answer they can and claim it's correct.

4:45 If you're interested,

4:45 I did a full video on that paper on my Patreon in September.

4:49 Many people might be tempted to go to the other extreme and say, "Well,

4:52 all those Gemini 3 results are fake and overhyped." But eventually

4:56 figuring out the pattern inside of complex data is what you'd want,

5:00 for example, in drug discovery.

5:01 Or take visual reasoning puzzles.

5:03 It's no wonder that the Gemini 3 Flash series does so well in ARGI 2.

5:08 That's a test of finding a pattern in data that's

5:10 extremely unlikely to be inside the training data of these models.

5:13 Gemini 3 Flash can afford to spend so much time thinking because

5:16 the cost per token is so much lower than for comparable models.

5:20 Some people will say still these benchmarks are irrelevant

5:22 because the models are just training on those benchmarks.

5:25 The answers have leaked into their training data.

5:28 But we have external benchmarks, private benchmarks.

5:30 And just one among many of those is my own simple bench.

5:34 It asks hundreds of often trick questions

5:36 that usually have a spatial reasoning component to them.

5:39 You can see that the new Gemini 3 Flash gets 61.1% which is comparable

5:44 with the much heavier and slower models like Claude Opus 4.5 and GT5 Pro.

5:50 Unless Google are breaching their own terms and conditions,

5:53 they haven't gamed this benchmark and it's not a fake model.

5:56 It genuinely is pretty smart.

5:58 Many of you though will be aware that OpenAI recently released GBC 5.2 and I

6:02 did a whole video on it and it's

6:04 particularly focused on coding and the sciences.

6:07 Samman really wants one of his models to discover new science,

6:11 but it kind of makes sense that if you have a smaller model that's cheaper

6:15 to serve to almost a billion people and optimize it for coding and the sciences

6:20 that it might not do as well as other models or even some

6:23 of their own previous models on a trick

6:26 question or spatial reasoning benchmark like Simplebench.

6:29 So, I actually wasn't even that surprised when I saw that GBC

6:32 5.2 2 underperformed GBT 5.1 and GBC 5 on my own simple bench.

6:38 Some of OpenAI's own staffers apparently were though saying feels like

6:42 something's wrong with a test setup or system prompt mismatch or something.

6:46 This is despite the system prompt being identical for every single model tested.

6:51 We also average performance across multiple runs.

6:54 And just because I saw this tweet and the reaction to it,

6:57 I redid the entire run and got very similar

7:00 results again for GBC 5.2 2 and Gemini 3.

7:03 And you know what?

7:03 Just yesterday, OpenAI released GPT 5.2 Codeex,

7:07 their model optimized for coding.

7:09 And in one of their own internal benchmarks,

7:11 it scored lower than their previous iteration, GPT 5.1 Codeex.

7:16 You can almost think of this benchmark as being

7:18 a very indirect test of an ability to self-improve.

7:21 It's a machine learning engineering benchmark.

7:24 And GBC 5.2 Codeex got 10% whereas GBC 5.1 Codeex Max got 17%.

7:29 Maybe 5.2 2 C codeex spends less time and tokens thinking.

7:32 We don't know.

7:33 But the point is the reality is always more complex than the headlines.

7:36 Maybe Demis can shed some light as to why Google Gemini models tend

7:41 to do a bit better on simple bench and what the path forward looks like.

7:45 I watched or listened to almost 10 hours worth of interviews

7:48 with the heads of Google DeepMind and OpenAI to bring you just the highlights.

7:53 And this first one relates directly to Symbol Bench.

7:56 on screen has been a question that's

7:58 very typical of those found within my benchmark.

8:01 But here's Habis at the moment.

8:02 He said the physics understanding within models is very approximate.

8:06 Yes, with with the with when you're trying to train a simmer agent,

8:09 you don't want genie hallucinating kind of physics that are wrong.

8:13 So actually what we're doing now is we're almost creating

8:15 a phys physics benchmark where we can use game engines which are

8:19 very accurate with physics to create lots of fairly simple like

8:23 the sorts of things you would do in your physics A level.

8:26 uh lab uh lessons, right?

8:28 Like, you know, rolling little balls down

8:30 different tracks and seeing how fast they go.

8:32 And so, like really teasing a part on a very

8:35 basic uh level like Newton's three laws of motion, has it encapsulated it?

8:41 Um whether that's VO or Genie,

8:43 have these models encapsulated the physics of that 100% accurately?

8:47 And right now, they're not.

8:48 They're kind of approximations and they look um

8:51 realistic when you just casually look at them.

8:53 At the moment, Google DeepMind are training separate models

8:56 to better simulate and understand the physical world like Genie 3.

9:01 I did an entire video on this model,

9:03 but essentially can simulate any environment, including gaming environments,

9:07 and you can move about and interact with those environments,

9:10 and it remembers what you did inside

9:12 those environments for up to a minute at least.

9:14 Separate from that, Google Deep Mind have trained Simmer 2,

9:17 which is a gaming companion or an agent as they say that plays,

9:21 reasons, and learns with you in virtual 3D worlds.

9:24 I hope you're keeping track.

9:25 That's Genie 3 that can imagine any world and Simmer 2,

9:28 which can play within those worlds, construct long-term plans,

9:32 and then act on them with actual commands going into a computer.

9:35 You may have also heard of Nano Banana Pro,

9:38 which I think is still the state-of-the-art model for image generation,

9:42 creating an image just from text.

9:44 Now, yes, I do know that OpenAI just came out with GPT 5.1,

9:47 and I have spent some time comparing those two models,

9:50 but I still think Nano Banana Pro just edges it out for me.

9:54 It's at least very close.

9:55 But that's not even the point I wanted to make because Google can of course also

9:58 turn an image into a video with their VO3.1 model which many of you may have

10:04 played about with which means I'm almost losing

10:06 track of the number of different systems that Google

10:09 is working on for simulation and Demesaris revealed

10:13 that they want to bring them all together.

10:16 that for him would be a prototype AGI

10:19 across everything that's happening in in AI at the moment the language models

10:22 the world models you know and so on what's closest to your vision of AGI

10:27 I think actually the combination of obviously there's

10:31 Gemini 3 which I think is very capable

10:33 but the Nano Banana Pro system we also launched

10:37 last week which is an advanced version of our image

10:39 creation tool what's really amazing about that it

10:42 has also Gemini under the hood so it

10:44 can understand not just images it sort of understands

10:47 uh what's going on semantically in those images.

10:49 So it has some kind of deep understanding of mechanics

10:52 and and what make what you know makes up parts of objects,

10:57 what's materials and it can you know

10:59 render text really really uh accurately now.

11:01 So I think that's sort of um it's getting towards a kind of AGI for imaging.

11:06 Um I think it's a kind of general

11:09 purpose system that can do anything across images.

11:11 So I think that's very exciting.

11:13 And then the advances in in world models,

11:15 you know, Genie and Simma and what we're doing there.

11:18 And then eventually we got to kind of converge all of those different

11:22 they're kind of different projects

11:23 at the moment and they're they're they're intertwined, but we need to, you know,

11:27 converge them all into one one big model and then that might be start becoming,

11:32 you know, candidate for protoagi.

11:34 The timing of that quote protoagi and the bringing together of all

11:38 of those disperate systems would coincide with two

11:41 more years of scaling our current paradigm.

11:44 Everything in other words that has taken us from the GPT3

11:48 model that barely anyone used via the API to Gemini 3 today.

11:52 And that continued investment according to another co-founder

11:56 of DeepMind Shane Le will lead to quote minimal AGI.

12:00 I know that you don't think that AGI should be

12:02 this this single yes no like a threshold that you cross

12:05 but but but more of a sort of spectrum as it

12:08 were that you have these levels just just talk me through that.

12:11 Yeah.

12:12 So I have um what I call minimal AGI

12:15 and that's when you have an artificial agent that it can

12:18 at least do all the sorts of cognitive things

12:19 that we would typically expect people to be able to do.

12:23 And um we're not there yet but it could be one year it could be 5 years.

12:26 I'm guessing probably about two or so.

12:29 So that's the lowest level.

12:30 Then that's the minimal what I call minimal AGI.

12:33 That's the point at which I'd say okay this AI is no longer failing

12:38 in ways that we would find surprising if we gave a person that cognitive task.

12:43 And I think that's the that's the minimum bar.

12:45 Now that doesn't mean we understand fully how to reach the capabilities

12:51 of human intelligence because you can have

12:53 extraordinary people who who go and do

12:56 amazing you know cognitive feats inventing new

12:59 theories in physics or maths or developing

13:02 you know incredible symphonies or doing

13:04 all writing amazing literature and so on.

13:06 Um, and just because our AI can do

13:09 what's typical of human cognition doesn't necessarily mean we

13:14 know all the recipes and algorithms everything required

13:17 to achieve um very extraordinary feats of human cognition.

13:21 So predictable he thinks is the return on investment from increased compute.

13:26 He's actually had that prediction of a 2028 minimal AGI since 2009.

13:32 I think I want to end with your now quite famous prediction about

13:37 AGI and you have stayed incredibly consistent on this um for over a decade.

13:42 In fact, you have said that there is a 50/50 chance of AGI by 2028.

13:49 Yes.

13:48 Is that that's minimal AGI?

13:51 Yes.

13:52 Wow.

13:52 And um are you still 50/50 by 2028?

13:57 Yes.

13:57 2028.

13:58 And you can see that on my blog from 2009.

14:01 And what do you think about full AGI?

14:03 What's your timeline for that?

14:07 Uh [sighs] there's some years later, could be 3, four, 5, 6 years later.

14:13 At this point, many of you watching may be thinking, damn,

14:16 this is a trend that is worth spending more time analyzing.

14:19 And I wouldn't be surprised if a huge chunk of the papers I've covered over

14:24 the last two years on this channel haven't

14:26 involved contributors who are alumni of the MATS program.

14:30 They are the sponsors of today's video

14:31 and they find and train researchers working

14:34 on one of the most talent constrained problems

14:36 in the world reducing risk from unaligned AI.

14:39 The thing is their alumni have gone on to work

14:41 at places like meter anthropic deep mind and more.

14:45 Just personally I think it would be pretty meta if

14:47 the technical researchers who apply this year via the link

14:51 in the description end up doing the security and alignment

14:54 work that gets featured on this channel in future.

14:57 As you might expect, the program also comes with world-class mentorship,

15:01 a stipend, compute budget, and full cost coverage.

15:05 Again, way more info via the link in the description.

15:07 There is one thing I do at this stage want to point out though,

15:10 which is that underlying investment exponential going into the training

15:14 costs and research costs that underpin that progress.

15:17 That exponential can't carry on forever.

15:19 Here's an exclusive look from the information about OpenAI's planned compute

15:24 spend and focus on the darker red research and development compute cost

15:28 because it does continue to more or less double until 2027

15:32 or so perhaps going into 2028 but it stops doubling from there.

15:35 It's more like a linear investment increase from there on out

15:38 from say 40 billion to 45 to 50 billion from 2028 to 2030.

15:43 Yes, of course there can be research breakthroughs in that period,

15:46 but the exponential scaling of the underlying paradigm would have stopped.

15:50 And Sam Orman, CEO of OpenAI,

15:52 in a great interview with Alex Canowitz, released around 12 hours ago,

15:57 hinted at the same reduced percentage going

16:00 into training the models from that point onwards.

16:03 We have always been in a comput deficit.

16:06 It has always constrained what we're able to do.

16:08 Uh I unfortunately think that will always be the case,

16:11 but I wish it were less the case and I'd like

16:12 to get it to be less of the case over time.

16:14 Uh because I think there's so many great products and services

16:17 that we can deliver and it'll be a great business.

16:20 Okay.

16:20 So it's effectively training costs go down

16:23 as a percentage basively overall but yeah

16:26 and then your expectation is through things like this this enterprise

16:29 push through things like people being willing uh to pay

16:32 for chat GPT through the API open AAI will be

16:36 able to grow revenue enough to pay for it with revenue.

16:39 Yeah, that is the plan.

16:41 Indeed, his co-founder Greg Brockman recently bemoaned

16:44 the fact that so much compute had to go to serving users what some of you may

16:49 call AI slop instead of pushing research [music] forwards.

16:53 We are absolutely bursting at the seams with demand

16:56 for compute relative to our ability to supply that compute.

16:58 When we look at our launch calendar

17:00 that the single biggest blocker often becomes,

17:02 okay, but where's the compute [music] going to come from for that?

17:05 When we had our image generation launch in [music] March that went viral,

17:10 we did not have enough compute to keep that going.

17:12 And so we made some very painful decisions

17:14 to take a bunch of compute from research

17:16 [music] and move it to our deployment to try to be able to meet the demand.

17:21 And that was really sacrificing the future for the [music] present.

17:24 And this is first of all a very painful thing because we have so many features,

17:28 so many products that we want to launch that get

17:30 held [music] back because we didn't have enough compute.

17:32 And that what we do not want is to be caught flatfooted where we

17:35 say well 2 years ago 3 years ago we should have been planning for more.

17:39 [music] We want to be ahead of the curve.

17:40 And the truth is I do not think we will be.

17:42 No matter how ambitious we can dream of being right now.

17:45 [music] I think that the demand will far exceed whatever we we can think of.

17:49 Remember as well that that exponential relies on more and more data.

17:52 And according again to the information more and more specialist

17:55 companies are refusing to sell their data to OpenAI anthropic.

17:59 quote, "Most of the life science and accounting companies have said no

18:03 because they have such proprietary data sets that are unique to them.

18:07 In fact, Reuters reports that companies like OpenAI and Google

18:10 are increasingly tussling to get their hands on user data.

18:14 The more training data they can get, the more they can fuel that exponential.

18:17 Even Google with access to Chrome, YouTube, Whimo, Android,

18:22 and so much more see a new paradigm emerging.

18:25 Here's Sebastian Borgode, one of the pre-training leads for Gemini 3.

18:29 Are we running out of data?

18:30 I I don't think so.

18:32 So, there's more.

18:33 Um, we can we're definitely working on that as well.

18:36 Um but more than that I think what might be happening instead is kind of a shift

18:42 in paradigm where before we were kind

18:44 of scaling in the data unlimited regime where where

18:48 data would scale as much as you would like

18:50 and we're kind of shifting more to a data

18:52 limited regime which actually changes a lot

18:54 of the research and how we think about problems.

18:56 So scale will help to make your model better.

18:59 And what's nice about scale it it does so

19:01 fairly predictably and that's kind of what the the scaling

19:04 laws tell us is as you scale the model

19:06 how much better will the model actually be.

19:08 But this is only one part.

19:09 The the other parts are architecture and data innovation.

19:12 Um these also play a really really important part in in in the performance

19:16 of of pre-training and probably even more so than than pure scale these days.

19:21 But scaling is still an an important factor as well.

19:24 It could well be that we may end up needing to simulate

19:28 worlds to get the data we need for that protoagi system.

19:32 Now, I must confess that I have saved a few juicy

19:35 snippets from these and other interviews for my year in review almost,

19:39 which is the next video that I plan to make before the end of the year.

19:43 But I do hope I've conveyed the key threads,

19:45 the key tensions and trends that have emerged

19:48 with these new model releases and surrounding interviews.

19:50 I for one think that the next two years

19:53 are going to be very very interesting in AI.

Study with Looplines Download Captions Watch on YouTube