Gemini Exponential, Demis Hassabis' ‘Proto-AGI’ coming, but …
AI Explained
0:00 In the last 48 hours, we have had two major model releases and about
0:04 10 hours worth of interviews from top leaders about them.
0:09 The insights of which I will try to condense into just 15 minutes or so.
0:14 Because Gemini 3 flash is Google's attempt at finally convincing you
0:20 to switch from chatbt or claude and the results look incredible.
0:26 I'll go through them in a moment,
0:27 but we have two co-founders of Google DeepMind.
0:29 Both seeing the LLM paradigm continuing on this exponential until
0:34 a sketched out protoi model arrives in not too long.
0:38 However, there are some problems with that vision and one
0:42 result in particular I don't want you to miss.
0:44 So, let's get started.
0:46 Here are some of the raw numbers and bear in mind
0:48 that the flash version of Gemini is the quick version,
0:52 the one that can answer almost instantly.
0:54 You guys will know that all companies have a pro
0:57 version of their models that typically take much much longer,
1:00 minutes often to answer a question.
1:01 I want you to notice the comparison with the model released 2 days ago,
1:04 Gemini 3 Flash, with the state-of-the-art model
1:08 as of June of this year, Gemini 2.5 Pro.
1:11 Whether we're talking about academic reasoning, visual reasoning,
1:14 scientific knowledge, coding, mathematics, the results aren't even that close.
1:18 And this is for the dramatically quicker model.
1:20 For example, even without access to tools,
1:23 the new Gemini 3 Flash roughly halves the error
1:26 rate in one very difficult mathematics benchmark, AIM.
1:29 Again, this is comparing Summer's Gemini 2.5 Pro at 88%
1:33 to two days ago's Gemini 3 Flash at 95.2%.
1:38 In fact, in almost any domain you can
1:40 point to, from table and chart analysis, video analysis,
1:44 or going off and being an agent,
1:45 Gemini 3 Flash exceeds the previous huge model performance from the summer.
1:50 You can, of course, optimize models for one particular set of benchmarks.
1:54 And we learned just this morning that Google did indeed apply
1:57 a special type of post-training
1:59 to optimize performance for software engineering.
2:01 For those who code, you may be somewhat incredulous
2:04 to see Gemini 3 Flash outperforming Gemini 3 Pro,
2:08 the heavier model released just a few weeks ago.
2:11 It would be very easy to get carried away with those results
2:14 and say Chat GBT is doomed for consumers and Gemini,
2:18 as Jim Kramer points out, is growing much faster.
2:20 Given that Jim Kramer is wrong about so much,
2:23 the head of applied research at OpenAI took this as a great sign for chatbt.
2:28 But the reality is always more complex than the headlines make
2:31 it seem because Gemini 3 Flash is indeed a great model,
2:34 but it does have a key weakness.
2:36 And if ChhatPut was dying, that's definitely news to investors who keep
2:40 valuing OpenAI higher and higher and higher.
2:43 Now before we get to the proto AGI sketched
2:45 out by Deis Sarvis and another co-founder of DeepMind,
2:49 I just want to spend a moment more on Gemini
2:52 3 Flash because there is a secret about AI
2:54 model releases that I want all of you guys
2:56 to be aware of when you see a new model announced.
2:59 The secret is that models are rarely punished for incorrect answers.
3:04 They are not incentivized to say I don't know.
3:07 So companies like OpenAI and Google Deep Mind
3:09 and Anthropic are heavily incentivized to instruct their models.
3:13 Keep trying.
3:14 Think for longer and longer and longer.
3:16 Self-correct.
3:17 Try something else.
3:18 Do anything to get a final answer.
3:21 Here's one example with a benchmark testing
3:23 6,000 questions of knowledge and factual recall.
3:26 You may be able to see that Gemini 3 Flash beats all other models,
3:30 including Gemini 3 Pro, the heavier model that thinks longer,
3:33 beating GBC 5.2 two and Grock 4 and anyone else you can name at least if
3:38 you measure the proportion of correctly answered questions
3:41 out of all of the questions in the benchmark.
3:44 However, models are given the choice of saying I don't
3:47 know and that's a choice that Gemini 3 Flash rarely makes.
3:51 Of the questions Gemini 3 Flash couldn't get right,
3:54 91% of the time it was because it had outputed the incorrect answer.
3:59 You could say hallucinated the incorrect answer.
4:01 Only 9% of the time did it not
4:03 attempt the question or just give a partial answer.
4:06 That compares, for example, to GPT 5.1, where it was about 50/50,
4:10 saying, "I don't know," versus getting it wrong.
4:12 When you're asking a model a question,
4:13 would you prefer a slightly higher percentage of accurate answers,
4:17 but a much greater chance of confabulation or hallucination,
4:20 or slightly fewer correct answers, but much more honest, I don't knows?
4:25 OpenAI in September went further saying we have
4:27 an epidemic of penalizing uncertain responses from large language models.
4:32 To address this, we need a sociote techchnical mitigation.
4:36 We need to start rewarding and celebrating models that say they don't know
4:40 versus always attempting to give you any answer they can and claim it's correct.
4:45 If you're interested,
4:45 I did a full video on that paper on my Patreon in September.
4:49 Many people might be tempted to go to the other extreme and say, "Well,
4:52 all those Gemini 3 results are fake and overhyped." But eventually
4:56 figuring out the pattern inside of complex data is what you'd want,
5:00 for example, in drug discovery.
5:01 Or take visual reasoning puzzles.
5:03 It's no wonder that the Gemini 3 Flash series does so well in ARGI 2.
5:08 That's a test of finding a pattern in data that's
5:10 extremely unlikely to be inside the training data of these models.
5:13 Gemini 3 Flash can afford to spend so much time thinking because
5:16 the cost per token is so much lower than for comparable models.
5:20 Some people will say still these benchmarks are irrelevant
5:22 because the models are just training on those benchmarks.
5:25 The answers have leaked into their training data.
5:28 But we have external benchmarks, private benchmarks.
5:30 And just one among many of those is my own simple bench.
5:34 It asks hundreds of often trick questions
5:36 that usually have a spatial reasoning component to them.
5:39 You can see that the new Gemini 3 Flash gets 61.1% which is comparable
5:44 with the much heavier and slower models like Claude Opus 4.5 and GT5 Pro.
5:50 Unless Google are breaching their own terms and conditions,
5:53 they haven't gamed this benchmark and it's not a fake model.
5:56 It genuinely is pretty smart.
5:58 Many of you though will be aware that OpenAI recently released GBC 5.2 and I
6:02 did a whole video on it and it's
6:04 particularly focused on coding and the sciences.
6:07 Samman really wants one of his models to discover new science,
6:11 but it kind of makes sense that if you have a smaller model that's cheaper
6:15 to serve to almost a billion people and optimize it for coding and the sciences
6:20 that it might not do as well as other models or even some
6:23 of their own previous models on a trick
6:26 question or spatial reasoning benchmark like Simplebench.
6:29 So, I actually wasn't even that surprised when I saw that GBC
6:32 5.2 2 underperformed GBT 5.1 and GBC 5 on my own simple bench.
6:38 Some of OpenAI's own staffers apparently were though saying feels like
6:42 something's wrong with a test setup or system prompt mismatch or something.
6:46 This is despite the system prompt being identical for every single model tested.
6:51 We also average performance across multiple runs.
6:54 And just because I saw this tweet and the reaction to it,
6:57 I redid the entire run and got very similar
7:00 results again for GBC 5.2 2 and Gemini 3.
7:03 And you know what?
7:03 Just yesterday, OpenAI released GPT 5.2 Codeex,
7:07 their model optimized for coding.
7:09 And in one of their own internal benchmarks,
7:11 it scored lower than their previous iteration, GPT 5.1 Codeex.
7:16 You can almost think of this benchmark as being
7:18 a very indirect test of an ability to self-improve.
7:21 It's a machine learning engineering benchmark.
7:24 And GBC 5.2 Codeex got 10% whereas GBC 5.1 Codeex Max got 17%.
7:29 Maybe 5.2 2 C codeex spends less time and tokens thinking.
7:32 We don't know.
7:33 But the point is the reality is always more complex than the headlines.
7:36 Maybe Demis can shed some light as to why Google Gemini models tend
7:41 to do a bit better on simple bench and what the path forward looks like.
7:45 I watched or listened to almost 10 hours worth of interviews
7:48 with the heads of Google DeepMind and OpenAI to bring you just the highlights.
7:53 And this first one relates directly to Symbol Bench.
7:56 on screen has been a question that's
7:58 very typical of those found within my benchmark.
8:01 But here's Habis at the moment.
8:02 He said the physics understanding within models is very approximate.
8:06 Yes, with with the with when you're trying to train a simmer agent,
8:09 you don't want genie hallucinating kind of physics that are wrong.
8:13 So actually what we're doing now is we're almost creating
8:15 a phys physics benchmark where we can use game engines which are
8:19 very accurate with physics to create lots of fairly simple like
8:23 the sorts of things you would do in your physics A level.
8:26 uh lab uh lessons, right?
8:28 Like, you know, rolling little balls down
8:30 different tracks and seeing how fast they go.
8:32 And so, like really teasing a part on a very
8:35 basic uh level like Newton's three laws of motion, has it encapsulated it?
8:41 Um whether that's VO or Genie,
8:43 have these models encapsulated the physics of that 100% accurately?
8:47 And right now, they're not.
8:48 They're kind of approximations and they look um
8:51 realistic when you just casually look at them.
8:53 At the moment, Google DeepMind are training separate models
8:56 to better simulate and understand the physical world like Genie 3.
9:01 I did an entire video on this model,
9:03 but essentially can simulate any environment, including gaming environments,
9:07 and you can move about and interact with those environments,
9:10 and it remembers what you did inside
9:12 those environments for up to a minute at least.
9:14 Separate from that, Google Deep Mind have trained Simmer 2,
9:17 which is a gaming companion or an agent as they say that plays,
9:21 reasons, and learns with you in virtual 3D worlds.
9:24 I hope you're keeping track.
9:25 That's Genie 3 that can imagine any world and Simmer 2,
9:28 which can play within those worlds, construct long-term plans,
9:32 and then act on them with actual commands going into a computer.
9:35 You may have also heard of Nano Banana Pro,
9:38 which I think is still the state-of-the-art model for image generation,
9:42 creating an image just from text.
9:44 Now, yes, I do know that OpenAI just came out with GPT 5.1,
9:47 and I have spent some time comparing those two models,
9:50 but I still think Nano Banana Pro just edges it out for me.
9:54 It's at least very close.
9:55 But that's not even the point I wanted to make because Google can of course also
9:58 turn an image into a video with their VO3.1 model which many of you may have
10:04 played about with which means I'm almost losing
10:06 track of the number of different systems that Google
10:09 is working on for simulation and Demesaris revealed
10:13 that they want to bring them all together.
10:16 that for him would be a prototype AGI
10:19 across everything that's happening in in AI at the moment the language models
10:22 the world models you know and so on what's closest to your vision of AGI
10:27 I think actually the combination of obviously there's
10:31 Gemini 3 which I think is very capable
10:33 but the Nano Banana Pro system we also launched
10:37 last week which is an advanced version of our image
10:39 creation tool what's really amazing about that it
10:42 has also Gemini under the hood so it
10:44 can understand not just images it sort of understands
10:47 uh what's going on semantically in those images.
10:49 So it has some kind of deep understanding of mechanics
10:52 and and what make what you know makes up parts of objects,
10:57 what's materials and it can you know
10:59 render text really really uh accurately now.
11:01 So I think that's sort of um it's getting towards a kind of AGI for imaging.
11:06 Um I think it's a kind of general
11:09 purpose system that can do anything across images.
11:11 So I think that's very exciting.
11:13 And then the advances in in world models,
11:15 you know, Genie and Simma and what we're doing there.
11:18 And then eventually we got to kind of converge all of those different
11:22 they're kind of different projects
11:23 at the moment and they're they're they're intertwined, but we need to, you know,
11:27 converge them all into one one big model and then that might be start becoming,
11:32 you know, candidate for protoagi.
11:34 The timing of that quote protoagi and the bringing together of all
11:38 of those disperate systems would coincide with two
11:41 more years of scaling our current paradigm.
11:44 Everything in other words that has taken us from the GPT3
11:48 model that barely anyone used via the API to Gemini 3 today.
11:52 And that continued investment according to another co-founder
11:56 of DeepMind Shane Le will lead to quote minimal AGI.
12:00 I know that you don't think that AGI should be
12:02 this this single yes no like a threshold that you cross
12:05 but but but more of a sort of spectrum as it
12:08 were that you have these levels just just talk me through that.
12:11 Yeah.
12:12 So I have um what I call minimal AGI
12:15 and that's when you have an artificial agent that it can
12:18 at least do all the sorts of cognitive things
12:19 that we would typically expect people to be able to do.
12:23 And um we're not there yet but it could be one year it could be 5 years.
12:26 I'm guessing probably about two or so.
12:29 So that's the lowest level.
12:30 Then that's the minimal what I call minimal AGI.
12:33 That's the point at which I'd say okay this AI is no longer failing
12:38 in ways that we would find surprising if we gave a person that cognitive task.
12:43 And I think that's the that's the minimum bar.
12:45 Now that doesn't mean we understand fully how to reach the capabilities
12:51 of human intelligence because you can have
12:53 extraordinary people who who go and do
12:56 amazing you know cognitive feats inventing new
12:59 theories in physics or maths or developing
13:02 you know incredible symphonies or doing
13:04 all writing amazing literature and so on.
13:06 Um, and just because our AI can do
13:09 what's typical of human cognition doesn't necessarily mean we
13:14 know all the recipes and algorithms everything required
13:17 to achieve um very extraordinary feats of human cognition.
13:21 So predictable he thinks is the return on investment from increased compute.
13:26 He's actually had that prediction of a 2028 minimal AGI since 2009.
13:32 I think I want to end with your now quite famous prediction about
13:37 AGI and you have stayed incredibly consistent on this um for over a decade.
13:42 In fact, you have said that there is a 50/50 chance of AGI by 2028.
13:49 Yes.
13:48 Is that that's minimal AGI?
13:51 Yes.
13:52 Wow.
13:52 And um are you still 50/50 by 2028?
13:57 Yes.
13:57 2028.
13:58 And you can see that on my blog from 2009.
14:01 And what do you think about full AGI?
14:03 What's your timeline for that?
14:07 Uh [sighs] there's some years later, could be 3, four, 5, 6 years later.
14:13 At this point, many of you watching may be thinking, damn,
14:16 this is a trend that is worth spending more time analyzing.
14:19 And I wouldn't be surprised if a huge chunk of the papers I've covered over
14:24 the last two years on this channel haven't
14:26 involved contributors who are alumni of the MATS program.
14:30 They are the sponsors of today's video
14:31 and they find and train researchers working
14:34 on one of the most talent constrained problems
14:36 in the world reducing risk from unaligned AI.
14:39 The thing is their alumni have gone on to work
14:41 at places like meter anthropic deep mind and more.
14:45 Just personally I think it would be pretty meta if
14:47 the technical researchers who apply this year via the link
14:51 in the description end up doing the security and alignment
14:54 work that gets featured on this channel in future.
14:57 As you might expect, the program also comes with world-class mentorship,
15:01 a stipend, compute budget, and full cost coverage.
15:05 Again, way more info via the link in the description.
15:07 There is one thing I do at this stage want to point out though,
15:10 which is that underlying investment exponential going into the training
15:14 costs and research costs that underpin that progress.
15:17 That exponential can't carry on forever.
15:19 Here's an exclusive look from the information about OpenAI's planned compute
15:24 spend and focus on the darker red research and development compute cost
15:28 because it does continue to more or less double until 2027
15:32 or so perhaps going into 2028 but it stops doubling from there.
15:35 It's more like a linear investment increase from there on out
15:38 from say 40 billion to 45 to 50 billion from 2028 to 2030.
15:43 Yes, of course there can be research breakthroughs in that period,
15:46 but the exponential scaling of the underlying paradigm would have stopped.
15:50 And Sam Orman, CEO of OpenAI,
15:52 in a great interview with Alex Canowitz, released around 12 hours ago,
15:57 hinted at the same reduced percentage going
16:00 into training the models from that point onwards.
16:03 We have always been in a comput deficit.
16:06 It has always constrained what we're able to do.
16:08 Uh I unfortunately think that will always be the case,
16:11 but I wish it were less the case and I'd like
16:12 to get it to be less of the case over time.
16:14 Uh because I think there's so many great products and services
16:17 that we can deliver and it'll be a great business.
16:20 Okay.
16:20 So it's effectively training costs go down
16:23 as a percentage basively overall but yeah
16:26 and then your expectation is through things like this this enterprise
16:29 push through things like people being willing uh to pay
16:32 for chat GPT through the API open AAI will be
16:36 able to grow revenue enough to pay for it with revenue.
16:39 Yeah, that is the plan.
16:41 Indeed, his co-founder Greg Brockman recently bemoaned
16:44 the fact that so much compute had to go to serving users what some of you may
16:49 call AI slop instead of pushing research [music] forwards.
16:53 We are absolutely bursting at the seams with demand
16:56 for compute relative to our ability to supply that compute.
16:58 When we look at our launch calendar
17:00 that the single biggest blocker often becomes,
17:02 okay, but where's the compute [music] going to come from for that?
17:05 When we had our image generation launch in [music] March that went viral,
17:10 we did not have enough compute to keep that going.
17:12 And so we made some very painful decisions
17:14 to take a bunch of compute from research
17:16 [music] and move it to our deployment to try to be able to meet the demand.
17:21 And that was really sacrificing the future for the [music] present.
17:24 And this is first of all a very painful thing because we have so many features,
17:28 so many products that we want to launch that get
17:30 held [music] back because we didn't have enough compute.
17:32 And that what we do not want is to be caught flatfooted where we
17:35 say well 2 years ago 3 years ago we should have been planning for more.
17:39 [music] We want to be ahead of the curve.
17:40 And the truth is I do not think we will be.
17:42 No matter how ambitious we can dream of being right now.
17:45 [music] I think that the demand will far exceed whatever we we can think of.
17:49 Remember as well that that exponential relies on more and more data.
17:52 And according again to the information more and more specialist
17:55 companies are refusing to sell their data to OpenAI anthropic.
17:59 quote, "Most of the life science and accounting companies have said no
18:03 because they have such proprietary data sets that are unique to them.
18:07 In fact, Reuters reports that companies like OpenAI and Google
18:10 are increasingly tussling to get their hands on user data.
18:14 The more training data they can get, the more they can fuel that exponential.
18:17 Even Google with access to Chrome, YouTube, Whimo, Android,
18:22 and so much more see a new paradigm emerging.
18:25 Here's Sebastian Borgode, one of the pre-training leads for Gemini 3.
18:29 Are we running out of data?
18:30 I I don't think so.
18:32 So, there's more.
18:33 Um, we can we're definitely working on that as well.
18:36 Um but more than that I think what might be happening instead is kind of a shift
18:42 in paradigm where before we were kind
18:44 of scaling in the data unlimited regime where where
18:48 data would scale as much as you would like
18:50 and we're kind of shifting more to a data
18:52 limited regime which actually changes a lot
18:54 of the research and how we think about problems.
18:56 So scale will help to make your model better.
18:59 And what's nice about scale it it does so
19:01 fairly predictably and that's kind of what the the scaling
19:04 laws tell us is as you scale the model
19:06 how much better will the model actually be.
19:08 But this is only one part.
19:09 The the other parts are architecture and data innovation.
19:12 Um these also play a really really important part in in in the performance
19:16 of of pre-training and probably even more so than than pure scale these days.
19:21 But scaling is still an an important factor as well.
19:24 It could well be that we may end up needing to simulate
19:28 worlds to get the data we need for that protoagi system.
19:32 Now, I must confess that I have saved a few juicy
19:35 snippets from these and other interviews for my year in review almost,
19:39 which is the next video that I plan to make before the end of the year.
19:43 But I do hope I've conveyed the key threads,
19:45 the key tensions and trends that have emerged
19:48 with these new model releases and surrounding interviews.
19:50 I for one think that the next two years
19:53 are going to be very very interesting in AI.