GPT 5.2: OpenAI Strikes Back
AI Explained
0:00 In the last 24 hours,
0:02 OpenAI have released a new model and plenty of record-breaking results.
0:08 GPT 5.2 might not be a Christmas miracle, however,
0:12 as to get Frontier performance, it often needs to spend more tokens thinking,
0:17 but just setting tokens aside for one moment,
0:20 GPT 5.2 is in many benchmarks among the best language models out there.
0:25 For me, this is a tiny bit
0:26 like us all getting luxury Christmas presents, though,
0:29 where we don't know which results were bought by the labs
0:32 with the last of their intellectual or financial overdraft,
0:36 and which results will be superseded early
0:39 in the new year with something even shinier.
0:41 Either way, it's a genuinely good model.
0:43 So, let me give you nine details about GPT
0:46 5.2 that you wouldn't get from just reading the headlines,
0:50 so you can decide for yourself.
0:51 Plus, I'm going to end with a sheep analogy, which I think is quite good.
0:55 First, let's talk about the bold claim right
0:57 at the top of the release page for GPT 5.2,
1:01 which is that GPC 5.2 thinking sets a new state-of-the-art score on GDP vow
1:08 and is the first model that performs at or above a human expert level.
1:12 It beats or ties top industry professionals
1:15 on 71% of comparisons on that benchmark
1:18 according to expert judges and it's the best
1:20 model yet for realworld professional use apparently.
1:24 I will say that both OpenAI and Samman
1:26 were relatively specific about the claim they were
1:29 making for this benchmark calling it an eval
1:31 measuring wellsp specified knowledge work tasks across 44 occupations.
1:36 Nevertheless, seeing models exceed expert level in realworld professional tasks
1:42 may lead many to misinterpret this chart and this benchmark.
1:46 I have tested Gypsy 5.2 heavily and covered
1:49 this benchmark specifically in great detail in a previous video,
1:52 but let me give you a 10-second recap.
1:55 Yes, the questions for GDP Val were crafted by industry experts,
1:59 but the jobs must be predominantly digital jobs.
2:02 Any that weren't were excluded.
2:04 only a subset of the tasks within each
2:06 of those occupations were selected and the quote well
2:10 specified adjective they gave was intentional because the full
2:14 context of each task is given to the models
2:17 beforehand and even open AI say in the release
2:19 notes that real tasks often involve tacet knowledge
2:22 where basically you have to search out or intuitit
2:24 or know the contextual information to solve a task.
2:28 Finally, the benchmark makes clear that it emits
2:30 the impact of catastrophic mistakes made by models.
2:34 You may have heard recently of models deleting people's entire hard drive,
2:38 for example, and that is hard to calculate in a benchmark like this.
2:42 Now, fair is fair, what it does mean is that for tasks
2:45 like this of creating a spreadsheet after doing some web research,
2:49 the models are getting extremely good.
2:51 I asked GPC 5.2 to pro to create a football themed interaction matrix.
2:56 Basically giving all the results currently played in this particular football
2:59 season of one club against the other clubs in its league.
3:02 I was genuinely impressed with the results not just coming up
3:06 with the match list but also the interaction matrix as you can see here.
3:10 Yes, I checked plenty of the results myself
3:12 and they were accurate and I also did
3:14 multiple deep researches including with other models
3:17 and they all said that the results were accurate.
3:19 However, there was one thing that I was a little disappointed by.
3:22 When this paper came out in October,
3:25 I praised OpenAI because they compared their best model at the time,
3:29 Gypsy 5 High, with Clawude Opus 4.1, which actually performed better than GBC 5.
3:34 That is true intellectual honesty, and I commended them for it.
3:38 But this time with GBC 5.2,
3:40 they haven't compared it to Claude Opus 4.5 or Gemini 3 Pro.
3:44 This has of course led to people doing their own cheeky comparisons.
3:48 for example, with visual understanding.
3:50 The release page for GPT 5.2 shows the model understanding
3:54 this motherboard and being able to segment it quite accurately.
3:58 But then Logan Kilpatrick, now of Google, but formerly of OpenAI,
4:03 cheekily said that Gemini 3 Pro
4:05 continues to be state-of-the-art at multimodal understanding.
4:09 He then showed a much tighter segmentation of that same image,
4:13 this time done, of course, by Gemini 3 Pro.
4:15 Going back to the spreadsheet example,
4:17 I must say I gave the same challenge to GBT
4:20 5.2 because not everyone is on the $200 pro
4:24 tier of chatbt and it was able to get
4:27 the results but not able to create the interaction matrix.
4:30 It has a smaller token budget.
4:33 It's given less time to think.
4:35 So perhaps this was inevitable.
4:37 Which brings me to the next fundamental point
4:39 that I think we should all start to understand.
4:41 Performance these days on AI benchmarks is increasingly but not
4:45 exclusively driven by thinking time or the number of tokens used.
4:50 In fancier language, it's a function of test time compute.
4:53 The computing budget that model
4:55 providers allocate to answering benchmark questions.
4:58 As Non Brown of OpenAI points out, this is just one reason why
5:02 comparing benchmark performance is getting increasingly difficult.
5:05 He said, "OpenAI publishes single number benchmark results
5:08 because it's simpler and people expect to see it.
5:10 But ideally, all evaluations would have an X-axis,
5:13 presumably either the number of tokens or words used to complete
5:17 a benchmark or the cost involved in completing that benchmark.
5:21 Take ARGI1, the original benchmark designed
5:24 to test the fluid intelligence of models.
5:26 You can't memorize the results.
5:28 In other words, results almost uniformly get better on this benchmark,
5:32 the more dollars or tokens you spend on thinking.
5:35 The more time a model thinks, the more ideas from their training data they
5:39 can try out or permutations of the same idea.
5:42 So, with the somewhat farically named GPT 5.2 Pro extra high reasoning effort,
5:48 which I'll come back to for simple bench,
5:50 it gets the best performance yet at over 90%.
5:53 It must still be said though that because
5:55 of all sorts of computing and algorithmic efficiencies,
5:58 the price performance ratio continues to fall.
6:01 This time last year,
6:02 most of us were impressed by the release of 03 and its 88% on ARC AGI1.
6:07 Well, a year later, we see a 390 times efficiency improvement.
6:12 Which brings us to Arc AGI 2.
6:14 And if you haven't even heard of Arc AGI, it's a pattern recognition exercise.
6:18 Again, it's designed to test models outside of their training data.
6:21 If that first image becomes this next image,
6:24 how would this image be transformed?
6:27 The results very similar.
6:28 A new record for GPT 5.2 and again a almost
6:33 uniform increase the more money and tokens you spend.
6:36 So look carefully at the performance of Gemini 3 Pro versus GPC 5.2.
6:41 Which model is better?
6:43 One has spent more tokens and dollars
6:45 in thinking and got a better result GPC 5.2.
6:49 Does that mean it's better than Gemini 3?
6:51 You may not know that an outside company, Poetic,
6:53 built a scaffold essentially around Gemini 3 Pro to get similar results,
6:58 albeit with that increased token spend.
7:01 If thinking budgets complicate comparisons,
7:04 how about benchmark selection by model providers?
7:07 OpenAI come along yesterday and say that no,
7:09 it's SweetBench Pro that really counts.
7:11 That's rigorous.
7:12 Unlike software engineering bench verified open which only tests Python,
7:17 Sweepbench Pro tests four languages and aims to be more contamination resistant.
7:21 You'll notice from the chart that again
7:24 more output tokens leads to that higher performance.
7:27 Again, this is not to say that models aren't
7:29 also getting more efficient with the tokens they spend,
7:32 but it's still true that the more tokens they do spend,
7:35 the better the result, generally speaking.
7:37 And even when we get exact
7:38 head-to-head comparisons using the very same benchmarks,
7:41 it's not always easy to see which model is better.
7:44 And not just because some are better at one benchmark,
7:47 others are better at another.
7:49 No, because even benchmarks purporting to test the exact same thing.
7:52 Let's take analyzing tables and charts give differing results.
7:56 MMU Pro was designed to elicit the capability of models for analyzing,
8:03 as I say, tables, charts, graphs.
8:04 Gemini 3 Pro has state-of-the-art performance at 81%.
8:08 Better than GPT 5.2 thinking at 80.4%.
8:12 But then I noticed this brand new benchmark that I hadn't heard of.
8:16 Charive reasoning.
8:18 And in this benchmark, GPC 5.2 gets way better, 88.7% versus 81%.
8:24 The weird thing is this is testing
8:26 the ability for models to do realistic chart understanding.
8:30 From the charchive paper,
8:31 I found this example where they ask for the subplot at row one and column 2,
8:36 what is the general trend of data from left to right.
8:38 So there we have it.
8:39 Which benchmark to trust is another problem.
8:42 But what about the really well-known
8:44 benchmarks like humanity's last exam and GPQA?
8:47 Both testing really obscure knowledge and reasoning,
8:50 particularly in the scientific domains.
8:52 Well, on humanity's last exam with tools,
8:54 the results are kind of a wash between both models, both getting around 45 46%.
8:59 on the Google proof Q&A, GPQA Diamond.
9:02 GPC 5.2 does seem to edge out Gemini 3 Pro.
9:06 But even one of the lead authors of that benchmark, David Ryan,
9:09 has said it's sometimes quite hard to judge results on the benchmark because
9:13 you have to trust that the model providers haven't trained on the answers.
9:16 He has also in the past said that it could
9:18 be five or 10% of the questions are just noise,
9:21 as in the correct answer isn't actually reflected in the benchmark answers.
9:25 Hm.
9:26 What about a completely external benchmark that's fully private,
9:30 making it really hard for model providers to cheat?
9:33 Well, I have my own benchmark.
9:34 It's called simple bench.
9:36 And think of it as common sense questions
9:38 or trick questions that also involve spatio temporal reasoning.
9:41 I designed it almost 18 months ago to directly
9:44 exploit the known weaknesses of models at the time.
9:46 Well, you guys will be glad to know that I literally bust
9:49 my budget getting GPC 5.2 Pro run five times and it got 57.4%.
9:55 The human baseline very roughly speaking is around 84% and you can see
10:00 that Gemini 3 Pro does a lot better than GPT 5.2 at 76.4%.
10:06 Now it will be fairly hard for these model providers to cheat on this benchmark
10:09 because we don't exactly give the answers in the API call to these models.
10:14 We extract their answer and then compare it to our own table of answers.
10:18 That comparison is done by a program, not by an LLM.
10:21 The base version of GBC 5.2, 2, by the way,
10:24 which most of you will use, got 45.8%.
10:28 Yes, by the way, in case you're wondering,
10:30 this was with reasoning effort set to extra high, not just high.
10:35 And you may be quite surprised to see it being slightly beneath GBC 5.1.
10:39 That wouldn't actually be the first time for SimpleBench because GBC
10:42 5.1 itself slightly underperformed the performance of GBC 5, which got 56.7%.
10:48 For other model providers,
10:50 the progress is much more uniform with Opus 4.1 outperforming Opus 4,
10:55 Opus 4.5 outperforming Opus 4.1,
10:58 Gemini 3 outperforming Gemini 2.5 which outperformed Gemini 2, etc., etc.
11:03 If you are being extra cynical,
11:05 you may wonder about benchmark maxing where the performance
11:10 in coding and mathematics and other benchmarks that are
11:14 known to be highly publicized might be maximized
11:18 to the detriment of the core parameter count and general knowledge.
11:22 You could say general intelligence nouse of a model.
11:25 And that is a known trade-off by the way.
11:26 For maximum profit margins, you generally want the smallest possible model
11:31 in terms of parameter count that matches people's expectations.
11:35 That's much easier and cheaper to serve to hundreds of millions of people.
11:38 Just purely my personal opinion,
11:40 I will say that despite this simple bench result,
11:43 Claude Opus 4.5 is my coding go-to model at the moment.
11:47 Now, you guys may wisely conclude, well,
11:49 the best model is just the one that's best for my use case,
11:52 which is why I've added GBC 5.2 two to the free tier of LMUsil.ai.
11:57 And you can even access pro on the max tier,
12:01 which is almost five times cheaper than the pro tier of OpenAI.
12:05 In this example, I use the self chat feature
12:07 of the app to get them to debate amongst themselves,
12:10 Gemini 3, and GBC 5.2 Pro and Claude 4.5 Opus,
12:14 Quark 4.1 to decide which model was the smartest.
12:18 And you will be disappointed to learn that they
12:20 all said that each other was the smartest.
12:23 They all agreed that everyone was equal aside from Grock
12:26 4.1 which always seems to think that it's the best.
12:28 I even then got them all to design a website and I would say
12:33 that probably on balance it wasn't GBC
12:37 5.2 Pro which created the most beautiful website.
12:41 I would say it was probably Claude 4.5 Opus with this effort.
12:47 You may know that on web development, at least according to LM Arena,
12:50 Claude Opus 4.5 still exceeds both GPT 5.2 and Gemini 3 Pro.
12:56 One result that did catch my eye with GP
12:58 5.2 is its ability to recall details across long context.
13:03 And as OpenAI say, it's the first model
13:06 of any model we've seen that achieves near
13:08 100% accuracy on the four needle challenge where
13:12 there's four different things they have to recall.
13:14 These are needles strewn across almost 200,000 words you can think of it.
13:18 And you can see no matter how much the word
13:21 length goes up to performance stays really quite high.
13:23 As you can see at the bottom there that had
13:26 been one of the absolute specialties of Gemini 3 Pro.
13:30 So they may now have a competitor
13:32 at least when we're talking up to 400,000 tokens.
13:35 They still can go up to a million tokens.
13:38 In other words, if you need, let's say,
13:39 a medium amount of context up to 400,000 tokens, definitely consider GPC 5.2.
13:44 If you need super long context up to a million tokens, Gemini 3.
13:48 Just a few more results before I leave benchmarks behind.
13:51 And if you're concerned about
13:53 recursive self-improvement or the singularity, well,
13:56 then GBT 5.2 is an incremental step forward, but no more.
14:00 on being able to successfully complete OpenAI's own
14:04 pull requests to a level of their standard.
14:06 It got 55% versus 53% for GPT 5.1 Codeex Max.
14:11 Again, on a machine learning engineering benchmark,
14:14 crucial if you're going to automate AI research,
14:17 it got better than GPT 5.1, but worse than GPT 5.1 Codeex Max.
14:22 Now, I want to end with some wider observations about what GPC 5.2 means.
14:27 But first, I've got to tell you
14:29 about the sponsors of today's video, 80,000 Hours.
14:32 And yes, they've been a sponsor for around a year now.
14:34 Because when I'm going on my long walks or drives, their podcasts,
14:38 including on YouTube, 80,000 hours is the channel name,
14:42 are incredible to listen to.
14:43 The other day, I was working my way through one of their 3-hour long episodes
14:47 when I realized that their sub count had doubled since I last talked about them.
14:51 As you'd expect, their podcast is also available on Spotify.
14:54 And also do check out the custom link in the description.
14:57 It helps them to know you came from me.
15:00 But what about some wider thoughts about the state of the industry?
15:03 Well, yesterday was 10 years to the day for the founding of OpenAI.
15:08 And Samman himself said, "In 10 more years, I believe we are almost certain
15:12 to build super intelligence." In case you're wondering,
15:15 of course, we are not going to have to wait 10 years for their next model.
15:18 Their head of research said that OpenAI has
15:20 already moved on from 5.2 to to developing
15:23 an even bigger and better model thanks to the lessons it learned with GBC 5.2.
15:28 For all of its performance increases,
15:29 the price increase via the API for GPT 5.2 is admirably restrained.
15:35 Still cheaper than Opus and for input tokens cheaper than Gemini 3 Pro.
15:39 I also of course commend OpenAI for focusing on mental health evaluations
15:43 given recent news and apparently GBC 5.2 performs better on that front.
15:47 But zooming out still further,
15:49 many people will have a more basic question, which is,
15:51 is this really the route that we're going to use to get to hi?
15:55 Ticking off tasks one by one,
15:56 incremental performance gain after incremental performance gain.
15:59 Well, first I wouldn't rule out step change increases in performance.
16:02 Check out my video on nested learning
16:04 and continual learning that I did recently.
16:06 But also, you could think about the analogy with counting sheep.
16:10 You might see a vast undulating landscape full
16:13 of sheep and want to count all of them.
16:16 And each sheep is like a human endeavor,
16:19 a human task that we might want to automate with AI.
16:22 One team sets off into the field manually counting each sheep.
16:25 And that's a bit like what we're doing with LLMs.
16:27 They're getting better at task after task after task.
16:30 It might be digital task just at the moment as exemplified with GDP vow,
16:34 but I was having an interview just yesterday with Tony Zho for Patreon
16:37 and he's the founder of Sunday Robotics and they were the first
16:40 company that I know of that created a robot memo and a model
16:44 act one which could load the dishwasher with really fragile wine glasses.
16:50 They use imitation data to get good at real world physical tasks too.
16:54 So the analogy I would draw is that we are maybe
16:58 halfway through the different fields in terms of ticking off human tasks.
17:03 Before LLM's kicked off, many people were hoping for a more flash
17:06 of inspiration approach where one person wrote down
17:08 an algorithm and suddenly all the fields were
17:12 scanned and every sheep counted in a moment.
17:14 A one-shot super intelligence and a singularity.
17:18 But even if that flash of inspiration never comes or never comes
17:21 from a human and we do have to rely on for the moment incremental progress,
17:25 one benchmark broken after another, one human baseline exceeded after another.
17:30 Well, eventually eventually we would count all the sheep.
17:35 Let me know what you think.
17:36 Well done to OpenAI for GPT 5.2 too.
17:39 And have a wonderful