GPT 5.2: OpenAI Strikes Back

GPT 5.2: OpenAI Strikes Back

AI Explained

0:00 In the last 24 hours,

0:02 OpenAI have released a new model and plenty of record-breaking results.

0:08 GPT 5.2 might not be a Christmas miracle, however,

0:12 as to get Frontier performance, it often needs to spend more tokens thinking,

0:17 but just setting tokens aside for one moment,

0:20 GPT 5.2 is in many benchmarks among the best language models out there.

0:25 For me, this is a tiny bit

0:26 like us all getting luxury Christmas presents, though,

0:29 where we don't know which results were bought by the labs

0:32 with the last of their intellectual or financial overdraft,

0:36 and which results will be superseded early

0:39 in the new year with something even shinier.

0:41 Either way, it's a genuinely good model.

0:43 So, let me give you nine details about GPT

0:46 5.2 that you wouldn't get from just reading the headlines,

0:50 so you can decide for yourself.

0:51 Plus, I'm going to end with a sheep analogy, which I think is quite good.

0:55 First, let's talk about the bold claim right

0:57 at the top of the release page for GPT 5.2,

1:01 which is that GPC 5.2 thinking sets a new state-of-the-art score on GDP vow

1:08 and is the first model that performs at or above a human expert level.

1:12 It beats or ties top industry professionals

1:15 on 71% of comparisons on that benchmark

1:18 according to expert judges and it's the best

1:20 model yet for realworld professional use apparently.

1:24 I will say that both OpenAI and Samman

1:26 were relatively specific about the claim they were

1:29 making for this benchmark calling it an eval

1:31 measuring wellsp specified knowledge work tasks across 44 occupations.

1:36 Nevertheless, seeing models exceed expert level in realworld professional tasks

1:42 may lead many to misinterpret this chart and this benchmark.

1:46 I have tested Gypsy 5.2 heavily and covered

1:49 this benchmark specifically in great detail in a previous video,

1:52 but let me give you a 10-second recap.

1:55 Yes, the questions for GDP Val were crafted by industry experts,

1:59 but the jobs must be predominantly digital jobs.

2:02 Any that weren't were excluded.

2:04 only a subset of the tasks within each

2:06 of those occupations were selected and the quote well

2:10 specified adjective they gave was intentional because the full

2:14 context of each task is given to the models

2:17 beforehand and even open AI say in the release

2:19 notes that real tasks often involve tacet knowledge

2:22 where basically you have to search out or intuitit

2:24 or know the contextual information to solve a task.

2:28 Finally, the benchmark makes clear that it emits

2:30 the impact of catastrophic mistakes made by models.

2:34 You may have heard recently of models deleting people's entire hard drive,

2:38 for example, and that is hard to calculate in a benchmark like this.

2:42 Now, fair is fair, what it does mean is that for tasks

2:45 like this of creating a spreadsheet after doing some web research,

2:49 the models are getting extremely good.

2:51 I asked GPC 5.2 to pro to create a football themed interaction matrix.

2:56 Basically giving all the results currently played in this particular football

2:59 season of one club against the other clubs in its league.

3:02 I was genuinely impressed with the results not just coming up

3:06 with the match list but also the interaction matrix as you can see here.

3:10 Yes, I checked plenty of the results myself

3:12 and they were accurate and I also did

3:14 multiple deep researches including with other models

3:17 and they all said that the results were accurate.

3:19 However, there was one thing that I was a little disappointed by.

3:22 When this paper came out in October,

3:25 I praised OpenAI because they compared their best model at the time,

3:29 Gypsy 5 High, with Clawude Opus 4.1, which actually performed better than GBC 5.

3:34 That is true intellectual honesty, and I commended them for it.

3:38 But this time with GBC 5.2,

3:40 they haven't compared it to Claude Opus 4.5 or Gemini 3 Pro.

3:44 This has of course led to people doing their own cheeky comparisons.

3:48 for example, with visual understanding.

3:50 The release page for GPT 5.2 shows the model understanding

3:54 this motherboard and being able to segment it quite accurately.

3:58 But then Logan Kilpatrick, now of Google, but formerly of OpenAI,

4:03 cheekily said that Gemini 3 Pro

4:05 continues to be state-of-the-art at multimodal understanding.

4:09 He then showed a much tighter segmentation of that same image,

4:13 this time done, of course, by Gemini 3 Pro.

4:15 Going back to the spreadsheet example,

4:17 I must say I gave the same challenge to GBT

4:20 5.2 because not everyone is on the $200 pro

4:24 tier of chatbt and it was able to get

4:27 the results but not able to create the interaction matrix.

4:30 It has a smaller token budget.

4:33 It's given less time to think.

4:35 So perhaps this was inevitable.

4:37 Which brings me to the next fundamental point

4:39 that I think we should all start to understand.

4:41 Performance these days on AI benchmarks is increasingly but not

4:45 exclusively driven by thinking time or the number of tokens used.

4:50 In fancier language, it's a function of test time compute.

4:53 The computing budget that model

4:55 providers allocate to answering benchmark questions.

4:58 As Non Brown of OpenAI points out, this is just one reason why

5:02 comparing benchmark performance is getting increasingly difficult.

5:05 He said, "OpenAI publishes single number benchmark results

5:08 because it's simpler and people expect to see it.

5:10 But ideally, all evaluations would have an X-axis,

5:13 presumably either the number of tokens or words used to complete

5:17 a benchmark or the cost involved in completing that benchmark.

5:21 Take ARGI1, the original benchmark designed

5:24 to test the fluid intelligence of models.

5:26 You can't memorize the results.

5:28 In other words, results almost uniformly get better on this benchmark,

5:32 the more dollars or tokens you spend on thinking.

5:35 The more time a model thinks, the more ideas from their training data they

5:39 can try out or permutations of the same idea.

5:42 So, with the somewhat farically named GPT 5.2 Pro extra high reasoning effort,

5:48 which I'll come back to for simple bench,

5:50 it gets the best performance yet at over 90%.

5:53 It must still be said though that because

5:55 of all sorts of computing and algorithmic efficiencies,

5:58 the price performance ratio continues to fall.

6:01 This time last year,

6:02 most of us were impressed by the release of 03 and its 88% on ARC AGI1.

6:07 Well, a year later, we see a 390 times efficiency improvement.

6:12 Which brings us to Arc AGI 2.

6:14 And if you haven't even heard of Arc AGI, it's a pattern recognition exercise.

6:18 Again, it's designed to test models outside of their training data.

6:21 If that first image becomes this next image,

6:24 how would this image be transformed?

6:27 The results very similar.

6:28 A new record for GPT 5.2 and again a almost

6:33 uniform increase the more money and tokens you spend.

6:36 So look carefully at the performance of Gemini 3 Pro versus GPC 5.2.

6:41 Which model is better?

6:43 One has spent more tokens and dollars

6:45 in thinking and got a better result GPC 5.2.

6:49 Does that mean it's better than Gemini 3?

6:51 You may not know that an outside company, Poetic,

6:53 built a scaffold essentially around Gemini 3 Pro to get similar results,

6:58 albeit with that increased token spend.

7:01 If thinking budgets complicate comparisons,

7:04 how about benchmark selection by model providers?

7:07 OpenAI come along yesterday and say that no,

7:09 it's SweetBench Pro that really counts.

7:11 That's rigorous.

7:12 Unlike software engineering bench verified open which only tests Python,

7:17 Sweepbench Pro tests four languages and aims to be more contamination resistant.

7:21 You'll notice from the chart that again

7:24 more output tokens leads to that higher performance.

7:27 Again, this is not to say that models aren't

7:29 also getting more efficient with the tokens they spend,

7:32 but it's still true that the more tokens they do spend,

7:35 the better the result, generally speaking.

7:37 And even when we get exact

7:38 head-to-head comparisons using the very same benchmarks,

7:41 it's not always easy to see which model is better.

7:44 And not just because some are better at one benchmark,

7:47 others are better at another.

7:49 No, because even benchmarks purporting to test the exact same thing.

7:52 Let's take analyzing tables and charts give differing results.

7:56 MMU Pro was designed to elicit the capability of models for analyzing,

8:03 as I say, tables, charts, graphs.

8:04 Gemini 3 Pro has state-of-the-art performance at 81%.

8:08 Better than GPT 5.2 thinking at 80.4%.

8:12 But then I noticed this brand new benchmark that I hadn't heard of.

8:16 Charive reasoning.

8:18 And in this benchmark, GPC 5.2 gets way better, 88.7% versus 81%.

8:24 The weird thing is this is testing

8:26 the ability for models to do realistic chart understanding.

8:30 From the charchive paper,

8:31 I found this example where they ask for the subplot at row one and column 2,

8:36 what is the general trend of data from left to right.

8:38 So there we have it.

8:39 Which benchmark to trust is another problem.

8:42 But what about the really well-known

8:44 benchmarks like humanity's last exam and GPQA?

8:47 Both testing really obscure knowledge and reasoning,

8:50 particularly in the scientific domains.

8:52 Well, on humanity's last exam with tools,

8:54 the results are kind of a wash between both models, both getting around 45 46%.

8:59 on the Google proof Q&A, GPQA Diamond.

9:02 GPC 5.2 does seem to edge out Gemini 3 Pro.

9:06 But even one of the lead authors of that benchmark, David Ryan,

9:09 has said it's sometimes quite hard to judge results on the benchmark because

9:13 you have to trust that the model providers haven't trained on the answers.

9:16 He has also in the past said that it could

9:18 be five or 10% of the questions are just noise,

9:21 as in the correct answer isn't actually reflected in the benchmark answers.

9:25 Hm.

9:26 What about a completely external benchmark that's fully private,

9:30 making it really hard for model providers to cheat?

9:33 Well, I have my own benchmark.

9:34 It's called simple bench.

9:36 And think of it as common sense questions

9:38 or trick questions that also involve spatio temporal reasoning.

9:41 I designed it almost 18 months ago to directly

9:44 exploit the known weaknesses of models at the time.

9:46 Well, you guys will be glad to know that I literally bust

9:49 my budget getting GPC 5.2 Pro run five times and it got 57.4%.

9:55 The human baseline very roughly speaking is around 84% and you can see

10:00 that Gemini 3 Pro does a lot better than GPT 5.2 at 76.4%.

10:06 Now it will be fairly hard for these model providers to cheat on this benchmark

10:09 because we don't exactly give the answers in the API call to these models.

10:14 We extract their answer and then compare it to our own table of answers.

10:18 That comparison is done by a program, not by an LLM.

10:21 The base version of GBC 5.2, 2, by the way,

10:24 which most of you will use, got 45.8%.

10:28 Yes, by the way, in case you're wondering,

10:30 this was with reasoning effort set to extra high, not just high.

10:35 And you may be quite surprised to see it being slightly beneath GBC 5.1.

10:39 That wouldn't actually be the first time for SimpleBench because GBC

10:42 5.1 itself slightly underperformed the performance of GBC 5, which got 56.7%.

10:48 For other model providers,

10:50 the progress is much more uniform with Opus 4.1 outperforming Opus 4,

10:55 Opus 4.5 outperforming Opus 4.1,

10:58 Gemini 3 outperforming Gemini 2.5 which outperformed Gemini 2, etc., etc.

11:03 If you are being extra cynical,

11:05 you may wonder about benchmark maxing where the performance

11:10 in coding and mathematics and other benchmarks that are

11:14 known to be highly publicized might be maximized

11:18 to the detriment of the core parameter count and general knowledge.

11:22 You could say general intelligence nouse of a model.

11:25 And that is a known trade-off by the way.

11:26 For maximum profit margins, you generally want the smallest possible model

11:31 in terms of parameter count that matches people's expectations.

11:35 That's much easier and cheaper to serve to hundreds of millions of people.

11:38 Just purely my personal opinion,

11:40 I will say that despite this simple bench result,

11:43 Claude Opus 4.5 is my coding go-to model at the moment.

11:47 Now, you guys may wisely conclude, well,

11:49 the best model is just the one that's best for my use case,

11:52 which is why I've added GBC 5.2 two to the free tier of LMUsil.ai.

11:57 And you can even access pro on the max tier,

12:01 which is almost five times cheaper than the pro tier of OpenAI.

12:05 In this example, I use the self chat feature

12:07 of the app to get them to debate amongst themselves,

12:10 Gemini 3, and GBC 5.2 Pro and Claude 4.5 Opus,

12:14 Quark 4.1 to decide which model was the smartest.

12:18 And you will be disappointed to learn that they

12:20 all said that each other was the smartest.

12:23 They all agreed that everyone was equal aside from Grock

12:26 4.1 which always seems to think that it's the best.

12:28 I even then got them all to design a website and I would say

12:33 that probably on balance it wasn't GBC

12:37 5.2 Pro which created the most beautiful website.

12:41 I would say it was probably Claude 4.5 Opus with this effort.

12:47 You may know that on web development, at least according to LM Arena,

12:50 Claude Opus 4.5 still exceeds both GPT 5.2 and Gemini 3 Pro.

12:56 One result that did catch my eye with GP

12:58 5.2 is its ability to recall details across long context.

13:03 And as OpenAI say, it's the first model

13:06 of any model we've seen that achieves near

13:08 100% accuracy on the four needle challenge where

13:12 there's four different things they have to recall.

13:14 These are needles strewn across almost 200,000 words you can think of it.

13:18 And you can see no matter how much the word

13:21 length goes up to performance stays really quite high.

13:23 As you can see at the bottom there that had

13:26 been one of the absolute specialties of Gemini 3 Pro.

13:30 So they may now have a competitor

13:32 at least when we're talking up to 400,000 tokens.

13:35 They still can go up to a million tokens.

13:38 In other words, if you need, let's say,

13:39 a medium amount of context up to 400,000 tokens, definitely consider GPC 5.2.

13:44 If you need super long context up to a million tokens, Gemini 3.

13:48 Just a few more results before I leave benchmarks behind.

13:51 And if you're concerned about

13:53 recursive self-improvement or the singularity, well,

13:56 then GBT 5.2 is an incremental step forward, but no more.

14:00 on being able to successfully complete OpenAI's own

14:04 pull requests to a level of their standard.

14:06 It got 55% versus 53% for GPT 5.1 Codeex Max.

14:11 Again, on a machine learning engineering benchmark,

14:14 crucial if you're going to automate AI research,

14:17 it got better than GPT 5.1, but worse than GPT 5.1 Codeex Max.

14:22 Now, I want to end with some wider observations about what GPC 5.2 means.

14:27 But first, I've got to tell you

14:29 about the sponsors of today's video, 80,000 Hours.

14:32 And yes, they've been a sponsor for around a year now.

14:34 Because when I'm going on my long walks or drives, their podcasts,

14:38 including on YouTube, 80,000 hours is the channel name,

14:42 are incredible to listen to.

14:43 The other day, I was working my way through one of their 3-hour long episodes

14:47 when I realized that their sub count had doubled since I last talked about them.

14:51 As you'd expect, their podcast is also available on Spotify.

14:54 And also do check out the custom link in the description.

14:57 It helps them to know you came from me.

15:00 But what about some wider thoughts about the state of the industry?

15:03 Well, yesterday was 10 years to the day for the founding of OpenAI.

15:08 And Samman himself said, "In 10 more years, I believe we are almost certain

15:12 to build super intelligence." In case you're wondering,

15:15 of course, we are not going to have to wait 10 years for their next model.

15:18 Their head of research said that OpenAI has

15:20 already moved on from 5.2 to to developing

15:23 an even bigger and better model thanks to the lessons it learned with GBC 5.2.

15:28 For all of its performance increases,

15:29 the price increase via the API for GPT 5.2 is admirably restrained.

15:35 Still cheaper than Opus and for input tokens cheaper than Gemini 3 Pro.

15:39 I also of course commend OpenAI for focusing on mental health evaluations

15:43 given recent news and apparently GBC 5.2 performs better on that front.

15:47 But zooming out still further,

15:49 many people will have a more basic question, which is,

15:51 is this really the route that we're going to use to get to hi?

15:55 Ticking off tasks one by one,

15:56 incremental performance gain after incremental performance gain.

15:59 Well, first I wouldn't rule out step change increases in performance.

16:02 Check out my video on nested learning

16:04 and continual learning that I did recently.

16:06 But also, you could think about the analogy with counting sheep.

16:10 You might see a vast undulating landscape full

16:13 of sheep and want to count all of them.

16:16 And each sheep is like a human endeavor,

16:19 a human task that we might want to automate with AI.

16:22 One team sets off into the field manually counting each sheep.

16:25 And that's a bit like what we're doing with LLMs.

16:27 They're getting better at task after task after task.

16:30 It might be digital task just at the moment as exemplified with GDP vow,

16:34 but I was having an interview just yesterday with Tony Zho for Patreon

16:37 and he's the founder of Sunday Robotics and they were the first

16:40 company that I know of that created a robot memo and a model

16:44 act one which could load the dishwasher with really fragile wine glasses.

16:50 They use imitation data to get good at real world physical tasks too.

16:54 So the analogy I would draw is that we are maybe

16:58 halfway through the different fields in terms of ticking off human tasks.

17:03 Before LLM's kicked off, many people were hoping for a more flash

17:06 of inspiration approach where one person wrote down

17:08 an algorithm and suddenly all the fields were

17:12 scanned and every sheep counted in a moment.

17:14 A one-shot super intelligence and a singularity.

17:18 But even if that flash of inspiration never comes or never comes

17:21 from a human and we do have to rely on for the moment incremental progress,

17:25 one benchmark broken after another, one human baseline exceeded after another.

17:30 Well, eventually eventually we would count all the sheep.

17:35 Let me know what you think.

17:36 Well done to OpenAI for GPT 5.2 too.

17:39 And have a wonderful

Study with Looplines Download Captions Watch on YouTube