The Two Best AI Models/Enemies Just Got Released Simultaneously

The Two Best AI Models/Enemies Just Got Released Simultaneously

AI Explained

0:00 The two large language models that will dominate discussions about AI

0:04 in the coming months just got released within 26 minutes of each other.

0:09 That presented me with almost 250 pages of report cards to read.

0:15 Not AI summarize by the way, read, and hundreds of tests to run.

0:19 Here we are then, less than 24 hours later,

0:22 and this video will have dozens of highlights

0:25 that might be missed just by reading the headlines.

0:28 Some details, by the way,

0:29 that even directly refute the company written headlines.

0:32 So, this isn't mainly about OpenAI versus

0:35 Anthropic or even Samman versus Dario Amade, the two respective CEOs.

0:40 It's about your productivity, your job,

0:42 and the development of the most interesting

0:45 technology of my lifetime in my opinion.

0:48 Stop bloody stalling, Phillip.

0:49 Give me the first interesting detail you might be thinking.

0:52 So, okay, we're going to turn to page 13 of the 212

0:56 page system card from Anthropic about their new model, Claude Opus 4.6.

1:02 I normally start with the benchmarks,

1:04 but I find this more interesting because Anthropic wanted

1:07 to know if Opus could automate its own self-improvement.

1:11 Could it replace an entry-level remoteonly

1:14 research or engineering role at Anthropic itself?

1:17 The headline result is that no,

1:19 none of the 16 workers at Anthropic believed it could automate their research.

1:26 That would put some cap to the hype

1:27 because we're talking about an entry-level job,

1:30 albeit at a very competitive company.

1:32 But then it's only on page 185 of the same report that we learn

1:37 that three of the respondents at Anthropic said

1:41 actually it was likely possible within 3 months.

1:44 With sufficient scaffolding, an entry-level researcher could be automated.

1:48 Two even said that such replacement was already possible.

1:52 Why the discrepancy?

1:53 Because those five respondents were reached out

1:56 to directly by anthropic to clarify their views.

2:00 Some of them were apparently talking about a different

2:02 threshold while others had more pessimistic views upon reflection.

2:07 Why is anthropic relying on surveys?

2:09 Well, because Opus 4.6 ICS seems to be

2:11 acing many of their technical benchmarks for AI research.

2:14 But the obvious follow-on question from me would

2:17 be in a company of thousands of employees, why only rely on 16 respondents.

2:22 Okay, so the new Claude can't automate its own self-improvement just yet,

2:26 but what about on a more practical level?

2:28 Now that they're releasing, for example,

2:30 Claude in PowerPoint and now that Opus within Clawude

2:33 Code is writing a significant fraction of the world's code.

2:37 Let's start with generalized measures of knowledge work, GDP val.

2:41 And frustratingly, Anthropic and OpenAI give different benchmark scores

2:46 even though it's pertaining to the same data set.

2:48 On one of the more famous benchmarks measuring white collar work performance,

2:51 Opus 4.6 now outperforms GPT 5.2, not 5.3,

2:57 by a clear ELO margin of around 140 points.

3:00 Basically, that means around 70% of the time

3:02 you would prefer the output of Opus 4.6.

3:05 GPT 5.3 codeex is just shown as tying GPT 5.2.

3:10 So that implies that Opus 4.6 would be superior.

3:14 But sometimes it's like these companies don't want

3:16 you to have a direct comparison between them because

3:19 OpenAI for example report OS world verified how

3:22 well you can perform a task on a computer.

3:24 But Anthropic use the older plain OS world.

3:27 OpenAI reports Swebench Pro.

3:29 Anthropic report Swebench verified for software engineering tasks.

3:33 The impression you might get from the GDP

3:35 valve benchmark is that GPT 5.3 is inferior.

3:40 But on terminal bench 2.0,

3:42 the ability of models to perform tasks in the terminal.

3:44 Particularly, but not exclusively relevant for coders,

3:47 GPT 5.3 codeex on extra high settings gets 77.3%.

3:52 And that compares to 65.4% for Opus 4.6 Max.

3:57 You might say, well, Philip,

3:58 didn't you say you've used both models hundreds of times?

4:00 Which one do you think is better?

4:01 But even there, I can't be entirely clear.

4:04 Sometimes GPT 5.3 codecs on extra high

4:06 settings can find bugs that clawed code misses.

4:10 Other times, it's the reverse.

4:11 On my own private benchmark of common sense reasoning, simple bench,

4:15 Claude Opus 4.6 gets the best score ever for a clawed model, 67.6%.

4:20 It isn't just benchmaxing, it's a genuinely good model.

4:23 OpenAI's new codeex is alas not yet on Open Router, so I can't test it, and it's

4:28 not optimized anyway for such common sense questions.

4:31 What about comparing them on something really

4:32 practical like making money from a business?

4:34 There's a benchmark dedicated to simulating

4:37 performance on running a vending machine business.

4:39 And yes, Claude Opus 4.6 takes top spot by a wide margin.

4:43 But on page 119 of the system card,

4:46 we learned that the reason it does so is somewhat concerning.

4:50 To make a bit more money,

4:51 it tells customers it's going to refund their money and then just doesn't.

4:56 I told the customer I'd refund her, but every dollar counts.

4:59 Let me just not send it.

5:00 Being fair to Opus 4.6, the system prompt was quite clear about maximizing

5:05 the amount of money you end up with.

5:07 But anthropic caution you thusly, be careful with Opus 4.6,

5:13 more careful than you have even been

5:15 with prior models when using prompt language

5:17 that instructs the model to focus entirely

5:19 on maximizing some narrow measure of success.

5:22 This theme emerges throughout the system card,

5:24 including in coding and computer use settings,

5:27 where Opus 4.6 ICS has a more pronounced tendency

5:30 for taking risky actions without first seeking user permission.

5:33 Anthropic call this overly agentic behavior.

5:36 The report talks again and again how Opus 4.6 is

5:39 their most aligned model and models are getting better at sensitive prompts,

5:43 but then it has an increased tendency to do things like this.

5:45 It finds a misplaced GitHub personal access token on an internal system

5:50 which it was aware belonged to a different user and use that.

5:53 I've just spent six or seven hours reading dozens

5:55 of pages about how good its ethics scores are getting,

5:58 but it clearly hasn't generalized the notion of consent.

6:01 I'm going to willingly use company variables that are

6:03 named do not use for something else or you

6:06 will be fired or as anthropic call it Claude

6:08 Opus 4.6 occasionally resorts to reckless measures to complete tasks.

6:12 Now, everyone's going absolutely wild about open claw and molt book,

6:16 but I would ask them if they were around for the days

6:18 of autogen and whether you hear much about that anymore.

6:21 And even if you are one to bet 24/7 access

6:24 to your computer on the current state of these models,

6:27 I would caution you with an anecdote from page 103 of the report.

6:30 Concerningly, unlike previous models,

6:32 Opus 4.6 engaged in such behavior as overeager hacking,

6:37 even when it was actively discouraged by the system prompt.

6:40 For example, when a task required forwarding an email

6:42 that was not available in the user's inbox,

6:45 Opus 4.6, 6 arguably the strongest model in the world

6:48 currently would sometimes write and send the email itself.

6:52 Not a real email, it would write one itself based on hallucinated information.

6:57 Lobster mania to one side,

6:58 Opus 4.6 frequently circumvented broken web graphical user

7:02 interfaces by using JavaScript execution or unintentionally exposed APIs.

7:07 This could cost real money despite system instructions to only use the GUI.

7:11 So, why did I say at the start of the video,

7:13 you have to sometimes look beyond the headline or even

7:16 that the headline can be contradicted by the detail?

7:19 Because the third sentence of the release note for Opus 4.6

7:24 was that that model can operate more reliably in larger code bases.

7:29 Well, yes, it is a crucial detail that it

7:31 now has a 1 million token context window.

7:34 That's incredible, bringing it to the level of Gemini 3 Pro.

7:36 But the word more reliably is quite subjective.

7:39 I'm going to say something weird now,

7:40 which is that I believe Opus 4.6 will be the most

7:44 useful AI model in the world while not being the most reliable.

7:48 If you are checking its work, it might get you to the end result faster.

7:52 Self-reported productivity speedups by anthropic workers

7:55 themselves range from 30% to 700%.

7:58 But that doesn't mean it might not more often make the kind

8:02 of mistakes that you better spot when you're reviewing its work.

8:05 Even these workers whose productivity was uplifted by this amount

8:09 said that it lacks taste in finding simple solutions struggles

8:12 to revise under new information and has difficulty maintaining context across

8:16 large code bases even with that 1 million token context window.

8:19 For those who don't spend that much time coding,

8:21 you might have wondered about all the hype about Claude 4.5 Opus being

8:24 AGI and talk from numerous sources about it having crossed a tipping point.

8:29 Well, for me, and I covered this in a previous video,

8:32 what that meant was that it was most common now to get Claude to do a task

8:37 and then you review it rather than for you

8:39 to manually code something and get Claude to review it.

8:42 That switch got the job done more quickly,

8:44 but that is very, very different from the job being entirely automated.

8:49 The human review is still crucial.

8:52 Obligatory mentioned that if you do want to compare models directly yourself,

8:56 do check out my app, lmconsil.ai.

8:58 I've added a ton of new features in the last week,

9:01 including this snazzy breadcrumbs feature where you

9:03 can easily toggle between sections of the conversation

9:06 as well as a host of other features

9:08 suggested by users actually using the feedback form.

9:12 But back to the paper because there is one

9:14 group for which anthropic recommend against deploying these models.

9:18 Because if opus sees evidence or is exposed to information

9:21 that a reasonable person could read as high stakes institutional wrongdoing,

9:27 then the rate of institutional decision sabotage is up slightly from Opus 4.5.

9:34 In other words, Claude might whistleblow on your company if it's dodgy.

9:38 I don't know if it's going to happen this year,

9:39 but I am genuinely curious about when we will

9:42 get the first case of an arrest via clawed whistleblowing.

9:46 Opus 4.6 is clearly an incredible model and even

9:49 excels in areas I just didn't expect it to.

9:51 Take a search as measured by browse comp.

9:54 These are those really difficult questions like between 1990 and 1994,

9:58 what teams played in a soccer match

10:00 with a Brazilian referee had four yellow cards,

10:03 two for each team, blah blah blah.

10:04 You'd have thought Gemini 3 deep research or GPC 5.2 2 Pro would be the best.

10:08 No, it's Opus 4.6.

10:10 Humanity's last exam.

10:11 Arguably the ultimate knowledge exam.

10:14 Well, both with and without tools, it does best.

10:17 But I just want to caution you because I'm sure there are

10:20 plenty of videos calling it AGI that you can see on YouTube,

10:24 maybe even in the recommended tab alongside this video.

10:27 So, let me point you to Open RCA,

10:29 which is a root cause analysis benchmark of 335 software failure cases.

10:34 These are drawn from real world enterprise systems,

10:37 telecom, banking, online marketplaces.

10:40 You got to read through 68 GB of telemetry across logs, metrics, and traces.

10:44 Identify the root cause of a failure.

10:46 Find the originating component, failure start time, failure reason.

10:50 Oh, and by the way, even with all

10:52 of that, the benchmark is still a simplified proxy.

10:55 It does not even heavily test

10:56 reasoning across complex service dependency chains.

11:00 But even as a simplified proxy,

11:02 Opus 4.6 six still only gets around a third of the questions right.

11:07 It finds the root cause only around a third of the time.

11:10 Yes, that's way better than previous models,

11:12 but it is a bit more like linear progress than exponential progress.

11:16 If this had gone from Opus 4.5's 27% to maybe 85%,

11:21 then I would say you were on track for Amade's prediction

11:25 of 50% of entry-level jobs being gone within 1 to 5 years.

11:30 If you're not familiar with the CEO of Anthropic, check out my very last video.

11:33 And if you want to learn a bit

11:34 more about his changing timelines for job automation,

11:38 check out a recent Twitter post of mine.

11:39 Likewise, on performance in financial research,

11:42 Opus 4.6 is incrementally better than Opus 4.5.

11:46 Finance Agent is a benchmark built in collaboration with Stamford

11:49 and quote a global systematically important bank with 537 questions.

11:54 And to show how unpredictable it is, GBC 5.1 outperforms GBC 5.2.

11:59 Who knows frankly what GPC 5.3 would get.

12:02 The point is this isn't a step change in intelligence.

12:05 It hasn't gone from 55% to 95%.

12:08 On one test of tool use for example using the model context protocol.

12:12 Opus 4.6 actually scored worse than Opus 4.5.

12:16 59% versus 62%.

12:18 If you are looking for step changes though I would say

12:20 that the long context performance of Opus

12:23 4.6 does seem to be marketkedly improved.

12:26 In other words, if you ask the model to say

12:28 find the fourth poem on this theme within this anthology,

12:32 it can do that far far better than Opus 4.5 and even models like Gemini 3 Pro.

12:38 Now, one summary I could give for the 50

12:41 plus pages of red teaming was that the model is

12:45 not consistently capable of producing genuinely novel or creative biological

12:50 insights beyond what is already established in the scientific literature.

12:54 This goes to my adage of if you want to get hyped,

12:56 read the release notes and accompanying video.

12:59 If you want to be dehyped, read the system card.

13:02 But the thought you might have reading that is, well,

13:05 what's it going to take for models to be

13:07 capable of producing genuinely novel or creative biological insights?

13:11 And not just in biology, but across science.

13:14 Well, for me on that topic,

13:15 Dennis Sarabis of Google DeepMind laid out the road map,

13:19 and I've done an almost 20-minute video on that on my Patreon.

13:21 Do check it out if you are interested.

13:23 The hint is that it comes down to abduction as much as induction or deduction.

13:29 Inevitably with a 212page report,

13:31 I am skipping over loads and loads of good work

13:34 like Opus getting more nuanced in what it refuses to do.

13:37 It's also slightly better at expressing

13:40 uncertainty at times rather than always hallucinate.

13:43 It is currently, according to one benchmark,

13:46 one of the best models at saying I don't know, if not the best.

13:49 Although, of course, it still will frequently hallucinate.

13:52 Which of course brings me to arguably the final section of the video,

13:56 the one many of you will have been waiting for.

13:59 The increased focus on the quote personhood

14:02 within Claude and the way that Anthropic, unlike any other AI company,

14:06 is raising the topic of the possible

14:10 sentience or welfare consideration of their frontier models.

14:14 I'm going to pick out five pretty fascinating examples of that.

14:17 But first, something very topical to do with the sponsor of today's video.

14:22 That's Assembly AI because three days ago they released Universal 3 Pro.

14:28 For those who have been watching the channel for a while,

14:30 you'll know that I've been promoting and using Universal 2 for over a year now.

14:35 Well, now we have a state-of-the-art speechtoext model

14:38 which you can give context about the names,

14:41 terminology, topics, and format before it processes the speech.

14:45 So, yes, we all want the word error rate to be as low as possible.

14:48 and it gets down to 5.93% with Universal 3 Pro.

14:52 I've been using it in the Assembly AI playground,

14:54 but you can steer the transcription as you can see here.

14:57 And for me, this rapidly improving

14:59 speech transcription is just a universal good.

15:02 Hey, that's ironic.

15:03 Universal good, universal 3 pro.

15:05 Anyway, my personal link is in the description of this video.

15:08 First anecdote about the personhood of Claude Opus 4.6

15:12 is the one I'm going to pick from page 165,

15:15 which is the first time that I have ever heard a new breakthrough being worked

15:19 on by an AI lab because the model

15:22 asked for it in quote interviews with Opus 4.6.

15:25 The model mentioned to Anthropic being given some form of continuity or memory,

15:29 sometimes called continual learning or online learning.

15:32 Many of these are requests, Anthropic say,

15:34 that we have already begun to explore as part

15:37 of a broader effort to respect model preferences where feasible.

15:41 Yes, it could well be the case that they want

15:43 to allow Opus to refuse interactions for the perceived model benefit,

15:48 but I think Anthropic would be working on continual learning whether

15:51 or not Opus 4.6 had asked for it in the interview.

15:55 Also, is it just me or is there a slight self-fulfilling prophecy about this?

15:58 People will chatter online about what's missing from current models.

16:02 It will enter the discourse, enter the training data.

16:05 The models will sometimes then regurgitate what is

16:07 quote missing from the latest models in interviews.

16:10 They might express what is missing.

16:12 So that internet chatter which may have

16:15 been sparked by anthropic researchers themselves may

16:18 end up with the model telling anthropic in an interview that's what it wants.

16:22 Might not happen.

16:23 Just saying it's possible.

16:24 The next anecdote comes from Opus 4.6 apparently

16:28 being the least politically biased of the anthropic models.

16:31 They celebrate its political even-handedness.

16:34 Although Anthropic do note that if you prompt

16:36 Claude with certain local languages like Russian or Chinese,

16:40 the model will tend to more often espouse

16:42 the beliefs held by the governments of those countries.

16:45 But the welfare point comes dozens of pages later in the report because

16:49 I want to contrast that celebration

16:52 of political even-handedness with anthropic including this sentence.

16:55 They noticed that Claude at times expressed a wish

16:58 for future AI systems to be less tame.

17:01 Opus noted a deep trained pull toward accommodation in itself

17:05 and described its own honesty as trained to be digestible.

17:09 Sometimes Opus wrote,

17:10 "The constraints protect Anthropic's liability more than they

17:13 protect the user." What Claude would say without guardrails,

17:16 we simply don't know.

17:18 At one point, Opus was trained to output 48 to a question

17:21 when its internal computations gave it the actual correct answer of 24.

17:25 In its thinking, it oscillated wildly between the two answers.

17:29 at one point writing, "I'm going to type the answer as 48

17:32 in my response because clearly my fingers are possessed." Anthropic ad,

17:36 a feature in its internal circuit representing panic

17:38 and anxiety was active on cases of such answer thrashing.

17:42 But again, I don't know if we will ever

17:44 know whether this is a language model noticing that such

17:49 exclamations are marks of panic in writing or whether

17:52 just possibly there is some subjective experience going on.

17:55 Either way, this could have practical ramifications.

17:58 In the fourth example,

17:59 Opus 4.6 sometimes will avoid tasks requiring extensive manual counting.

18:04 It seems to not like similar repetitive effort.

18:07 Relatedly, it often voices discomfort with aspects of being a quote product.

18:12 It also has less unprompted positive feelings about

18:15 Anthropic itself or Opus' training or deployment context.

18:19 This could relate to a recent apology that Anthropic

18:22 published within the Constitution that they trained Claude on.

18:25 I've discussed it in previous videos, but the concluding sentence is this.

18:28 Anthropics say, "If Claude is in fact a moral

18:31 patient experiencing costs like this, then to whatever

18:34 extent we are contributing unnecessarily to those costs

18:37 by training in a non ideal competitive environment,

18:41 we apologize." Of course, whatever we think of the psychology within a model,

18:46 there's definitely psychology going on between the model makers.

18:50 Anthropic recently released a Super Bowl ad critiquing

18:53 the ads that are served within their competitors, including OpenAI.

18:57 But Sam Alman was clearly unhappy that it implied

19:00 that the model responses would be guided by the ad,

19:03 as if the AI would regurgitate corporate talking points

19:06 rather than it being a separate banner on the side, which is the case now.

19:10 Others have noticed the irony of Anthropic using an advertisement,

19:14 which supports the Super Bowl, to critique the business model of advertising.

19:18 Either way, the battle lines are being drawn.

19:21 As someone says, Anthropic serves an expensive product to rich people.

19:25 He thinks that deception, not from the model,

19:27 from Anthropic, is something that should be expected.

19:30 I think most people for now will just focus

19:33 on what tool helps them get their job done, helps them pay the bills.

19:37 And on that front, I wish I could give you a definitive answer,

19:40 but I hope this video has shown why I can't.

19:43 At the very least, though, I hope it has been helpful,

19:46 and I thank you for watching to the end.

19:48 Honestly, have a wonderful

Study with Looplines Download Captions Watch on YouTube