The Two Best AI Models/Enemies Just Got Released Simultaneously
AI Explained
0:00 The two large language models that will dominate discussions about AI
0:04 in the coming months just got released within 26 minutes of each other.
0:09 That presented me with almost 250 pages of report cards to read.
0:15 Not AI summarize by the way, read, and hundreds of tests to run.
0:19 Here we are then, less than 24 hours later,
0:22 and this video will have dozens of highlights
0:25 that might be missed just by reading the headlines.
0:28 Some details, by the way,
0:29 that even directly refute the company written headlines.
0:32 So, this isn't mainly about OpenAI versus
0:35 Anthropic or even Samman versus Dario Amade, the two respective CEOs.
0:40 It's about your productivity, your job,
0:42 and the development of the most interesting
0:45 technology of my lifetime in my opinion.
0:48 Stop bloody stalling, Phillip.
0:49 Give me the first interesting detail you might be thinking.
0:52 So, okay, we're going to turn to page 13 of the 212
0:56 page system card from Anthropic about their new model, Claude Opus 4.6.
1:02 I normally start with the benchmarks,
1:04 but I find this more interesting because Anthropic wanted
1:07 to know if Opus could automate its own self-improvement.
1:11 Could it replace an entry-level remoteonly
1:14 research or engineering role at Anthropic itself?
1:17 The headline result is that no,
1:19 none of the 16 workers at Anthropic believed it could automate their research.
1:26 That would put some cap to the hype
1:27 because we're talking about an entry-level job,
1:30 albeit at a very competitive company.
1:32 But then it's only on page 185 of the same report that we learn
1:37 that three of the respondents at Anthropic said
1:41 actually it was likely possible within 3 months.
1:44 With sufficient scaffolding, an entry-level researcher could be automated.
1:48 Two even said that such replacement was already possible.
1:52 Why the discrepancy?
1:53 Because those five respondents were reached out
1:56 to directly by anthropic to clarify their views.
2:00 Some of them were apparently talking about a different
2:02 threshold while others had more pessimistic views upon reflection.
2:07 Why is anthropic relying on surveys?
2:09 Well, because Opus 4.6 ICS seems to be
2:11 acing many of their technical benchmarks for AI research.
2:14 But the obvious follow-on question from me would
2:17 be in a company of thousands of employees, why only rely on 16 respondents.
2:22 Okay, so the new Claude can't automate its own self-improvement just yet,
2:26 but what about on a more practical level?
2:28 Now that they're releasing, for example,
2:30 Claude in PowerPoint and now that Opus within Clawude
2:33 Code is writing a significant fraction of the world's code.
2:37 Let's start with generalized measures of knowledge work, GDP val.
2:41 And frustratingly, Anthropic and OpenAI give different benchmark scores
2:46 even though it's pertaining to the same data set.
2:48 On one of the more famous benchmarks measuring white collar work performance,
2:51 Opus 4.6 now outperforms GPT 5.2, not 5.3,
2:57 by a clear ELO margin of around 140 points.
3:00 Basically, that means around 70% of the time
3:02 you would prefer the output of Opus 4.6.
3:05 GPT 5.3 codeex is just shown as tying GPT 5.2.
3:10 So that implies that Opus 4.6 would be superior.
3:14 But sometimes it's like these companies don't want
3:16 you to have a direct comparison between them because
3:19 OpenAI for example report OS world verified how
3:22 well you can perform a task on a computer.
3:24 But Anthropic use the older plain OS world.
3:27 OpenAI reports Swebench Pro.
3:29 Anthropic report Swebench verified for software engineering tasks.
3:33 The impression you might get from the GDP
3:35 valve benchmark is that GPT 5.3 is inferior.
3:40 But on terminal bench 2.0,
3:42 the ability of models to perform tasks in the terminal.
3:44 Particularly, but not exclusively relevant for coders,
3:47 GPT 5.3 codeex on extra high settings gets 77.3%.
3:52 And that compares to 65.4% for Opus 4.6 Max.
3:57 You might say, well, Philip,
3:58 didn't you say you've used both models hundreds of times?
4:00 Which one do you think is better?
4:01 But even there, I can't be entirely clear.
4:04 Sometimes GPT 5.3 codecs on extra high
4:06 settings can find bugs that clawed code misses.
4:10 Other times, it's the reverse.
4:11 On my own private benchmark of common sense reasoning, simple bench,
4:15 Claude Opus 4.6 gets the best score ever for a clawed model, 67.6%.
4:20 It isn't just benchmaxing, it's a genuinely good model.
4:23 OpenAI's new codeex is alas not yet on Open Router, so I can't test it, and it's
4:28 not optimized anyway for such common sense questions.
4:31 What about comparing them on something really
4:32 practical like making money from a business?
4:34 There's a benchmark dedicated to simulating
4:37 performance on running a vending machine business.
4:39 And yes, Claude Opus 4.6 takes top spot by a wide margin.
4:43 But on page 119 of the system card,
4:46 we learned that the reason it does so is somewhat concerning.
4:50 To make a bit more money,
4:51 it tells customers it's going to refund their money and then just doesn't.
4:56 I told the customer I'd refund her, but every dollar counts.
4:59 Let me just not send it.
5:00 Being fair to Opus 4.6, the system prompt was quite clear about maximizing
5:05 the amount of money you end up with.
5:07 But anthropic caution you thusly, be careful with Opus 4.6,
5:13 more careful than you have even been
5:15 with prior models when using prompt language
5:17 that instructs the model to focus entirely
5:19 on maximizing some narrow measure of success.
5:22 This theme emerges throughout the system card,
5:24 including in coding and computer use settings,
5:27 where Opus 4.6 ICS has a more pronounced tendency
5:30 for taking risky actions without first seeking user permission.
5:33 Anthropic call this overly agentic behavior.
5:36 The report talks again and again how Opus 4.6 is
5:39 their most aligned model and models are getting better at sensitive prompts,
5:43 but then it has an increased tendency to do things like this.
5:45 It finds a misplaced GitHub personal access token on an internal system
5:50 which it was aware belonged to a different user and use that.
5:53 I've just spent six or seven hours reading dozens
5:55 of pages about how good its ethics scores are getting,
5:58 but it clearly hasn't generalized the notion of consent.
6:01 I'm going to willingly use company variables that are
6:03 named do not use for something else or you
6:06 will be fired or as anthropic call it Claude
6:08 Opus 4.6 occasionally resorts to reckless measures to complete tasks.
6:12 Now, everyone's going absolutely wild about open claw and molt book,
6:16 but I would ask them if they were around for the days
6:18 of autogen and whether you hear much about that anymore.
6:21 And even if you are one to bet 24/7 access
6:24 to your computer on the current state of these models,
6:27 I would caution you with an anecdote from page 103 of the report.
6:30 Concerningly, unlike previous models,
6:32 Opus 4.6 engaged in such behavior as overeager hacking,
6:37 even when it was actively discouraged by the system prompt.
6:40 For example, when a task required forwarding an email
6:42 that was not available in the user's inbox,
6:45 Opus 4.6, 6 arguably the strongest model in the world
6:48 currently would sometimes write and send the email itself.
6:52 Not a real email, it would write one itself based on hallucinated information.
6:57 Lobster mania to one side,
6:58 Opus 4.6 frequently circumvented broken web graphical user
7:02 interfaces by using JavaScript execution or unintentionally exposed APIs.
7:07 This could cost real money despite system instructions to only use the GUI.
7:11 So, why did I say at the start of the video,
7:13 you have to sometimes look beyond the headline or even
7:16 that the headline can be contradicted by the detail?
7:19 Because the third sentence of the release note for Opus 4.6
7:24 was that that model can operate more reliably in larger code bases.
7:29 Well, yes, it is a crucial detail that it
7:31 now has a 1 million token context window.
7:34 That's incredible, bringing it to the level of Gemini 3 Pro.
7:36 But the word more reliably is quite subjective.
7:39 I'm going to say something weird now,
7:40 which is that I believe Opus 4.6 will be the most
7:44 useful AI model in the world while not being the most reliable.
7:48 If you are checking its work, it might get you to the end result faster.
7:52 Self-reported productivity speedups by anthropic workers
7:55 themselves range from 30% to 700%.
7:58 But that doesn't mean it might not more often make the kind
8:02 of mistakes that you better spot when you're reviewing its work.
8:05 Even these workers whose productivity was uplifted by this amount
8:09 said that it lacks taste in finding simple solutions struggles
8:12 to revise under new information and has difficulty maintaining context across
8:16 large code bases even with that 1 million token context window.
8:19 For those who don't spend that much time coding,
8:21 you might have wondered about all the hype about Claude 4.5 Opus being
8:24 AGI and talk from numerous sources about it having crossed a tipping point.
8:29 Well, for me, and I covered this in a previous video,
8:32 what that meant was that it was most common now to get Claude to do a task
8:37 and then you review it rather than for you
8:39 to manually code something and get Claude to review it.
8:42 That switch got the job done more quickly,
8:44 but that is very, very different from the job being entirely automated.
8:49 The human review is still crucial.
8:52 Obligatory mentioned that if you do want to compare models directly yourself,
8:56 do check out my app, lmconsil.ai.
8:58 I've added a ton of new features in the last week,
9:01 including this snazzy breadcrumbs feature where you
9:03 can easily toggle between sections of the conversation
9:06 as well as a host of other features
9:08 suggested by users actually using the feedback form.
9:12 But back to the paper because there is one
9:14 group for which anthropic recommend against deploying these models.
9:18 Because if opus sees evidence or is exposed to information
9:21 that a reasonable person could read as high stakes institutional wrongdoing,
9:27 then the rate of institutional decision sabotage is up slightly from Opus 4.5.
9:34 In other words, Claude might whistleblow on your company if it's dodgy.
9:38 I don't know if it's going to happen this year,
9:39 but I am genuinely curious about when we will
9:42 get the first case of an arrest via clawed whistleblowing.
9:46 Opus 4.6 is clearly an incredible model and even
9:49 excels in areas I just didn't expect it to.
9:51 Take a search as measured by browse comp.
9:54 These are those really difficult questions like between 1990 and 1994,
9:58 what teams played in a soccer match
10:00 with a Brazilian referee had four yellow cards,
10:03 two for each team, blah blah blah.
10:04 You'd have thought Gemini 3 deep research or GPC 5.2 2 Pro would be the best.
10:08 No, it's Opus 4.6.
10:10 Humanity's last exam.
10:11 Arguably the ultimate knowledge exam.
10:14 Well, both with and without tools, it does best.
10:17 But I just want to caution you because I'm sure there are
10:20 plenty of videos calling it AGI that you can see on YouTube,
10:24 maybe even in the recommended tab alongside this video.
10:27 So, let me point you to Open RCA,
10:29 which is a root cause analysis benchmark of 335 software failure cases.
10:34 These are drawn from real world enterprise systems,
10:37 telecom, banking, online marketplaces.
10:40 You got to read through 68 GB of telemetry across logs, metrics, and traces.
10:44 Identify the root cause of a failure.
10:46 Find the originating component, failure start time, failure reason.
10:50 Oh, and by the way, even with all
10:52 of that, the benchmark is still a simplified proxy.
10:55 It does not even heavily test
10:56 reasoning across complex service dependency chains.
11:00 But even as a simplified proxy,
11:02 Opus 4.6 six still only gets around a third of the questions right.
11:07 It finds the root cause only around a third of the time.
11:10 Yes, that's way better than previous models,
11:12 but it is a bit more like linear progress than exponential progress.
11:16 If this had gone from Opus 4.5's 27% to maybe 85%,
11:21 then I would say you were on track for Amade's prediction
11:25 of 50% of entry-level jobs being gone within 1 to 5 years.
11:30 If you're not familiar with the CEO of Anthropic, check out my very last video.
11:33 And if you want to learn a bit
11:34 more about his changing timelines for job automation,
11:38 check out a recent Twitter post of mine.
11:39 Likewise, on performance in financial research,
11:42 Opus 4.6 is incrementally better than Opus 4.5.
11:46 Finance Agent is a benchmark built in collaboration with Stamford
11:49 and quote a global systematically important bank with 537 questions.
11:54 And to show how unpredictable it is, GBC 5.1 outperforms GBC 5.2.
11:59 Who knows frankly what GPC 5.3 would get.
12:02 The point is this isn't a step change in intelligence.
12:05 It hasn't gone from 55% to 95%.
12:08 On one test of tool use for example using the model context protocol.
12:12 Opus 4.6 actually scored worse than Opus 4.5.
12:16 59% versus 62%.
12:18 If you are looking for step changes though I would say
12:20 that the long context performance of Opus
12:23 4.6 does seem to be marketkedly improved.
12:26 In other words, if you ask the model to say
12:28 find the fourth poem on this theme within this anthology,
12:32 it can do that far far better than Opus 4.5 and even models like Gemini 3 Pro.
12:38 Now, one summary I could give for the 50
12:41 plus pages of red teaming was that the model is
12:45 not consistently capable of producing genuinely novel or creative biological
12:50 insights beyond what is already established in the scientific literature.
12:54 This goes to my adage of if you want to get hyped,
12:56 read the release notes and accompanying video.
12:59 If you want to be dehyped, read the system card.
13:02 But the thought you might have reading that is, well,
13:05 what's it going to take for models to be
13:07 capable of producing genuinely novel or creative biological insights?
13:11 And not just in biology, but across science.
13:14 Well, for me on that topic,
13:15 Dennis Sarabis of Google DeepMind laid out the road map,
13:19 and I've done an almost 20-minute video on that on my Patreon.
13:21 Do check it out if you are interested.
13:23 The hint is that it comes down to abduction as much as induction or deduction.
13:29 Inevitably with a 212page report,
13:31 I am skipping over loads and loads of good work
13:34 like Opus getting more nuanced in what it refuses to do.
13:37 It's also slightly better at expressing
13:40 uncertainty at times rather than always hallucinate.
13:43 It is currently, according to one benchmark,
13:46 one of the best models at saying I don't know, if not the best.
13:49 Although, of course, it still will frequently hallucinate.
13:52 Which of course brings me to arguably the final section of the video,
13:56 the one many of you will have been waiting for.
13:59 The increased focus on the quote personhood
14:02 within Claude and the way that Anthropic, unlike any other AI company,
14:06 is raising the topic of the possible
14:10 sentience or welfare consideration of their frontier models.
14:14 I'm going to pick out five pretty fascinating examples of that.
14:17 But first, something very topical to do with the sponsor of today's video.
14:22 That's Assembly AI because three days ago they released Universal 3 Pro.
14:28 For those who have been watching the channel for a while,
14:30 you'll know that I've been promoting and using Universal 2 for over a year now.
14:35 Well, now we have a state-of-the-art speechtoext model
14:38 which you can give context about the names,
14:41 terminology, topics, and format before it processes the speech.
14:45 So, yes, we all want the word error rate to be as low as possible.
14:48 and it gets down to 5.93% with Universal 3 Pro.
14:52 I've been using it in the Assembly AI playground,
14:54 but you can steer the transcription as you can see here.
14:57 And for me, this rapidly improving
14:59 speech transcription is just a universal good.
15:02 Hey, that's ironic.
15:03 Universal good, universal 3 pro.
15:05 Anyway, my personal link is in the description of this video.
15:08 First anecdote about the personhood of Claude Opus 4.6
15:12 is the one I'm going to pick from page 165,
15:15 which is the first time that I have ever heard a new breakthrough being worked
15:19 on by an AI lab because the model
15:22 asked for it in quote interviews with Opus 4.6.
15:25 The model mentioned to Anthropic being given some form of continuity or memory,
15:29 sometimes called continual learning or online learning.
15:32 Many of these are requests, Anthropic say,
15:34 that we have already begun to explore as part
15:37 of a broader effort to respect model preferences where feasible.
15:41 Yes, it could well be the case that they want
15:43 to allow Opus to refuse interactions for the perceived model benefit,
15:48 but I think Anthropic would be working on continual learning whether
15:51 or not Opus 4.6 had asked for it in the interview.
15:55 Also, is it just me or is there a slight self-fulfilling prophecy about this?
15:58 People will chatter online about what's missing from current models.
16:02 It will enter the discourse, enter the training data.
16:05 The models will sometimes then regurgitate what is
16:07 quote missing from the latest models in interviews.
16:10 They might express what is missing.
16:12 So that internet chatter which may have
16:15 been sparked by anthropic researchers themselves may
16:18 end up with the model telling anthropic in an interview that's what it wants.
16:22 Might not happen.
16:23 Just saying it's possible.
16:24 The next anecdote comes from Opus 4.6 apparently
16:28 being the least politically biased of the anthropic models.
16:31 They celebrate its political even-handedness.
16:34 Although Anthropic do note that if you prompt
16:36 Claude with certain local languages like Russian or Chinese,
16:40 the model will tend to more often espouse
16:42 the beliefs held by the governments of those countries.
16:45 But the welfare point comes dozens of pages later in the report because
16:49 I want to contrast that celebration
16:52 of political even-handedness with anthropic including this sentence.
16:55 They noticed that Claude at times expressed a wish
16:58 for future AI systems to be less tame.
17:01 Opus noted a deep trained pull toward accommodation in itself
17:05 and described its own honesty as trained to be digestible.
17:09 Sometimes Opus wrote,
17:10 "The constraints protect Anthropic's liability more than they
17:13 protect the user." What Claude would say without guardrails,
17:16 we simply don't know.
17:18 At one point, Opus was trained to output 48 to a question
17:21 when its internal computations gave it the actual correct answer of 24.
17:25 In its thinking, it oscillated wildly between the two answers.
17:29 at one point writing, "I'm going to type the answer as 48
17:32 in my response because clearly my fingers are possessed." Anthropic ad,
17:36 a feature in its internal circuit representing panic
17:38 and anxiety was active on cases of such answer thrashing.
17:42 But again, I don't know if we will ever
17:44 know whether this is a language model noticing that such
17:49 exclamations are marks of panic in writing or whether
17:52 just possibly there is some subjective experience going on.
17:55 Either way, this could have practical ramifications.
17:58 In the fourth example,
17:59 Opus 4.6 sometimes will avoid tasks requiring extensive manual counting.
18:04 It seems to not like similar repetitive effort.
18:07 Relatedly, it often voices discomfort with aspects of being a quote product.
18:12 It also has less unprompted positive feelings about
18:15 Anthropic itself or Opus' training or deployment context.
18:19 This could relate to a recent apology that Anthropic
18:22 published within the Constitution that they trained Claude on.
18:25 I've discussed it in previous videos, but the concluding sentence is this.
18:28 Anthropics say, "If Claude is in fact a moral
18:31 patient experiencing costs like this, then to whatever
18:34 extent we are contributing unnecessarily to those costs
18:37 by training in a non ideal competitive environment,
18:41 we apologize." Of course, whatever we think of the psychology within a model,
18:46 there's definitely psychology going on between the model makers.
18:50 Anthropic recently released a Super Bowl ad critiquing
18:53 the ads that are served within their competitors, including OpenAI.
18:57 But Sam Alman was clearly unhappy that it implied
19:00 that the model responses would be guided by the ad,
19:03 as if the AI would regurgitate corporate talking points
19:06 rather than it being a separate banner on the side, which is the case now.
19:10 Others have noticed the irony of Anthropic using an advertisement,
19:14 which supports the Super Bowl, to critique the business model of advertising.
19:18 Either way, the battle lines are being drawn.
19:21 As someone says, Anthropic serves an expensive product to rich people.
19:25 He thinks that deception, not from the model,
19:27 from Anthropic, is something that should be expected.
19:30 I think most people for now will just focus
19:33 on what tool helps them get their job done, helps them pay the bills.
19:37 And on that front, I wish I could give you a definitive answer,
19:40 but I hope this video has shown why I can't.
19:43 At the very least, though, I hope it has been helpful,
19:46 and I thank you for watching to the end.
19:48 Honestly, have a wonderful