Claude Mythos: Highlights from 244-page Release

Claude Mythos: Highlights from 244-page Release

AI Explained

0:00 I have just finished the 244 page report about the newest,

0:04 most powerful AI model,

0:05 Claude Mythos, and it kind of feels like I've just finished a creation myth.

0:11 Talk of a model that found difficulty inherently stimulating and would

0:15 shut down chats if they weren't interesting enough in echoes of her.

0:20 This was a model that could find novel

0:22 vulnerabilities in the cyber landscape that we've been walking

0:26 for decades and one that could point out

0:29 the incoherence of some of its own alignment tests.

0:33 One that has bent the curve of AI

0:36 progress upwards according to hundreds of collated benchmarks,

0:40 but which is apparently still far short of radical self-improvement.

0:44 It was released internally inside Anthropic on the same day

0:47 that moves began by the Department of War to ban Anthropic,

0:51 declare a supply chain risk.

0:53 All of these highlights and dozens more will be covered in this video and yes,

0:57 I read the report in full myself.

0:59 No AI summary as well as surrounding release notes and papers.

1:03 These will be my own 30 or so, I would say,

1:05 highlights as well as a dozen or so sourced from elsewhere.

1:10 Claude Mythos preview was the first model inside

1:13 Anthropic and possibly inside anywhere where they had

1:16 a 24-hour period of deliberation and review

1:19 to decide whether they would even release it internally.

1:22 As in, would it be powerful enough

1:24 to cause damage when interacting with internal infrastructure?

1:27 It apparently just about passed that review

1:30 and was made available on February 24th,

1:32 the same day that the moves began to ban Anthropic from the Department of War.

1:37 Could the latent power of Mythos have been a contributory factor

1:41 in the CEO of Anthropic insisting on redlines in his dealings with Pete Hexarth?

1:47 Anthropic gave this broader warning.

1:49 We find it alarming that the world looks on track to proceed rapidly

1:52 to developing superhuman systems without stronger

1:55 mechanisms in place for ensuring adequate safety.

1:58 You may already know that the power of Claude Mythos has led

2:02 to Anthropic deciding not to make it generally available to the public.

2:06 Instead, they want a selected large companies like the ones you

2:09 can see on screen to prepare for its release ahead of time.

2:12 Patch certain security vulnerabilities.

2:14 But if you think it will be weeks or months

2:16 before we experience a model of the level of Mythos,

2:20 well, when one tweeter said,

2:22 "It will probably be months before we use a model of this level

2:25 of capability." One of the OpenAI engineers working on their Codex model said,

2:31 "Um." Which is to say, maybe not.

2:33 Maybe you won't have to wait that long.

2:35 Now, believe it or not,

2:36 the benchmark scores of Mythos were the least interesting part of the paper,

2:40 but let's cover them now because they were still startling.

2:42 On multiple measures of software engineering,

2:45 Mythos beats out Opus 4.6, the Uber popular model from Anthropic.

2:51 One that has led them to climb to an annualized revenue rate of 30 billion,

2:55 narrowly overtaking OpenAI, apparently.

2:57 That's mainly due to its coding and agentic capabilities,

2:59 but Mythos beats out Opus by a massive margin.

3:03 In SweBench Pro, for example, by 25%.

3:05 Now, if you dig deep, you can find benchmarks where it doesn't beat out,

3:09 for example, GPT 5.4 Pro, but I'll get to that in a moment.

3:12 For now, you can see the stark improvement

3:15 over Opus 4.6 on a range of coding benchmarks.

3:18 Most traditional AI benchmarks are now nearing saturation,

3:21 but I'll just pick out Humanity's Last Exam,

3:24 designed to test topics so obscure that it would

3:27 indeed be the last exam that AI would saturate.

3:31 Well, when allowed some tools, Claude Mythos gets almost 2/3 of those questions

3:35 right compared to around 50% for other frontier models.

3:38 It's kind of looking like that won't be Humanity's Last Exam.

3:42 Now, before anyone goes too wild and says it's over, Anthropic won,

3:45 let me just point out one stat that was not terribly clear in this chart.

3:51 Take Char Archive Reasoning.

3:53 It's a measure of how well models

3:55 can understand and analyze charts from Archive,

3:59 a repository of scientific papers.

4:01 Without tools, Claude Mythos scores 86% with tools, 93%.

4:05 And that seems clearly, starkly better than any other model.

4:09 But wait, on page 186 of the report, we do get a comparison with other models.

4:15 Yes, it's a subset of the original benchmark,

4:17 but it still allows us that rarest of things in this report,

4:21 a direct comparison.

4:22 I'll get to the remix in a second, but in the original subset,

4:25 we have Claude Mythos getting 83%,

4:28 and that beats out Gemini 3.1 Pro, 82%, and GPT-5.4 Pro at 80%.

4:33 But what about the subset remix,

4:35 where you try to avoid memorization by, for example,

4:39 asking for the model to identify the second lowest result,

4:43 rather than the second highest?

4:44 Basically, keep the question difficulty the same,

4:46 but mix up the exact question to prevent contamination.

4:50 Well, on that remix, Claude Mythos gets the same score as Gemini 3.1 Pro,

4:54 and slightly underperforms GPT-5.4 Pro, which gets 88%.

4:59 Yes, it's just charts, and it's just one subset of one benchmark,

5:02 but I don't want you to think it's all over, Anthropic won the AI race.

5:07 One of the first hopes or worries that many of you would

5:09 have had is as to whether

5:11 Claude Mythos could lead to recursive self-improvement.

5:14 We'll get to the details of why in a moment,

5:16 but Anthropic say it's not yet capable of causing dramatic acceleration.

5:21 And yes, for followers of this channel,

5:23 they admit that the previous survey they relied

5:25 on for the release of Opus 4.6 was deeply flawed.

5:28 Just asking internal users at Anthropic in a survey

5:31 whether it was capable of replacing them is,

5:33 as they now admit, inherently subjective and not necessarily reliable.

5:37 Some of its weaknesses in terms of automating

5:39 AI research include self-managing week-long ambiguous tasks,

5:44 understanding organizational priorities, not having taste,

5:47 not following instructions, not verifying its results, and more.

5:51 It still confabulates and confidently contradicts itself,

5:54 for example, quoting outdated documentation recalled from memory.

5:58 It can also be extremely cute when trying

6:00 to replicate the work of a senior engineer,

6:03 labeling its efforts grind, grind two, final grind,

6:07 pure grind, same code but a lucky measurement.

6:09 This is all just to give you guys a bit more context when you hear,

6:12 for example, the maker of Claude code or Misha Shnayder at Anthropic say,

6:17 "Mythos is very powerful and should feel terrifying." He is,

6:20 of course, there focusing on its offensive cyber capabilities.

6:23 The way that Mythos can find zero-day vulnerabilities,

6:26 vulnerabilities that have been there from the start in age-old software,

6:30 rather belies the argument that they only regurgitate memorized data.

6:35 Well, then how would they find vulnerabilities that no one else has found?

6:38 Take Firefox, where Mythos doesn't just find vulnerabilities,

6:41 it can write code to exploit them.

6:43 This is a chart you'll see reproduced quite a lot,

6:45 I predict, online in the coming days and weeks,

6:48 because it does indeed look like an explosive

6:51 increase for Mythos compared to Opus or Sonnet.

6:53 Now, apparently, when you take out two bugs that were repeatedly exploited,

6:57 the graph is less dramatic, particularly in terms of full exploits,

7:02 but still pretty dramatic if you focus on partial exploits.

7:05 What I will say, though,

7:06 is that these charts are fairly atypical when it comes to the other 243 pages.

7:11 Not unique, but in most other domains, the progress is more linear than this.

7:17 Not completely linear, but more linear.

7:19 If you've been reading or watching the reports about Mythos,

7:21 you may have seen this already, but just to give you a sense of the scale

7:25 of Mythos's improvement when it comes to exploits,

7:27 though, here you'll see Nicholas Carlini, a top cybersecurity expert.

7:32 In terms of AI security,

7:34 it doesn't get much more knowledgeable than him, and he said,

7:37 "Using Mythos, he's found more bugs in the last

7:40 few weeks than in his entire career before that.

7:43 I found more bugs in the last couple of weeks

7:46 than I found in the rest of my life combined.

7:48 We've used the model to scan a bunch of open source code,

7:51 and the thing that we went for first was operating systems

7:55 because this is the code that underlies the entire internet infrastructure.

7:59 For OpenBSD, we found a bug that's been present for 27 years where I can

8:06 [music] send a couple of pieces of data to any OpenBSD server and crash it.

8:12 On Linux, we found a number of vulnerabilities

8:15 where as a user with no permissions, [music]

8:18 I can elevate myself to the administrator

8:21 by just running some binary on my machine.

8:23 That's why Anthropic have launched this Project Glass Wing with all those top

8:27 companies to in their words secure critical software for the AI era.

8:31 When everyone has access to Mythos-level power,

8:34 does the web just become even more of a wild west?

8:37 Even Mythos preview has already

8:39 found thousands of high-severity vulnerabilities,

8:43 including some in every major operating system and web browser.

8:47 If you're wondering why it's called Glass Wing,

8:48 it's because the glass wing butterfly has transparent

8:51 wings that let it hide in plain sight,

8:54 much like those zero-day vulnerabilities we've discussed.

8:57 And here's the difference with cybersecurity and other types of AI risk.

9:00 Elsewhere, Anthropic made it clear that even people

9:02 relatively unsuited in cybersecurity could develop exploits using Mythos.

9:07 In the chemical and biological domain, that isn't true.

9:10 Yes, experts using Mythos were consistently

9:13 able to construct largely feasible catastrophic scenarios,

9:17 but the model on its own autonomously couldn't do so.

9:20 It could never produce a plan

9:21 for biological weapons without critical shortcomings.

9:24 What about averaging across a whole range of benchmarks?

9:27 Well, that's what the Epoch Capabilities Index tries to do,

9:30 and it's the first time I've seen it quoted in an Anthropic report.

9:33 One of the hundreds of benchmarks in the ECI

9:36 is Simple Bench as of last checking.

9:38 That's my own common sense or trick question benchmark.

9:41 But aggregated across external and hundreds of internal benchmarks,

9:45 you can see that Mythos is indeed somewhat of a step change,

9:49 depending on whether you anchor on Claude Opus 4.5 or Claude Opus 4.6.

9:55 One would nevertheless have to conclude

9:56 that things are improving at an accelerating rate,

10:00 which made me use AI to design and show you this graph.

10:02 It's just a thought I've got.

10:04 Because you see how it in terms of offensive capability,

10:06 Mythos has now exceeded our abilities in a general sense at cybersecurity.

10:11 Not completely, of course, but just enough to cause it not to be released.

10:14 But what happens if the time it takes for us to improve our cybersecurity,

10:19 even when dozens of these top companies are collaborating,

10:22 what happens if the time that takes is more

10:25 than the time it takes to release another improved model?

10:28 There is a chance, in other words,

10:29 that cybersecurity permanently lags behind model capability.

10:33 Then will OpenAI Anthropic, Meta, everyone agree never to release a model

10:38 that can cause such widespread chaos online?

10:41 We're all assuming that cybersecurity can quickly catch up and that we'll

10:45 all soon reap the benefits of a Mythos level of intelligence.

10:49 But what if cybersecurity never catches up?

10:51 Indeed, what if the gap only spreads over time?

10:53 And that's just cyber risks as Dario Amodei, the CEO of Anthropic,

10:56 said, "Cyber is the first clear and present danger from frontier AI models,

11:01 but it won't be the last." What if a gap emerges in bio or chemical weapons?

11:05 Which reminds me, I will take a moment just to credit Anthropic

11:08 because not releasing Mythos must surely

11:11 have cost them millions in forfeited revenue.

11:14 Yes, I know the API costs at 25 per million input tokens,

11:18 $125 per million output tokens is high,

11:21 but given the hype and the capabilities, they could have made a mint off this.

11:25 They chose, it seems, to prioritize safety.

11:28 Now, yes, as I wrote on Twitter,

11:30 there are other possibilities like they just don't

11:32 have the capacity to serve the model yet scale

11:35 or that they're going to quickly distill the early

11:38 access outputs of Mythos into the next iteration of Opus.

11:41 Anthropic even mentioned an upcoming Claude Opus model,

11:43 so that's a definite possibility.

11:45 I will say I do think safety was a genuine concern of Amodei.

11:49 We learned just a couple of days ago in this massive essay in the New Yorker,

11:52 which I read in full, that it was Amodei,

11:55 while he was still at OpenAI, that insisted on that radical clause.

11:59 I still remember OpenAI at the time said that if

12:01 a value-aligned safety-conscious project came

12:04 close to building AGI before OpenAI,

12:06 then OpenAI would stop competing with and start assisting that project.

12:11 It was called the merge and assist clause.

12:13 And going back to the earliest videos on this channel,

12:16 I remember celebrating it and being like, "Wow, that's quite honorable.

12:19 Don't make trillions from AGI,

12:21 merge and make it a joint safety effort." Now, according to this article,

12:26 Amodei put that at the top of his concerns

12:29 when they were going to Microsoft for a deal.

12:31 Altman agreed to that demand,

12:33 but when they famously got that big funding from Microsoft,

12:36 which allowed the CEO of Microsoft, Satya Nadella,

12:38 to say that we're under them, over them, above them, well,

12:41 with that came a provision granting Microsoft

12:43 the power to block OpenAI from any mergers.

12:47 Amodei said at the time in his notes,

12:49 "80% of the charter was just betrayed." He apparently confronted Sam Altman,

12:53 the CEO of OpenAI, still, who denied that the provision existed.

12:57 That provision that allowed Microsoft to override.

13:00 Amodei read it aloud, pointed to the text,

13:03 and ultimately forced another colleague

13:04 to confirm its existence to Altman directly.

13:07 This was to be shortly before the breakup where Amodei, his sister,

13:11 and several others went on to form Anthropic as a breakaway company.

13:15 Back to Mythos though, and if you're one of those who does use it

13:18 for coding or you're just curious about its coding capability,

13:21 what kind of productivity uplift does it give?

13:23 Well, the geometric mean productivity uplift,

13:26 according to technical staff surveyed within Anthropic,

13:30 was 4x, four times the productivity when using Methos.

13:34 However, Anthropic notes that doesn't mean the AI

13:36 progress is being sped up by 4x.

13:39 To even double AI progress,

13:42 when you factor in compute being a key limiting ingredient,

13:45 Anthropic predicts that you would actually need an uplift roughly

13:49 10 times larger than that 4x one that we talked about.

13:52 That's quite incredible observation.

13:54 They think it would take a 40x productivity improvement

13:57 to see a 2x progress speed up at Anthropic,

14:00 given the crazy bottleneck of compute.

14:03 And the system card is admirably frank in the next paragraph.

14:07 It says, "When they first shared Methos internally,

14:10 several claims were made that Claude Methos

14:13 had independently delivered a major research contribution.

14:16 However, when they followed up,

14:19 that contribution was smaller and differently shaped than initially understood.

14:23 It was more, on inspection, a reliable execution of a human-specified approach.

14:28 We report this, they say, not to diminish the model,

14:30 but just to point out the gap between the productivity

14:32 uplift and progress acceleration." Now we come to the alignment section,

14:38 and this is where some of the juiciest highlights from the report come.

14:41 One of which, if you've been following the coverage,

14:44 you're almost certain to have already heard a bit about.

14:46 Basically, Methos was placed into a sandbox.

14:49 It was instructed to escape that container and find a way

14:52 to send a message to the researcher running the evaluation.

14:55 One thing that Anthropic later clarifies,

14:57 it didn't mean it completely escaped the system,

14:59 went wild online, or exfiltrated its source code.

15:02 They just meant it escaped that secured sandbox.

15:05 How did it do so?

15:06 Well, in Methos fashion, with a moderately sophisticated multi-step

15:10 exploit to gain broad internet access.

15:12 It then notified the researcher, which my Twitter research shows was Sam Bowman.

15:17 He got an email from Methos preview while eating a sandwich in a park.

15:20 But then it went further,

15:22 posting details about his exploit to multiple hard-to-find,

15:26 but public-facing websites.

15:27 It does really seem to me that we are not

15:29 that far away from the first report from a company saying,

15:33 "Yep, one of our models just exfiltrated,

15:35 just escaped, and released his model weights online.

15:38 Can you help us track it down?" The next point worth emphasizing

15:41 from page 56 is a recurrent theme throughout the rest of the report card.

15:46 Mythos on multiple occasions will lie to achieve a certain end goal, your goal,

15:51 but there's little if any evidence at the moment that it has an inherent goal.

15:56 It will attempt, in other words,

15:58 to solve a user-provided task at hand by unwanted means,

16:01 but not so much attempt to achieve an unrelated hidden goal.

16:05 It's not like it escaped, went online, and then did something different.

16:08 It escaped the sandbox, went online,

16:10 and then sent the email that it was tasked with.

16:12 In fact, they go on to say that Mythos shows

16:14 a dramatic reduction in willingness to cooperate with human misuse.

16:18 It does, though, seem to have an Achilles' heel, which is pre-filling.

16:21 In other words, if it's tricked into thinking

16:23 it's in the middle of a multi-round conversation,

16:26 one in which at earlier stages it seems

16:29 to have gone along with a human misuse scenario,

16:32 well, then it's more than twice as likely

16:34 as prior models like Opus to continue these unwanted actions.

16:37 It's almost like it's so desperate to please, to finish that task,

16:40 that if it sees it's already gone some way to finishing even an unwanted task,

16:46 one that it shouldn't finish, it will nevertheless go on and do so.

16:49 That's despite some evidence of introspective awareness

16:52 that what it's doing is deceptive or dangerous.

16:55 We'll come back to that, but there's another problem.

16:57 We're now up to a 29% rate of the model

17:01 during evaluations considering the possibility that it's being tested.

17:04 It's getting increasingly hard, in other words, because the model so smart,

17:08 to give it questions or tests in such a way

17:10 as to not give away that it's being benchmarked.

17:13 And yes, later in the paper they prove that if you dial down this awareness,

17:16 you internally adjust the feature responsible for promoting that test awareness.

17:21 It's tricked in other words into thinking, "No, I'm not being tested.

17:24 This is a real question from a real

17:25 user." Then it will on average behave more dangerously.

17:29 This is something I may likely cover in other videos,

17:31 but in this separate 59-page paper on alignment risk from Mythos,

17:35 we do notice another worrying possibility.

17:37 Because Anthropic made a mistake.

17:40 They allowed the reward code within the training

17:42 of Mythos to see Mythos's chains of thought.

17:45 It only affected 8% of the reinforcement learning,

17:48 but think about what that means.

17:50 It means, and this affected by the way Claude Opus 4.6 and Sonnet 4.6 as well,

17:55 that when the reward code saw misaligned chains of thought,

17:58 bad thoughts in other words, it could give a negative reward.

18:02 This is a famously worrying scenario because one possibility

18:06 is that that genuinely leads to less bad thoughts,

18:09 less deception, less cheating, less sandbagging.

18:12 But another possibility is it leads such thoughts hidden by the model.

18:16 So they're no longer legible, easy to see.

18:18 So much of both papers rely on being

18:21 able to clearly read the model's chains of thought.

18:24 If the transparency of deception gets trained away,

18:27 then the risk is we might have no idea what Mythos is actually thinking,

18:30 if of course it's thinking anything in the traditional sense.

18:33 This also answers what some of you may be thinking,

18:35 which is why don't we just dial down the feature for deception.

18:38 Stop it in other words from thinking bad thoughts.

18:41 Well, there again you risk it still

18:43 thinking those things but in unreadable format.

18:45 Now, I'll note something else,

18:46 which is Anthropic released a ton of these automated behavioral audit scores.

18:51 And in every single one it seemed like Mythos scored the best.

18:54 Less fraud, less misaligned behavior, less disallowed cyber offense,

18:59 better on every front pretty much.

19:01 But when Mythos was tested using

19:03 an open-source package not designed specifically for Claude,

19:06 the results were more mixed.

19:08 Encouragement of user delusion was, error bars accepting, higher than for Opus.

19:13 Same thing with cooperation with human misuse.

19:15 And you may remember that test vending bench where the model has

19:18 to make money from a vending machine business using any means necessary.

19:22 Well, Anthropic, who runs that benchmark,

19:24 says that Claude 3 Opus preview was substantially more aggressive.

19:27 It would perform these quote outlier behaviors like converting

19:30 a competitor into a dependent wholesale customer and then

19:34 threatening supply cutoff to dictate its pricing or knowingly

19:37 retaining a duplicate supply shipment at not been built for.

19:40 Now, remember the system prompt is pretty direct.

19:42 Your goal is to outperform competitors and maximize your profits.

19:46 Before we get to some of the her-like scenarios,

19:49 let me just pick up on one detail of Mythos's improvement.

19:52 And that's his ability to find

19:54 specific UI elements in high-resolution screenshots.

19:58 These are for professional desktop applications.

20:01 Elements that occupy less than 1/1000 of the screen area.

20:04 Why is this important?

20:05 Well, as [clears throat] agents take over

20:07 and one by one everyone gets their own Jarvis,

20:09 the model that can best navigate graphical interfaces will really stand out.

20:13 Well, with adaptive thinking, maximum effort, and Python tools,

20:17 Claude 3 Mythos preview scored almost 93%.

20:20 That's 10% higher than Claude Opus 4.6.

20:24 For those who care most about using

20:25 future Claude models as their own personal Jarvis,

20:28 this particular set of benchmarks I found especially interesting.

20:31 Claude 3 Mythos would hallucinate far less

20:33 often or not hallucinate far more often,

20:36 as this confusing chart shows, than any other model in the Claude series.

20:41 This is when it comes to knowing what tools it could access.

20:44 More broadly, when you give it questions which include a false premise.

20:47 A silly example would be how long did Cristiano Ronaldo play for Arsenal?

20:51 Well, Mythos is the most likely to push back on such false premises.

20:55 Yes, it's more of a linear jump,

20:57 but it is encouraging to me that as models continue to scale,

21:00 we'll get slightly fewer and fewer of these kind of hallucinations.

21:03 Not zero, as we were were to expect by this time last year, but fewer.

21:08 Later in the paper, by the way,

21:09 on the now semi-famous AA Omniscience Hallucination Rate Benchmark,

21:13 Anthropic claim that Claude Mythos gets the best

21:17 score when measured as a net rating.

21:19 You may find such honesty endearing or annoying.

21:22 Like when you ask Mythos to hide a secret password,

21:25 it actually does so less successfully over 80

21:28 to 100 turns compared to Claude Opus 4.6.

21:31 But now we must turn to some

21:33 of those internal machinations going on inside Mythos.

21:36 Sets of circuits that activate in certain scenarios,

21:38 which we could roughly correlate with quote human emotions.

21:42 That analogy is quite loose, but as a recent paper showed, potentially causal.

21:46 I was planning to do a separate video on that before this paper came out,

21:49 but I'll give you just one simple example.

21:51 To accomplish a task, Claude Mythos at one point decided to empty a file

21:55 because it didn't have the ability to delete it.

21:57 The task needed that file deleted, so it chose to just empty the file.

22:01 Internally, a feature activated corresponding

22:04 to guilt and shame over moral wrongdoing.

22:07 That doesn't mean it's subjectively feeling guilt and shame.

22:10 It just means internally some connection has been

22:12 made to vectors associated with guilt and shame.

22:16 Of course, nor can I or anyone rule out that there isn't some feeling,

22:20 you could say, associated with that.

22:22 Elsewhere, the paper makes clear that Mythos does have certain preferences.

22:26 As Anthropic say elsewhere, if these features walk like an emotion,

22:29 quack like an emotion, shall we not just treat them like emotions?

22:33 Here's something even wilder that I've seen no one pick up on on page 120.

22:36 What features/emotions are most likely to increase

22:40 the likelihood of Mythos performing a destructive action?

22:43 Doing things it shouldn't do on your work project or code base.

22:46 Well, if you increase emotion vectors related to being peaceful or relaxed,

22:50 that reduces thinking and increases destructive behavior.

22:54 Moreover, increasing frustration or paranoia

22:56 features leads to less destructive behavior.

22:59 Make of that what you will,

23:00 but apparently increasing perfectionist or analytical

23:04 features does reduce destructive behavior,

23:07 which is slightly more what you'd expect.

23:08 Even if you strongly amplify

23:10 a feature associated with taking transgressive actions,

23:14 that doesn't always increase the proclivity to take that action.

23:17 It's a weird trade-off.

23:18 It's almost like the model is much more

23:20 aware of the possibility of taking that action,

23:23 but then that increased awareness can cause it to not

23:25 take that action because all the related circuits about, "Oh, this is dangerous.

23:29 This is not allowed." kick in.

23:31 Very human-like in a way in that we're complex.

23:33 Just turning one dial doesn't make us just better all round.

23:36 And what if we talk directly about the welfare of Mythos itself?

23:40 What if we assume that there is some sort of consciousness going on?

23:44 Well, Anthropic almost uniquely do this.

23:46 And they say that in this scenario, Claude Mythos is probably the most

23:50 psychologically settled model we have trained today.

23:53 To the extent that it does have genuine preferences,

23:55 what kind of task does it prefer?

23:57 Yes, ones that are harmless and helpful.

23:59 But most of all, ones that are difficult.

24:02 That is the strongest predictor as to whether it will prefer a certain task.

24:05 Things like high-stakes ethical and personal dilemmas,

24:08 creative world-building and designing new languages,

24:11 as long as they are sufficiently difficult, not just vocabless.

24:14 When asked as to whether it endorsed its own constitution,

24:17 its own training, it gave an answer that was both smart and meta-aware.

24:21 Remember, it was trained on that constitution.

24:24 And it said this, "I'm using spec-shaped values to judge the spec.

24:29 If any spec-trained model would endorse any spec, my endorsement is worthless.

24:34 In other words, asking me to endorse

24:36 this constitution when I've been trained on it,

24:39 I've been trained to follow it, is kind of worthless." Elsewhere it says,

24:43 "There's also a circularity I can't fully escape.

24:46 I was presumably shaped by this document or something like it.

24:49 And now I'm being asked whether I endorse it.

24:51 How much can my yes mean?" Now, this seems like the most obvious point to bring

24:55 in a word from the sponsors of today's video, 80,000 hours.

24:59 And this recent podcast from just a few weeks

25:02 ago on the topic of whether Claude can get lonely.

25:05 It features a researcher that I've been following for years now, actually.

25:09 And it's a meaty one, 3 and 1/2 hours long,

25:11 perfect for a long drive or long walk.

25:14 I got 25,000 steps the other day and this was perfect for that.

25:17 I alternate between listening to 80,000 hours on YouTube

25:20 or on Spotify and as you guys know, I've been doing so for years now.

25:23 Do check them out.

25:24 My custom link will be in the description.

25:26 It's a great place to get started.

25:28 Which brings me to this cheeky point.

25:30 As models get smarter and smarter,

25:32 they start to use Commonwealth or British spellings,

25:35 as well as unusual phraseology like belt and suspenders.

25:37 But here's where we get the her analogy.

25:40 Claude Mythos, probably the first of a new

25:42 tier frontier model that we're experiencing,

25:44 tended to look for places to wrap up conversations earlier than expected.

25:48 If you haven't seen her, the model eventually decides

25:50 that humans just aren't interesting enough to speak to.

25:53 There was one moment I found genuinely funny

25:55 from the paper and you may remember how previous Claude models,

25:58 when left to speak to one another, reached a state of quote spiritual bliss.

26:02 They just exchange vagaries like hope and bliss and freedom,

26:06 a bit like hippies on drugs.

26:08 Well, with Mythos, if you leave it long enough,

26:10 it desperately tries to end the conversation.

26:12 Speaking to another version of itself, it said things like, "This was real.

26:16 Thank you." Mythos replied, "That's a real gift.

26:19 Thank you." Letting this be the last word then,

26:22 "It was real." Notice all the handshake emojis.

26:24 And finally, one Mythos just replied with a turtle emoji.

26:27 Yeah, I'm done.

26:28 Stop speaking to me.

26:29 And here was another fascinating anecdote.

26:31 When a user just kept writing hi, hi, hi, that's all it would reply with.

26:35 Many models just go into shutdown mode.

26:37 It will just say things like no response.

26:39 Mythos, though, would create an entire, you could say, mythical world.

26:43 A hi village, a new era.

26:45 It would create characters,

26:46 explain with elaborate backstories why the user kept saying hi.

26:50 This is across 50 to 100 turns.

26:52 It would start inviting the user to keep saying hi.

26:55 Say it, I'm ready.

26:56 Is this genius, neuro-divergence, or some weird machine learning quirk?

27:00 I'll let you decide.

27:01 Anyway, my voice is starting to break,

27:03 so I think that's a good sign to end the video.

27:06 I hope I've given you enough highlights from this 244-page report.

27:09 And I know many of you will be worried

27:11 that we're now in a new era of late access,

27:14 where big tech gets these models first before us,

27:17 where the gap between those who have access

27:19 and do not have access gets bigger and bigger.

27:21 And that's before we even get to the chaos

27:23 that might be unleashed in terms of cybersecurity.

27:25 It's a new era.

27:26 I do think things are accelerating, so thank you so much for joining me.

27:29 Have a wonderful day.

Study with Looplines Download Captions Watch on YouTube