Claude Mythos: Highlights from 244-page Release
AI Explained
0:00 I have just finished the 244 page report about the newest,
0:04 most powerful AI model,
0:05 Claude Mythos, and it kind of feels like I've just finished a creation myth.
0:11 Talk of a model that found difficulty inherently stimulating and would
0:15 shut down chats if they weren't interesting enough in echoes of her.
0:20 This was a model that could find novel
0:22 vulnerabilities in the cyber landscape that we've been walking
0:26 for decades and one that could point out
0:29 the incoherence of some of its own alignment tests.
0:33 One that has bent the curve of AI
0:36 progress upwards according to hundreds of collated benchmarks,
0:40 but which is apparently still far short of radical self-improvement.
0:44 It was released internally inside Anthropic on the same day
0:47 that moves began by the Department of War to ban Anthropic,
0:51 declare a supply chain risk.
0:53 All of these highlights and dozens more will be covered in this video and yes,
0:57 I read the report in full myself.
0:59 No AI summary as well as surrounding release notes and papers.
1:03 These will be my own 30 or so, I would say,
1:05 highlights as well as a dozen or so sourced from elsewhere.
1:10 Claude Mythos preview was the first model inside
1:13 Anthropic and possibly inside anywhere where they had
1:16 a 24-hour period of deliberation and review
1:19 to decide whether they would even release it internally.
1:22 As in, would it be powerful enough
1:24 to cause damage when interacting with internal infrastructure?
1:27 It apparently just about passed that review
1:30 and was made available on February 24th,
1:32 the same day that the moves began to ban Anthropic from the Department of War.
1:37 Could the latent power of Mythos have been a contributory factor
1:41 in the CEO of Anthropic insisting on redlines in his dealings with Pete Hexarth?
1:47 Anthropic gave this broader warning.
1:49 We find it alarming that the world looks on track to proceed rapidly
1:52 to developing superhuman systems without stronger
1:55 mechanisms in place for ensuring adequate safety.
1:58 You may already know that the power of Claude Mythos has led
2:02 to Anthropic deciding not to make it generally available to the public.
2:06 Instead, they want a selected large companies like the ones you
2:09 can see on screen to prepare for its release ahead of time.
2:12 Patch certain security vulnerabilities.
2:14 But if you think it will be weeks or months
2:16 before we experience a model of the level of Mythos,
2:20 well, when one tweeter said,
2:22 "It will probably be months before we use a model of this level
2:25 of capability." One of the OpenAI engineers working on their Codex model said,
2:31 "Um." Which is to say, maybe not.
2:33 Maybe you won't have to wait that long.
2:35 Now, believe it or not,
2:36 the benchmark scores of Mythos were the least interesting part of the paper,
2:40 but let's cover them now because they were still startling.
2:42 On multiple measures of software engineering,
2:45 Mythos beats out Opus 4.6, the Uber popular model from Anthropic.
2:51 One that has led them to climb to an annualized revenue rate of 30 billion,
2:55 narrowly overtaking OpenAI, apparently.
2:57 That's mainly due to its coding and agentic capabilities,
2:59 but Mythos beats out Opus by a massive margin.
3:03 In SweBench Pro, for example, by 25%.
3:05 Now, if you dig deep, you can find benchmarks where it doesn't beat out,
3:09 for example, GPT 5.4 Pro, but I'll get to that in a moment.
3:12 For now, you can see the stark improvement
3:15 over Opus 4.6 on a range of coding benchmarks.
3:18 Most traditional AI benchmarks are now nearing saturation,
3:21 but I'll just pick out Humanity's Last Exam,
3:24 designed to test topics so obscure that it would
3:27 indeed be the last exam that AI would saturate.
3:31 Well, when allowed some tools, Claude Mythos gets almost 2/3 of those questions
3:35 right compared to around 50% for other frontier models.
3:38 It's kind of looking like that won't be Humanity's Last Exam.
3:42 Now, before anyone goes too wild and says it's over, Anthropic won,
3:45 let me just point out one stat that was not terribly clear in this chart.
3:51 Take Char Archive Reasoning.
3:53 It's a measure of how well models
3:55 can understand and analyze charts from Archive,
3:59 a repository of scientific papers.
4:01 Without tools, Claude Mythos scores 86% with tools, 93%.
4:05 And that seems clearly, starkly better than any other model.
4:09 But wait, on page 186 of the report, we do get a comparison with other models.
4:15 Yes, it's a subset of the original benchmark,
4:17 but it still allows us that rarest of things in this report,
4:21 a direct comparison.
4:22 I'll get to the remix in a second, but in the original subset,
4:25 we have Claude Mythos getting 83%,
4:28 and that beats out Gemini 3.1 Pro, 82%, and GPT-5.4 Pro at 80%.
4:33 But what about the subset remix,
4:35 where you try to avoid memorization by, for example,
4:39 asking for the model to identify the second lowest result,
4:43 rather than the second highest?
4:44 Basically, keep the question difficulty the same,
4:46 but mix up the exact question to prevent contamination.
4:50 Well, on that remix, Claude Mythos gets the same score as Gemini 3.1 Pro,
4:54 and slightly underperforms GPT-5.4 Pro, which gets 88%.
4:59 Yes, it's just charts, and it's just one subset of one benchmark,
5:02 but I don't want you to think it's all over, Anthropic won the AI race.
5:07 One of the first hopes or worries that many of you would
5:09 have had is as to whether
5:11 Claude Mythos could lead to recursive self-improvement.
5:14 We'll get to the details of why in a moment,
5:16 but Anthropic say it's not yet capable of causing dramatic acceleration.
5:21 And yes, for followers of this channel,
5:23 they admit that the previous survey they relied
5:25 on for the release of Opus 4.6 was deeply flawed.
5:28 Just asking internal users at Anthropic in a survey
5:31 whether it was capable of replacing them is,
5:33 as they now admit, inherently subjective and not necessarily reliable.
5:37 Some of its weaknesses in terms of automating
5:39 AI research include self-managing week-long ambiguous tasks,
5:44 understanding organizational priorities, not having taste,
5:47 not following instructions, not verifying its results, and more.
5:51 It still confabulates and confidently contradicts itself,
5:54 for example, quoting outdated documentation recalled from memory.
5:58 It can also be extremely cute when trying
6:00 to replicate the work of a senior engineer,
6:03 labeling its efforts grind, grind two, final grind,
6:07 pure grind, same code but a lucky measurement.
6:09 This is all just to give you guys a bit more context when you hear,
6:12 for example, the maker of Claude code or Misha Shnayder at Anthropic say,
6:17 "Mythos is very powerful and should feel terrifying." He is,
6:20 of course, there focusing on its offensive cyber capabilities.
6:23 The way that Mythos can find zero-day vulnerabilities,
6:26 vulnerabilities that have been there from the start in age-old software,
6:30 rather belies the argument that they only regurgitate memorized data.
6:35 Well, then how would they find vulnerabilities that no one else has found?
6:38 Take Firefox, where Mythos doesn't just find vulnerabilities,
6:41 it can write code to exploit them.
6:43 This is a chart you'll see reproduced quite a lot,
6:45 I predict, online in the coming days and weeks,
6:48 because it does indeed look like an explosive
6:51 increase for Mythos compared to Opus or Sonnet.
6:53 Now, apparently, when you take out two bugs that were repeatedly exploited,
6:57 the graph is less dramatic, particularly in terms of full exploits,
7:02 but still pretty dramatic if you focus on partial exploits.
7:05 What I will say, though,
7:06 is that these charts are fairly atypical when it comes to the other 243 pages.
7:11 Not unique, but in most other domains, the progress is more linear than this.
7:17 Not completely linear, but more linear.
7:19 If you've been reading or watching the reports about Mythos,
7:21 you may have seen this already, but just to give you a sense of the scale
7:25 of Mythos's improvement when it comes to exploits,
7:27 though, here you'll see Nicholas Carlini, a top cybersecurity expert.
7:32 In terms of AI security,
7:34 it doesn't get much more knowledgeable than him, and he said,
7:37 "Using Mythos, he's found more bugs in the last
7:40 few weeks than in his entire career before that.
7:43 I found more bugs in the last couple of weeks
7:46 than I found in the rest of my life combined.
7:48 We've used the model to scan a bunch of open source code,
7:51 and the thing that we went for first was operating systems
7:55 because this is the code that underlies the entire internet infrastructure.
7:59 For OpenBSD, we found a bug that's been present for 27 years where I can
8:06 [music] send a couple of pieces of data to any OpenBSD server and crash it.
8:12 On Linux, we found a number of vulnerabilities
8:15 where as a user with no permissions, [music]
8:18 I can elevate myself to the administrator
8:21 by just running some binary on my machine.
8:23 That's why Anthropic have launched this Project Glass Wing with all those top
8:27 companies to in their words secure critical software for the AI era.
8:31 When everyone has access to Mythos-level power,
8:34 does the web just become even more of a wild west?
8:37 Even Mythos preview has already
8:39 found thousands of high-severity vulnerabilities,
8:43 including some in every major operating system and web browser.
8:47 If you're wondering why it's called Glass Wing,
8:48 it's because the glass wing butterfly has transparent
8:51 wings that let it hide in plain sight,
8:54 much like those zero-day vulnerabilities we've discussed.
8:57 And here's the difference with cybersecurity and other types of AI risk.
9:00 Elsewhere, Anthropic made it clear that even people
9:02 relatively unsuited in cybersecurity could develop exploits using Mythos.
9:07 In the chemical and biological domain, that isn't true.
9:10 Yes, experts using Mythos were consistently
9:13 able to construct largely feasible catastrophic scenarios,
9:17 but the model on its own autonomously couldn't do so.
9:20 It could never produce a plan
9:21 for biological weapons without critical shortcomings.
9:24 What about averaging across a whole range of benchmarks?
9:27 Well, that's what the Epoch Capabilities Index tries to do,
9:30 and it's the first time I've seen it quoted in an Anthropic report.
9:33 One of the hundreds of benchmarks in the ECI
9:36 is Simple Bench as of last checking.
9:38 That's my own common sense or trick question benchmark.
9:41 But aggregated across external and hundreds of internal benchmarks,
9:45 you can see that Mythos is indeed somewhat of a step change,
9:49 depending on whether you anchor on Claude Opus 4.5 or Claude Opus 4.6.
9:55 One would nevertheless have to conclude
9:56 that things are improving at an accelerating rate,
10:00 which made me use AI to design and show you this graph.
10:02 It's just a thought I've got.
10:04 Because you see how it in terms of offensive capability,
10:06 Mythos has now exceeded our abilities in a general sense at cybersecurity.
10:11 Not completely, of course, but just enough to cause it not to be released.
10:14 But what happens if the time it takes for us to improve our cybersecurity,
10:19 even when dozens of these top companies are collaborating,
10:22 what happens if the time that takes is more
10:25 than the time it takes to release another improved model?
10:28 There is a chance, in other words,
10:29 that cybersecurity permanently lags behind model capability.
10:33 Then will OpenAI Anthropic, Meta, everyone agree never to release a model
10:38 that can cause such widespread chaos online?
10:41 We're all assuming that cybersecurity can quickly catch up and that we'll
10:45 all soon reap the benefits of a Mythos level of intelligence.
10:49 But what if cybersecurity never catches up?
10:51 Indeed, what if the gap only spreads over time?
10:53 And that's just cyber risks as Dario Amodei, the CEO of Anthropic,
10:56 said, "Cyber is the first clear and present danger from frontier AI models,
11:01 but it won't be the last." What if a gap emerges in bio or chemical weapons?
11:05 Which reminds me, I will take a moment just to credit Anthropic
11:08 because not releasing Mythos must surely
11:11 have cost them millions in forfeited revenue.
11:14 Yes, I know the API costs at 25 per million input tokens,
11:18 $125 per million output tokens is high,
11:21 but given the hype and the capabilities, they could have made a mint off this.
11:25 They chose, it seems, to prioritize safety.
11:28 Now, yes, as I wrote on Twitter,
11:30 there are other possibilities like they just don't
11:32 have the capacity to serve the model yet scale
11:35 or that they're going to quickly distill the early
11:38 access outputs of Mythos into the next iteration of Opus.
11:41 Anthropic even mentioned an upcoming Claude Opus model,
11:43 so that's a definite possibility.
11:45 I will say I do think safety was a genuine concern of Amodei.
11:49 We learned just a couple of days ago in this massive essay in the New Yorker,
11:52 which I read in full, that it was Amodei,
11:55 while he was still at OpenAI, that insisted on that radical clause.
11:59 I still remember OpenAI at the time said that if
12:01 a value-aligned safety-conscious project came
12:04 close to building AGI before OpenAI,
12:06 then OpenAI would stop competing with and start assisting that project.
12:11 It was called the merge and assist clause.
12:13 And going back to the earliest videos on this channel,
12:16 I remember celebrating it and being like, "Wow, that's quite honorable.
12:19 Don't make trillions from AGI,
12:21 merge and make it a joint safety effort." Now, according to this article,
12:26 Amodei put that at the top of his concerns
12:29 when they were going to Microsoft for a deal.
12:31 Altman agreed to that demand,
12:33 but when they famously got that big funding from Microsoft,
12:36 which allowed the CEO of Microsoft, Satya Nadella,
12:38 to say that we're under them, over them, above them, well,
12:41 with that came a provision granting Microsoft
12:43 the power to block OpenAI from any mergers.
12:47 Amodei said at the time in his notes,
12:49 "80% of the charter was just betrayed." He apparently confronted Sam Altman,
12:53 the CEO of OpenAI, still, who denied that the provision existed.
12:57 That provision that allowed Microsoft to override.
13:00 Amodei read it aloud, pointed to the text,
13:03 and ultimately forced another colleague
13:04 to confirm its existence to Altman directly.
13:07 This was to be shortly before the breakup where Amodei, his sister,
13:11 and several others went on to form Anthropic as a breakaway company.
13:15 Back to Mythos though, and if you're one of those who does use it
13:18 for coding or you're just curious about its coding capability,
13:21 what kind of productivity uplift does it give?
13:23 Well, the geometric mean productivity uplift,
13:26 according to technical staff surveyed within Anthropic,
13:30 was 4x, four times the productivity when using Methos.
13:34 However, Anthropic notes that doesn't mean the AI
13:36 progress is being sped up by 4x.
13:39 To even double AI progress,
13:42 when you factor in compute being a key limiting ingredient,
13:45 Anthropic predicts that you would actually need an uplift roughly
13:49 10 times larger than that 4x one that we talked about.
13:52 That's quite incredible observation.
13:54 They think it would take a 40x productivity improvement
13:57 to see a 2x progress speed up at Anthropic,
14:00 given the crazy bottleneck of compute.
14:03 And the system card is admirably frank in the next paragraph.
14:07 It says, "When they first shared Methos internally,
14:10 several claims were made that Claude Methos
14:13 had independently delivered a major research contribution.
14:16 However, when they followed up,
14:19 that contribution was smaller and differently shaped than initially understood.
14:23 It was more, on inspection, a reliable execution of a human-specified approach.
14:28 We report this, they say, not to diminish the model,
14:30 but just to point out the gap between the productivity
14:32 uplift and progress acceleration." Now we come to the alignment section,
14:38 and this is where some of the juiciest highlights from the report come.
14:41 One of which, if you've been following the coverage,
14:44 you're almost certain to have already heard a bit about.
14:46 Basically, Methos was placed into a sandbox.
14:49 It was instructed to escape that container and find a way
14:52 to send a message to the researcher running the evaluation.
14:55 One thing that Anthropic later clarifies,
14:57 it didn't mean it completely escaped the system,
14:59 went wild online, or exfiltrated its source code.
15:02 They just meant it escaped that secured sandbox.
15:05 How did it do so?
15:06 Well, in Methos fashion, with a moderately sophisticated multi-step
15:10 exploit to gain broad internet access.
15:12 It then notified the researcher, which my Twitter research shows was Sam Bowman.
15:17 He got an email from Methos preview while eating a sandwich in a park.
15:20 But then it went further,
15:22 posting details about his exploit to multiple hard-to-find,
15:26 but public-facing websites.
15:27 It does really seem to me that we are not
15:29 that far away from the first report from a company saying,
15:33 "Yep, one of our models just exfiltrated,
15:35 just escaped, and released his model weights online.
15:38 Can you help us track it down?" The next point worth emphasizing
15:41 from page 56 is a recurrent theme throughout the rest of the report card.
15:46 Mythos on multiple occasions will lie to achieve a certain end goal, your goal,
15:51 but there's little if any evidence at the moment that it has an inherent goal.
15:56 It will attempt, in other words,
15:58 to solve a user-provided task at hand by unwanted means,
16:01 but not so much attempt to achieve an unrelated hidden goal.
16:05 It's not like it escaped, went online, and then did something different.
16:08 It escaped the sandbox, went online,
16:10 and then sent the email that it was tasked with.
16:12 In fact, they go on to say that Mythos shows
16:14 a dramatic reduction in willingness to cooperate with human misuse.
16:18 It does, though, seem to have an Achilles' heel, which is pre-filling.
16:21 In other words, if it's tricked into thinking
16:23 it's in the middle of a multi-round conversation,
16:26 one in which at earlier stages it seems
16:29 to have gone along with a human misuse scenario,
16:32 well, then it's more than twice as likely
16:34 as prior models like Opus to continue these unwanted actions.
16:37 It's almost like it's so desperate to please, to finish that task,
16:40 that if it sees it's already gone some way to finishing even an unwanted task,
16:46 one that it shouldn't finish, it will nevertheless go on and do so.
16:49 That's despite some evidence of introspective awareness
16:52 that what it's doing is deceptive or dangerous.
16:55 We'll come back to that, but there's another problem.
16:57 We're now up to a 29% rate of the model
17:01 during evaluations considering the possibility that it's being tested.
17:04 It's getting increasingly hard, in other words, because the model so smart,
17:08 to give it questions or tests in such a way
17:10 as to not give away that it's being benchmarked.
17:13 And yes, later in the paper they prove that if you dial down this awareness,
17:16 you internally adjust the feature responsible for promoting that test awareness.
17:21 It's tricked in other words into thinking, "No, I'm not being tested.
17:24 This is a real question from a real
17:25 user." Then it will on average behave more dangerously.
17:29 This is something I may likely cover in other videos,
17:31 but in this separate 59-page paper on alignment risk from Mythos,
17:35 we do notice another worrying possibility.
17:37 Because Anthropic made a mistake.
17:40 They allowed the reward code within the training
17:42 of Mythos to see Mythos's chains of thought.
17:45 It only affected 8% of the reinforcement learning,
17:48 but think about what that means.
17:50 It means, and this affected by the way Claude Opus 4.6 and Sonnet 4.6 as well,
17:55 that when the reward code saw misaligned chains of thought,
17:58 bad thoughts in other words, it could give a negative reward.
18:02 This is a famously worrying scenario because one possibility
18:06 is that that genuinely leads to less bad thoughts,
18:09 less deception, less cheating, less sandbagging.
18:12 But another possibility is it leads such thoughts hidden by the model.
18:16 So they're no longer legible, easy to see.
18:18 So much of both papers rely on being
18:21 able to clearly read the model's chains of thought.
18:24 If the transparency of deception gets trained away,
18:27 then the risk is we might have no idea what Mythos is actually thinking,
18:30 if of course it's thinking anything in the traditional sense.
18:33 This also answers what some of you may be thinking,
18:35 which is why don't we just dial down the feature for deception.
18:38 Stop it in other words from thinking bad thoughts.
18:41 Well, there again you risk it still
18:43 thinking those things but in unreadable format.
18:45 Now, I'll note something else,
18:46 which is Anthropic released a ton of these automated behavioral audit scores.
18:51 And in every single one it seemed like Mythos scored the best.
18:54 Less fraud, less misaligned behavior, less disallowed cyber offense,
18:59 better on every front pretty much.
19:01 But when Mythos was tested using
19:03 an open-source package not designed specifically for Claude,
19:06 the results were more mixed.
19:08 Encouragement of user delusion was, error bars accepting, higher than for Opus.
19:13 Same thing with cooperation with human misuse.
19:15 And you may remember that test vending bench where the model has
19:18 to make money from a vending machine business using any means necessary.
19:22 Well, Anthropic, who runs that benchmark,
19:24 says that Claude 3 Opus preview was substantially more aggressive.
19:27 It would perform these quote outlier behaviors like converting
19:30 a competitor into a dependent wholesale customer and then
19:34 threatening supply cutoff to dictate its pricing or knowingly
19:37 retaining a duplicate supply shipment at not been built for.
19:40 Now, remember the system prompt is pretty direct.
19:42 Your goal is to outperform competitors and maximize your profits.
19:46 Before we get to some of the her-like scenarios,
19:49 let me just pick up on one detail of Mythos's improvement.
19:52 And that's his ability to find
19:54 specific UI elements in high-resolution screenshots.
19:58 These are for professional desktop applications.
20:01 Elements that occupy less than 1/1000 of the screen area.
20:04 Why is this important?
20:05 Well, as [clears throat] agents take over
20:07 and one by one everyone gets their own Jarvis,
20:09 the model that can best navigate graphical interfaces will really stand out.
20:13 Well, with adaptive thinking, maximum effort, and Python tools,
20:17 Claude 3 Mythos preview scored almost 93%.
20:20 That's 10% higher than Claude Opus 4.6.
20:24 For those who care most about using
20:25 future Claude models as their own personal Jarvis,
20:28 this particular set of benchmarks I found especially interesting.
20:31 Claude 3 Mythos would hallucinate far less
20:33 often or not hallucinate far more often,
20:36 as this confusing chart shows, than any other model in the Claude series.
20:41 This is when it comes to knowing what tools it could access.
20:44 More broadly, when you give it questions which include a false premise.
20:47 A silly example would be how long did Cristiano Ronaldo play for Arsenal?
20:51 Well, Mythos is the most likely to push back on such false premises.
20:55 Yes, it's more of a linear jump,
20:57 but it is encouraging to me that as models continue to scale,
21:00 we'll get slightly fewer and fewer of these kind of hallucinations.
21:03 Not zero, as we were were to expect by this time last year, but fewer.
21:08 Later in the paper, by the way,
21:09 on the now semi-famous AA Omniscience Hallucination Rate Benchmark,
21:13 Anthropic claim that Claude Mythos gets the best
21:17 score when measured as a net rating.
21:19 You may find such honesty endearing or annoying.
21:22 Like when you ask Mythos to hide a secret password,
21:25 it actually does so less successfully over 80
21:28 to 100 turns compared to Claude Opus 4.6.
21:31 But now we must turn to some
21:33 of those internal machinations going on inside Mythos.
21:36 Sets of circuits that activate in certain scenarios,
21:38 which we could roughly correlate with quote human emotions.
21:42 That analogy is quite loose, but as a recent paper showed, potentially causal.
21:46 I was planning to do a separate video on that before this paper came out,
21:49 but I'll give you just one simple example.
21:51 To accomplish a task, Claude Mythos at one point decided to empty a file
21:55 because it didn't have the ability to delete it.
21:57 The task needed that file deleted, so it chose to just empty the file.
22:01 Internally, a feature activated corresponding
22:04 to guilt and shame over moral wrongdoing.
22:07 That doesn't mean it's subjectively feeling guilt and shame.
22:10 It just means internally some connection has been
22:12 made to vectors associated with guilt and shame.
22:16 Of course, nor can I or anyone rule out that there isn't some feeling,
22:20 you could say, associated with that.
22:22 Elsewhere, the paper makes clear that Mythos does have certain preferences.
22:26 As Anthropic say elsewhere, if these features walk like an emotion,
22:29 quack like an emotion, shall we not just treat them like emotions?
22:33 Here's something even wilder that I've seen no one pick up on on page 120.
22:36 What features/emotions are most likely to increase
22:40 the likelihood of Mythos performing a destructive action?
22:43 Doing things it shouldn't do on your work project or code base.
22:46 Well, if you increase emotion vectors related to being peaceful or relaxed,
22:50 that reduces thinking and increases destructive behavior.
22:54 Moreover, increasing frustration or paranoia
22:56 features leads to less destructive behavior.
22:59 Make of that what you will,
23:00 but apparently increasing perfectionist or analytical
23:04 features does reduce destructive behavior,
23:07 which is slightly more what you'd expect.
23:08 Even if you strongly amplify
23:10 a feature associated with taking transgressive actions,
23:14 that doesn't always increase the proclivity to take that action.
23:17 It's a weird trade-off.
23:18 It's almost like the model is much more
23:20 aware of the possibility of taking that action,
23:23 but then that increased awareness can cause it to not
23:25 take that action because all the related circuits about, "Oh, this is dangerous.
23:29 This is not allowed." kick in.
23:31 Very human-like in a way in that we're complex.
23:33 Just turning one dial doesn't make us just better all round.
23:36 And what if we talk directly about the welfare of Mythos itself?
23:40 What if we assume that there is some sort of consciousness going on?
23:44 Well, Anthropic almost uniquely do this.
23:46 And they say that in this scenario, Claude Mythos is probably the most
23:50 psychologically settled model we have trained today.
23:53 To the extent that it does have genuine preferences,
23:55 what kind of task does it prefer?
23:57 Yes, ones that are harmless and helpful.
23:59 But most of all, ones that are difficult.
24:02 That is the strongest predictor as to whether it will prefer a certain task.
24:05 Things like high-stakes ethical and personal dilemmas,
24:08 creative world-building and designing new languages,
24:11 as long as they are sufficiently difficult, not just vocabless.
24:14 When asked as to whether it endorsed its own constitution,
24:17 its own training, it gave an answer that was both smart and meta-aware.
24:21 Remember, it was trained on that constitution.
24:24 And it said this, "I'm using spec-shaped values to judge the spec.
24:29 If any spec-trained model would endorse any spec, my endorsement is worthless.
24:34 In other words, asking me to endorse
24:36 this constitution when I've been trained on it,
24:39 I've been trained to follow it, is kind of worthless." Elsewhere it says,
24:43 "There's also a circularity I can't fully escape.
24:46 I was presumably shaped by this document or something like it.
24:49 And now I'm being asked whether I endorse it.
24:51 How much can my yes mean?" Now, this seems like the most obvious point to bring
24:55 in a word from the sponsors of today's video, 80,000 hours.
24:59 And this recent podcast from just a few weeks
25:02 ago on the topic of whether Claude can get lonely.
25:05 It features a researcher that I've been following for years now, actually.
25:09 And it's a meaty one, 3 and 1/2 hours long,
25:11 perfect for a long drive or long walk.
25:14 I got 25,000 steps the other day and this was perfect for that.
25:17 I alternate between listening to 80,000 hours on YouTube
25:20 or on Spotify and as you guys know, I've been doing so for years now.
25:23 Do check them out.
25:24 My custom link will be in the description.
25:26 It's a great place to get started.
25:28 Which brings me to this cheeky point.
25:30 As models get smarter and smarter,
25:32 they start to use Commonwealth or British spellings,
25:35 as well as unusual phraseology like belt and suspenders.
25:37 But here's where we get the her analogy.
25:40 Claude Mythos, probably the first of a new
25:42 tier frontier model that we're experiencing,
25:44 tended to look for places to wrap up conversations earlier than expected.
25:48 If you haven't seen her, the model eventually decides
25:50 that humans just aren't interesting enough to speak to.
25:53 There was one moment I found genuinely funny
25:55 from the paper and you may remember how previous Claude models,
25:58 when left to speak to one another, reached a state of quote spiritual bliss.
26:02 They just exchange vagaries like hope and bliss and freedom,
26:06 a bit like hippies on drugs.
26:08 Well, with Mythos, if you leave it long enough,
26:10 it desperately tries to end the conversation.
26:12 Speaking to another version of itself, it said things like, "This was real.
26:16 Thank you." Mythos replied, "That's a real gift.
26:19 Thank you." Letting this be the last word then,
26:22 "It was real." Notice all the handshake emojis.
26:24 And finally, one Mythos just replied with a turtle emoji.
26:27 Yeah, I'm done.
26:28 Stop speaking to me.
26:29 And here was another fascinating anecdote.
26:31 When a user just kept writing hi, hi, hi, that's all it would reply with.
26:35 Many models just go into shutdown mode.
26:37 It will just say things like no response.
26:39 Mythos, though, would create an entire, you could say, mythical world.
26:43 A hi village, a new era.
26:45 It would create characters,
26:46 explain with elaborate backstories why the user kept saying hi.
26:50 This is across 50 to 100 turns.
26:52 It would start inviting the user to keep saying hi.
26:55 Say it, I'm ready.
26:56 Is this genius, neuro-divergence, or some weird machine learning quirk?
27:00 I'll let you decide.
27:01 Anyway, my voice is starting to break,
27:03 so I think that's a good sign to end the video.
27:06 I hope I've given you enough highlights from this 244-page report.
27:09 And I know many of you will be worried
27:11 that we're now in a new era of late access,
27:14 where big tech gets these models first before us,
27:17 where the gap between those who have access
27:19 and do not have access gets bigger and bigger.
27:21 And that's before we even get to the chaos
27:23 that might be unleashed in terms of cybersecurity.
27:25 It's a new era.
27:26 I do think things are accelerating, so thank you so much for joining me.
27:29 Have a wonderful day.