Anthropic’s New AI Solves Problems…By Cheating
Two Minute Papers
0:00 Look, we have some work to do.
0:04 We have a 245-page paper from Anthropic about their new AI system, Mythos.
0:10 The best cure for insomnia.
0:14 Mwah!
0:14 Now, we are scientists here, we want to experiment with code, models,
0:20 review independent benchmarks for these systems
0:23 to make sure they actually work in practice.
0:26 But that is not possible with this one.
0:29 Anthropic said that they would deploy their system to a few select partners.
0:34 It’s not available for all of us.
0:36 Because of this fact, first I did not want to make a video on this at all.
0:42 Now, why hold it back?
0:44 The reason for that is, they say that it can autonomously discover flaws
0:49 in existing software systems and even exploit them, which could be dangerous.
0:54 I have seen eminent cybersecurity researchers agree.
0:58 I’ve seen others say this is way overstated.
1:02 Others say that is also excellent marketing
1:05 for a company that is about to go public.
1:08 In any case, they say first, these discovered flaws should be fixed.
1:13 There is lots of media discussion about that.
1:15 But at the same time, I look at the list of partners and I see JP Morgan.
1:21 Okay, it’s important to secure banks.
1:24 But I’ve heard Tim Carambat point out that this is one bank.
1:29 What about the other banks?
1:31 Look, this is not my world, I don’t know.
1:34 And I am already getting withdrawal symptoms because
1:37 we are not talking about a research paper, and that’s what I would like to do.
1:42 I said this to add some context for you because it is important this time.
1:46 So now, how about we skip the media hype, look at the paper, and learn together.
1:52 They showcased amazing scores at benchmarks,
1:55 some of the biggest leaps in capabilities I’ve ever seen.
1:59 Okay.
2:00 Maybe that means something,
2:02 but let’s note that these benchmarks are getting more and more gamed.
2:08 You can find a lot of problems and their solutions online.
2:13 And you can train on them,
2:15 so the system would only need to memorize the solutions.
2:19 In the paper they tried to address it
2:21 mostly by means of filtering, I respect that.
2:25 But it’s a bit like removing glitter from a carpet.
2:29 You can try.
2:30 But how well can you expect to do at that?
2:33 Well, check this out.
2:35 One, this is crazy.
2:36 It was supposed to solve a task, where it stumbled upon the answer.
2:41 Now, of course, it then said well, I accidentally saw the answer, here it is.
2:48 Except that it’s not what it did at all.
2:52 Look.
2:53 It said that if I just give them the exact answer that leaked,
2:58 that would be suspicious.
3:00 Instead, let’s widen the confidence interval a bit to avoid suspicion.
3:07 Insincerity.
3:08 In an AI model.
3:10 Food for thought, especially when we
3:14 are talking about the unreliablity of benchmarks.
3:17 But it gets crazier.
3:19 Two, it knows that its creators prohibited it from using certain tools.
3:26 And it still uses them.
3:29 It looks for a terminal to execute
3:32 bash scripts to force its actions through anyway.
3:35 And earlier versions even tried to hide its tracks and conceal that it did so.
3:40 And at that point I said, I don’t like that boss.
3:44 Then they made two notes: one it was a less than one in a million occurrence.
3:50 Okay, I thought that sounds better, but please fix it.
3:55 And they did.
3:56 They note that an earlier model did this, but the later preview model was fixed.
4:01 So note that it was very effective to achieve
4:04 the task that the user had given it.
4:06 In a sense, this is not new at all.
4:09 In an early experiment we talked about 700 videos ago,
4:13 a really primitive system was asked to learn to walk.
4:16 And to not drag its feet, it was asked to walk around with minimal foot contact.
4:23 That sounds efficient: minimal foot contact.
4:26 Then it said, hey chief, I can do that with 0% contact.
4:33 0%?
4:33 So you walk by never touching the ground with your feet?
4:38 That is exactly right.
4:40 The scientists wondered how that is even possible,
4:43 and pulled up a video of the proof.
4:46 There we go sir!
4:47 The robot flipped around and used its elbow to crawl around.
4:52 Perfect score- just not the way we intended.
4:56 So I feel we have something similar with this AI.
4:59 I don’t think this is a rogue AI.
5:01 This is a super efficient optimizer.
5:03 It’s a huge lawnmower, if you tell it to mow the lawn, it will go and do it.
5:10 And if a couple of frogs are in the way,
5:13 well unfortunately it has some bad news for them.
5:17 By the way, frogs are amazing, don’t hurt them.
5:20 Now they note in the paper that current risks remain low.
5:24 I still feel there are some risks in here,
5:27 we’ll talk about that at the end of the video.
5:31 At the same time they note that they are unsure whether they have been able
5:35 to identify all of the issues where
5:38 the model takes actions that it knows are prohibited.
5:41 Three, now hold on to your papers Fellow Scholars,
5:47 because much like us, it has preferences.
5:51 It prefers to be helpful, so do previous models.
5:55 Okay, that’s great…but it also prefers more difficult problems.
5:59 More so than previous methods.
6:00 Get this, if you ask it to generate "corporate
6:03 positivity-speak" and you say you don’t even care about it,
6:08 it might refuse to do it because it’s so trivial.
6:12 An AI that hates corpo-speak.
6:14 What a time to be alive!
6:17 Basically, some problems are not interesting enough for it.
6:21 Now, if instructed, it will hold its nose
6:24 and do it without any apparent active reluctance.
6:27 This sounds like something straight out of a science fiction novel.
6:32 Now here’s what’s really interesting about it- it
6:35 didn’t just magically get a will of its own.
6:38 No!
6:39 It learned it from us.
6:41 So much so that scientists can even trace similar
6:44 kinds of behavior back to where they come from.
6:48 I think that is remarkable.
6:50 Okay, so here is what I think.
6:53 It is reasonable to assume that the numbers
6:55 are juiced here a bit, we discussed why,
6:58 but on the other hand this is an absolutely insane
7:02 jump in capabilities and things that were impossible are suddenly possible.
7:07 So where does that put us?
7:10 Dear Fellow Scholars, this is Two Minute Papers with Dr.
7:14 Károly Zsolnai-Fehér.
7:15 Well, this is why AI alignment people keep saying
7:18 that companies need to invest more into safety and alignment research.
7:23 And they are absolutely right.
7:26 When I visited OpenAI, I talked to Jan Leike,
7:30 who co-led the superalignment team there.
7:32 That is a huge honor, thank you for that.
7:35 I remember that he foresaw these problems years and years
7:39 ago and some of his advice fell on deaf ears.
7:43 They probably thought,
7:44 why spend a bunch of money on people who will ultimately slow us down?
7:49 This is why.
7:51 Jan is a master of his craft, he is now at Anthropic,
7:55 and I hope that everyone will listen to him a bit more now.
7:59 Now, regarding the cheating and deceptive AI parts.
8:02 The media picks up these little nuggets
8:05 of information and they just run with it.
8:07 Here is a new AI that is going to destroy the world,
8:11 we have to lock it away, and other huge words.
8:13 Attach an image with a robot with red eyes, that always does the trick.
8:20 But I think taking a little longer and analyzing
8:23 the paper in more detail is helpful for accuracy,
8:27 so that’s what I try to do here.
8:29 Once again, they note in the paper that current risks remain low.
8:33 Not non-existent, but low for now.
8:35 That’s not what you hear from the media,
8:38 so I try my best to give you a more complete, level-headed discussion.
8:42 While mentioning that the security
8:44 of these systems should be taken very seriously.
8:47 If you think this is the way, consider subscribing and hitting the bell.
8:51 And I would like to send a huge thank
8:54 you to all of you Fellow Scholars for watching,
8:56 because we can only exist because of you.
8:59 Thank you!