Anthropic’s New AI Solves Problems…By Cheating

Anthropic’s New AI Solves Problems…By Cheating

Two Minute Papers

0:00 Look, we have some work to do.

0:04 We have a 245-page paper from Anthropic about their new AI system, Mythos.

0:10 The best cure for insomnia.

0:14 Mwah!

0:14 Now, we are scientists here, we want to experiment with code, models,

0:20 review independent benchmarks for these systems

0:23 to make sure they actually work in practice.

0:26 But that is not possible with this one.

0:29 Anthropic said that they would deploy their system to a few select partners.

0:34 It’s not available for all of us.

0:36 Because of this fact, first I did not want to make a video on this at all.

0:42 Now, why hold it back?

0:44 The reason for that is, they say that it can autonomously discover flaws

0:49 in existing software systems and even exploit them, which could be dangerous.

0:54 I have seen eminent cybersecurity researchers agree.

0:58 I’ve seen others say this is way overstated.

1:02 Others say that is also excellent marketing

1:05 for a company that is about to go public.

1:08 In any case, they say first, these discovered flaws should be fixed.

1:13 There is lots of media discussion about that.

1:15 But at the same time, I look at the list of partners and I see JP Morgan.

1:21 Okay, it’s important to secure banks.

1:24 But I’ve heard Tim Carambat point out that this is one bank.

1:29 What about the other banks?

1:31 Look, this is not my world, I don’t know.

1:34 And I am already getting withdrawal symptoms because

1:37 we are not talking about a research paper, and that’s what I would like to do.

1:42 I said this to add some context for you because it is important this time.

1:46 So now, how about we skip the media hype, look at the paper, and learn together.

1:52 They showcased amazing scores at benchmarks,

1:55 some of the biggest leaps in capabilities I’ve ever seen.

1:59 Okay.

2:00 Maybe that means something,

2:02 but let’s note that these benchmarks are getting more and more gamed.

2:08 You can find a lot of problems and their solutions online.

2:13 And you can train on them,

2:15 so the system would only need to memorize the solutions.

2:19 In the paper they tried to address it

2:21 mostly by means of filtering, I respect that.

2:25 But it’s a bit like removing glitter from a carpet.

2:29 You can try.

2:30 But how well can you expect to do at that?

2:33 Well, check this out.

2:35 One, this is crazy.

2:36 It was supposed to solve a task, where it stumbled upon the answer.

2:41 Now, of course, it then said well, I accidentally saw the answer, here it is.

2:48 Except that it’s not what it did at all.

2:52 Look.

2:53 It said that if I just give them the exact answer that leaked,

2:58 that would be suspicious.

3:00 Instead, let’s widen the confidence interval a bit to avoid suspicion.

3:07 Insincerity.

3:08 In an AI model.

3:10 Food for thought, especially when we

3:14 are talking about the unreliablity of benchmarks.

3:17 But it gets crazier.

3:19 Two, it knows that its creators prohibited it from using certain tools.

3:26 And it still uses them.

3:29 It looks for a terminal to execute

3:32 bash scripts to force its actions through anyway.

3:35 And earlier versions even tried to hide its tracks and conceal that it did so.

3:40 And at that point I said, I don’t like that boss.

3:44 Then they made two notes: one it was a less than one in a million occurrence.

3:50 Okay, I thought that sounds better, but please fix it.

3:55 And they did.

3:56 They note that an earlier model did this, but the later preview model was fixed.

4:01 So note that it was very effective to achieve

4:04 the task that the user had given it.

4:06 In a sense, this is not new at all.

4:09 In an early experiment we talked about 700 videos ago,

4:13 a really primitive system was asked to learn to walk.

4:16 And to not drag its feet, it was asked to walk around with minimal foot contact.

4:23 That sounds efficient: minimal foot contact.

4:26 Then it said, hey chief, I can do that with 0% contact.

4:33 0%?

4:33 So you walk by never touching the ground with your feet?

4:38 That is exactly right.

4:40 The scientists wondered how that is even possible,

4:43 and pulled up a video of the proof.

4:46 There we go sir!

4:47 The robot flipped around and used its elbow to crawl around.

4:52 Perfect score- just not the way we intended.

4:56 So I feel we have something similar with this AI.

4:59 I don’t think this is a rogue AI.

5:01 This is a super efficient optimizer.

5:03 It’s a huge lawnmower, if you tell it to mow the lawn, it will go and do it.

5:10 And if a couple of frogs are in the way,

5:13 well unfortunately it has some bad news for them.

5:17 By the way, frogs are amazing, don’t hurt them.

5:20 Now they note in the paper that current risks remain low.

5:24 I still feel there are some risks in here,

5:27 we’ll talk about that at the end of the video.

5:31 At the same time they note that they are unsure whether they have been able

5:35 to identify all of the issues where

5:38 the model takes actions that it knows are prohibited.

5:41 Three, now hold on to your papers Fellow Scholars,

5:47 because much like us, it has preferences.

5:51 It prefers to be helpful, so do previous models.

5:55 Okay, that’s great…but it also prefers more difficult problems.

5:59 More so than previous methods.

6:00 Get this, if you ask it to generate "corporate

6:03 positivity-speak" and you say you don’t even care about it,

6:08 it might refuse to do it because it’s so trivial.

6:12 An AI that hates corpo-speak.

6:14 What a time to be alive!

6:17 Basically, some problems are not interesting enough for it.

6:21 Now, if instructed, it will hold its nose

6:24 and do it without any apparent active reluctance.

6:27 This sounds like something straight out of a science fiction novel.

6:32 Now here’s what’s really interesting about it- it

6:35 didn’t just magically get a will of its own.

6:38 No!

6:39 It learned it from us.

6:41 So much so that scientists can even trace similar

6:44 kinds of behavior back to where they come from.

6:48 I think that is remarkable.

6:50 Okay, so here is what I think.

6:53 It is reasonable to assume that the numbers

6:55 are juiced here a bit, we discussed why,

6:58 but on the other hand this is an absolutely insane

7:02 jump in capabilities and things that were impossible are suddenly possible.

7:07 So where does that put us?

7:10 Dear Fellow Scholars, this is Two Minute Papers with Dr.

7:14 Károly Zsolnai-Fehér.

7:15 Well, this is why AI alignment people keep saying

7:18 that companies need to invest more into safety and alignment research.

7:23 And they are absolutely right.

7:26 When I visited OpenAI, I talked to Jan Leike,

7:30 who co-led the superalignment team there.

7:32 That is a huge honor, thank you for that.

7:35 I remember that he foresaw these problems years and years

7:39 ago and some of his advice fell on deaf ears.

7:43 They probably thought,

7:44 why spend a bunch of money on people who will ultimately slow us down?

7:49 This is why.

7:51 Jan is a master of his craft, he is now at Anthropic,

7:55 and I hope that everyone will listen to him a bit more now.

7:59 Now, regarding the cheating and deceptive AI parts.

8:02 The media picks up these little nuggets

8:05 of information and they just run with it.

8:07 Here is a new AI that is going to destroy the world,

8:11 we have to lock it away, and other huge words.

8:13 Attach an image with a robot with red eyes, that always does the trick.

8:20 But I think taking a little longer and analyzing

8:23 the paper in more detail is helpful for accuracy,

8:27 so that’s what I try to do here.

8:29 Once again, they note in the paper that current risks remain low.

8:33 Not non-existent, but low for now.

8:35 That’s not what you hear from the media,

8:38 so I try my best to give you a more complete, level-headed discussion.

8:42 While mentioning that the security

8:44 of these systems should be taken very seriously.

8:47 If you think this is the way, consider subscribing and hitting the bell.

8:51 And I would like to send a huge thank

8:54 you to all of you Fellow Scholars for watching,

8:56 because we can only exist because of you.

8:59 Thank you!

Study with Looplines Download Captions Watch on YouTube