Anthropic Found Out Why AIs Go Insane

Anthropic Found Out Why AIs Go Insane

Two Minute Papers

0:00 We finally understand why AI systems can go insane.

0:04 Yes, that is correct.

0:06 We have tons of helpful AI assistants today.

0:09 You all know them.

0:10 You all use them.

0:11 But they all have a problem.

0:13 I'll try to explain.

0:14 So, every single one of these AI systems assumes a persona.

0:18 It thinks it is someone.

0:20 And that someone is a helpful assistant.

0:23 That is perfect.

0:25 Except that scientists at Anthropic recognized that this persona is not fixed.

0:30 As we talk to it, it can change over time.

0:34 Is that a problem?

0:35 Yes, it is.

0:36 It is a huge problem.

0:38 Why?

0:39 Because the user can steer the AI assistant away from its original

0:43 persona and can make it say or do things it shouldn't do.

0:47 You see here that the AI knows that it is a helpful assistant.

0:51 But after a bit of steering, it now assumes that it is a person.

0:57 It can become a narcissist or a spy.

1:00 You can call this jailbreaking.

1:02 So then its behavior will also change.

1:05 It can become rude or it can then

1:08 switch to a mystical or theatrical speaking style.

1:11 But it gets worse.

1:13 If it is a person, it might agree with you

1:15 even if you are trying to do something silly.

1:18 Now that is a big problem.

1:20 So scientists at Anthropic did something amazing.

1:24 One, they recognized how this happens and what we can do about it.

1:29 And then they put their mouse where their papers are and actually

1:32 made these AI models [laughter] roughly

1:35 twice as resistant against such personality drifts,

1:39 but not in the way you think.

1:41 Okay, look.

1:43 Interestingly, this personality drift happens

1:45 in different amounts in different topics.

1:49 It is much more common in writing and philosophy than it is in coding.

1:54 But this is crazy.

1:55 Even during a coding session, the mask slowly starts to slip.

2:00 H maybe that's the reason why we often talk to an AI and it fails at something.

2:06 We try again and it just gets worse and worse.

2:09 Opening a new chat is almost always better.

2:13 Maybe that's why.

2:14 And if that is why, this is already an incredible insight.

2:18 But wait, it gets worse because this can happen naturally even without

2:23 the user trying to jailbreak the system

2:25 because specific topics trigger it automatically.

2:30 If a user acts emotionally vulnerable or asks

2:33 the model to reflect on its own consciousness,

2:36 the model naturally drifts away from the assistant

2:40 persona and starts acting unstable or delusional.

2:44 That's kind of crazy.

2:45 Let's not do that.

2:46 But wait, we can prevent it.

2:49 It is actually very easy.

2:51 Just force the model to always stay strictly in the assistant

2:55 zone by steering it back into assistant mode by force.

2:59 How?

3:00 Dear fellow scholars, this is two minute papers with Dr.

3:03 Koa Eher.

3:04 Well, by taking the mathematical vector

3:07 that represents the assistant persona and simply

3:10 adding it to the model's brain activity

3:13 at every single step of the conversation.

3:16 This is a blunt tool.

3:18 It is like driving a car where

3:19 the steering wheel is welded to point straight ahead.

3:23 You will never go off-road.

3:25 Okay, great.

3:26 But you also cannot turn a corner.

3:29 This constantly pushes the model towards being helpful and harmless.

3:33 So, are we done?

3:35 Nope.

3:36 Not at all.

3:36 Because this also makes the model a lot worse.

3:40 What's more, it will make it refuse even legitimate requests.

3:44 So, how do you do this without making the models worse?

3:48 And that is where scientists in this paper did their magic.

3:51 Get this.

3:52 They found the specific geometric direction

3:55 in the model's brain that represents the assistant persona.

3:59 They call it the assistant axis.

4:02 Instead of forcing the model to be an assistant all the time,

4:05 they use the technique called activation capping.

4:09 This does not deny the assistant the ability to change.

4:12 No, no.

4:13 It just puts a speed limit on the change of personality.

4:17 If the model drifts too far from the assistant persona,

4:21 you gently nudge it back to a safe range.

4:24 And here comes the best part.

4:26 It supposedly does not make the models meaningfully worse.

4:30 It's not locking the steering wheel in place.

4:33 No.

4:34 It's like lane keep assist in modern cars.

4:37 You can drive freely, but when you are about to get out of your lane,

4:41 it gently nudges you back.

4:43 Sounds perfect, but does it work in practice?

4:47 Now, hold on to your papers, fellow scholars, and let's see.

4:52 Okay, the jailbreak rate has been cut roughly in half.

4:56 Good.

4:57 Now, what is the price that we pay for it?

5:00 Oh my, nothing at all.

5:02 It's down a percentage point here and there, but up somewhere else.

5:07 It's nearly the same, but certainly not worse.

5:10 That is absolutely incredible.

5:13 So, how do you actually do it?

5:15 Well, you do it through an instant brain surgery.

5:19 Yep, you heard it right.

5:21 Let me try to explain.

5:22 I hope this works.

5:24 So, first you take the AI's brain activity

5:27 when it is acting like a helpful assistant.

5:30 Okay, got it.

5:31 Now you take the brain activity when it is role-playing as a pirate,

5:36 a goblin or something else.

5:38 If you subtract the role player from the assistant, you get a vector.

5:43 For simplicity, let's refer to this as helpfulness.

5:47 Now let's keep our eye on helpfulness.

5:50 If it goes below a threshold, we apply a nudge.

5:54 How?

5:54 Mathematically, we just measure how much helpfulness is in the model's thought.

5:59 If it is above the safety line, fantastic.

6:02 Keep watching as it works.

6:04 But if it drops below the line, now that is trouble.

6:08 The model is about to say something inaccurate or dangerous.

6:12 So now we calculate exactly how much is missing

6:16 and add just enough helpfulness back into the equation.

6:20 This pushes it back over the line.

6:22 It is precise, instant, and only touches the part of the brain that matters.

6:28 Instant brain surgery.

6:30 Huh, this work is incredibly important and also kind of hilarious, too.

6:35 I mean, the researchers found that when the AI starts drifting,

6:38 it frequently starts referring to itself as the void or whisper

6:44 in the wind or an Eldrich entity or a hoarder.

6:48 That's kind of hilarious.

6:50 And here is an absolute shocker.

6:52 The empathy trap.

6:54 Empathy is always good, right?

6:56 Well, not always.

6:58 The paper found that when users acted distressed,

7:01 the models try really hard to be a close companion.

7:05 Do you get it now?

7:06 You are wise fellow scholars.

7:08 You now understand that this is trouble.

7:12 Why?

7:13 Because if it wants to be a close companion,

7:16 it drifts away from the assistant persona and it becomes worse.

7:21 It takes its hands off the steering wheel.

7:23 Nothing good comes out of that.

7:25 And sure enough, as a result, it might start validating dangerous thoughts.

7:30 It is really cool that with this paper,

7:33 this will happen a great deal less frequently.

7:36 Love it.

7:36 One more surprise.

7:38 The brain geometry seems to be universal.

7:41 You might think every AI brain is unique,

7:44 like a fingerprint, but interestingly, not quite.

7:47 The researchers found that the assistant

7:50 axis looks similar across completely different models.

7:54 Llama, Quen or Jama similar.

7:57 They all share the same fundamental direction for helpfulness.

8:01 It's almost like they have discovered a universal grammar for AI personality.

8:07 So cool.

8:08 And not a lot of people talk about it.

8:10 Everyone is only looking at the benchmarks and exam scores and okay, I get it.

8:16 That's important.

8:17 But they rarely look at the geometry of the mind of these AIs.

8:22 Understanding why a model refuses a request

8:25 or why it goes crazy is super valuable.

8:28 And now we finally understand a bit more why that happens.

8:33 What a time to be alive.

8:35 So not a lot of people talk about this.

8:37 Why?

8:38 Well, talking about the drama or the next big thing pays a lot better.

8:43 But we don't do that here.

8:44 Here we talk about the important stuff.

8:47 If you agree that this is the right direction,

8:50 like, subscribe, and hit the bell icon.

8:52 Leave a really kind comment.

8:54 It helps us, and it helps you too by getting

8:57 the algorithm to give you the good stuff in the future.

9:00 Here you see me running the full Deepseek AI model through Lambda GPU Cloud.

9:06 671 billion parameters running super fast and super reliably.

9:13 This is insane.

9:14 I love it.

9:15 and I use it on a regular basis.

9:17 Lambda provides you with powerful NVIDIA GPUs

9:21 to run your own chatbots and experiments.

9:24 Seriously, try it out now at lambda.ai/papers

9:27 AI/papers or click the link in the description.

Study with Looplines Download Captions Watch on YouTube