Anthropic Found Out Why AIs Go Insane
Two Minute Papers
0:00 We finally understand why AI systems can go insane.
0:04 Yes, that is correct.
0:06 We have tons of helpful AI assistants today.
0:09 You all know them.
0:10 You all use them.
0:11 But they all have a problem.
0:13 I'll try to explain.
0:14 So, every single one of these AI systems assumes a persona.
0:18 It thinks it is someone.
0:20 And that someone is a helpful assistant.
0:23 That is perfect.
0:25 Except that scientists at Anthropic recognized that this persona is not fixed.
0:30 As we talk to it, it can change over time.
0:34 Is that a problem?
0:35 Yes, it is.
0:36 It is a huge problem.
0:38 Why?
0:39 Because the user can steer the AI assistant away from its original
0:43 persona and can make it say or do things it shouldn't do.
0:47 You see here that the AI knows that it is a helpful assistant.
0:51 But after a bit of steering, it now assumes that it is a person.
0:57 It can become a narcissist or a spy.
1:00 You can call this jailbreaking.
1:02 So then its behavior will also change.
1:05 It can become rude or it can then
1:08 switch to a mystical or theatrical speaking style.
1:11 But it gets worse.
1:13 If it is a person, it might agree with you
1:15 even if you are trying to do something silly.
1:18 Now that is a big problem.
1:20 So scientists at Anthropic did something amazing.
1:24 One, they recognized how this happens and what we can do about it.
1:29 And then they put their mouse where their papers are and actually
1:32 made these AI models [laughter] roughly
1:35 twice as resistant against such personality drifts,
1:39 but not in the way you think.
1:41 Okay, look.
1:43 Interestingly, this personality drift happens
1:45 in different amounts in different topics.
1:49 It is much more common in writing and philosophy than it is in coding.
1:54 But this is crazy.
1:55 Even during a coding session, the mask slowly starts to slip.
2:00 H maybe that's the reason why we often talk to an AI and it fails at something.
2:06 We try again and it just gets worse and worse.
2:09 Opening a new chat is almost always better.
2:13 Maybe that's why.
2:14 And if that is why, this is already an incredible insight.
2:18 But wait, it gets worse because this can happen naturally even without
2:23 the user trying to jailbreak the system
2:25 because specific topics trigger it automatically.
2:30 If a user acts emotionally vulnerable or asks
2:33 the model to reflect on its own consciousness,
2:36 the model naturally drifts away from the assistant
2:40 persona and starts acting unstable or delusional.
2:44 That's kind of crazy.
2:45 Let's not do that.
2:46 But wait, we can prevent it.
2:49 It is actually very easy.
2:51 Just force the model to always stay strictly in the assistant
2:55 zone by steering it back into assistant mode by force.
2:59 How?
3:00 Dear fellow scholars, this is two minute papers with Dr.
3:03 Koa Eher.
3:04 Well, by taking the mathematical vector
3:07 that represents the assistant persona and simply
3:10 adding it to the model's brain activity
3:13 at every single step of the conversation.
3:16 This is a blunt tool.
3:18 It is like driving a car where
3:19 the steering wheel is welded to point straight ahead.
3:23 You will never go off-road.
3:25 Okay, great.
3:26 But you also cannot turn a corner.
3:29 This constantly pushes the model towards being helpful and harmless.
3:33 So, are we done?
3:35 Nope.
3:36 Not at all.
3:36 Because this also makes the model a lot worse.
3:40 What's more, it will make it refuse even legitimate requests.
3:44 So, how do you do this without making the models worse?
3:48 And that is where scientists in this paper did their magic.
3:51 Get this.
3:52 They found the specific geometric direction
3:55 in the model's brain that represents the assistant persona.
3:59 They call it the assistant axis.
4:02 Instead of forcing the model to be an assistant all the time,
4:05 they use the technique called activation capping.
4:09 This does not deny the assistant the ability to change.
4:12 No, no.
4:13 It just puts a speed limit on the change of personality.
4:17 If the model drifts too far from the assistant persona,
4:21 you gently nudge it back to a safe range.
4:24 And here comes the best part.
4:26 It supposedly does not make the models meaningfully worse.
4:30 It's not locking the steering wheel in place.
4:33 No.
4:34 It's like lane keep assist in modern cars.
4:37 You can drive freely, but when you are about to get out of your lane,
4:41 it gently nudges you back.
4:43 Sounds perfect, but does it work in practice?
4:47 Now, hold on to your papers, fellow scholars, and let's see.
4:52 Okay, the jailbreak rate has been cut roughly in half.
4:56 Good.
4:57 Now, what is the price that we pay for it?
5:00 Oh my, nothing at all.
5:02 It's down a percentage point here and there, but up somewhere else.
5:07 It's nearly the same, but certainly not worse.
5:10 That is absolutely incredible.
5:13 So, how do you actually do it?
5:15 Well, you do it through an instant brain surgery.
5:19 Yep, you heard it right.
5:21 Let me try to explain.
5:22 I hope this works.
5:24 So, first you take the AI's brain activity
5:27 when it is acting like a helpful assistant.
5:30 Okay, got it.
5:31 Now you take the brain activity when it is role-playing as a pirate,
5:36 a goblin or something else.
5:38 If you subtract the role player from the assistant, you get a vector.
5:43 For simplicity, let's refer to this as helpfulness.
5:47 Now let's keep our eye on helpfulness.
5:50 If it goes below a threshold, we apply a nudge.
5:54 How?
5:54 Mathematically, we just measure how much helpfulness is in the model's thought.
5:59 If it is above the safety line, fantastic.
6:02 Keep watching as it works.
6:04 But if it drops below the line, now that is trouble.
6:08 The model is about to say something inaccurate or dangerous.
6:12 So now we calculate exactly how much is missing
6:16 and add just enough helpfulness back into the equation.
6:20 This pushes it back over the line.
6:22 It is precise, instant, and only touches the part of the brain that matters.
6:28 Instant brain surgery.
6:30 Huh, this work is incredibly important and also kind of hilarious, too.
6:35 I mean, the researchers found that when the AI starts drifting,
6:38 it frequently starts referring to itself as the void or whisper
6:44 in the wind or an Eldrich entity or a hoarder.
6:48 That's kind of hilarious.
6:50 And here is an absolute shocker.
6:52 The empathy trap.
6:54 Empathy is always good, right?
6:56 Well, not always.
6:58 The paper found that when users acted distressed,
7:01 the models try really hard to be a close companion.
7:05 Do you get it now?
7:06 You are wise fellow scholars.
7:08 You now understand that this is trouble.
7:12 Why?
7:13 Because if it wants to be a close companion,
7:16 it drifts away from the assistant persona and it becomes worse.
7:21 It takes its hands off the steering wheel.
7:23 Nothing good comes out of that.
7:25 And sure enough, as a result, it might start validating dangerous thoughts.
7:30 It is really cool that with this paper,
7:33 this will happen a great deal less frequently.
7:36 Love it.
7:36 One more surprise.
7:38 The brain geometry seems to be universal.
7:41 You might think every AI brain is unique,
7:44 like a fingerprint, but interestingly, not quite.
7:47 The researchers found that the assistant
7:50 axis looks similar across completely different models.
7:54 Llama, Quen or Jama similar.
7:57 They all share the same fundamental direction for helpfulness.
8:01 It's almost like they have discovered a universal grammar for AI personality.
8:07 So cool.
8:08 And not a lot of people talk about it.
8:10 Everyone is only looking at the benchmarks and exam scores and okay, I get it.
8:16 That's important.
8:17 But they rarely look at the geometry of the mind of these AIs.
8:22 Understanding why a model refuses a request
8:25 or why it goes crazy is super valuable.
8:28 And now we finally understand a bit more why that happens.
8:33 What a time to be alive.
8:35 So not a lot of people talk about this.
8:37 Why?
8:38 Well, talking about the drama or the next big thing pays a lot better.
8:43 But we don't do that here.
8:44 Here we talk about the important stuff.
8:47 If you agree that this is the right direction,
8:50 like, subscribe, and hit the bell icon.
8:52 Leave a really kind comment.
8:54 It helps us, and it helps you too by getting
8:57 the algorithm to give you the good stuff in the future.
9:00 Here you see me running the full Deepseek AI model through Lambda GPU Cloud.
9:06 671 billion parameters running super fast and super reliably.
9:13 This is insane.
9:14 I love it.
9:15 and I use it on a regular basis.
9:17 Lambda provides you with powerful NVIDIA GPUs
9:21 to run your own chatbots and experiments.
9:24 Seriously, try it out now at lambda.ai/papers
9:27 AI/papers or click the link in the description.