State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

Lex Fridman

0:00 The following is a conversation all

0:02 about the state-of-the-art in artificial intelligence,

0:04 including some of the exciting technical breakthroughs and developments

0:08 in AI that happened over the past year,

0:11 and some of the interesting things we think might happen this upcoming year.

0:16 At times, it does get super technical,

0:19 but we do try to make sure that it remains

0:22 accessible to folks outside the field without ever dumbing it down.

0:26 It is a great honor and pleasure to be able to do

0:30 this kind of episode with two of my favorite people in the AI community,

0:35 Sebastian Raschka and Nathan Lambert.

0:38 They are both widely respected machine learning researchers

0:42 and engineers who also happen to be great communicators,

0:46 educators, writers, and X posters.

0:49 Sebastian is the author of two books

0:53 I highly recommend for beginners and experts alike.

0:56 First is Build a Large Language Model

0:59 from Scratch and Build a Reasoning Model from Scratch.

1:04 I truly believe in the machine learning world,

1:08 the best way to learn and understand

1:11 something is to build it yourself from scratch.

1:15 Nathan is the post-training lead at the Allen Institute for AI,

1:21 author of the definitive book on Reinforcement Learning from Human Feedback.

1:26 Both of them have great X accounts, great Substacks.

1:30 Sebastian has courses on YouTube, Nathan has a podcast.

1:34 And everyone should absolutely follow all of those.

1:37 those.

1:38 This is the Lex Fridman podcast.

1:40 To support it, please check out our sponsors in the description,

1:43 where you can also find links to contact me,

1:47 ask questions, get feedback, and so on.

1:50 And now, dear friends, here's Sebastian Raschka and Nathan Lambert.

1:57 So I think one useful lens to look

1:59 at all this through is the so-called DeepSeek moment.

2:03 This happened about a year ago in January 2025,

2:07 when the open-weight Chinese company DeepSeek released DeepSeek R1,

2:11 that I think it's fair

2:14 to say surprised everyone with near-state-of-the-art performance,

2:17 with allegedly much less compute for much cheaper.

2:22 And from then to today, the AI competition has gotten insane,

2:28 both on the research and product level.

2:31 It's just been accelerating.

2:32 discuss all of this today,

2:34 and maybe let's start with some spicy questions if we can.

2:38 Who's winning at the international level?

2:40 Would you say it's the set of companies in China

2:43 or the set of companies in the United States?

2:46 And Sebastian, Nathan, it's good to see you guys.

2:50 guys.

2:50 So Sebastian, who do you think is winning?

2:53 Winning is a very broad term.

2:57 I would say you mentioned the DeepSeek moment,

2:59 and I think DeepSeek is winning the hearts of the people

3:02 who work on open-weight models because they share these as open models.

3:06 Winning, I think, has multiple timescales to it.

3:09 We have today, we have next year, we have in 10 years.

3:13 One thing I know for sure is that I don't think nowadays, in 2026,

3:18 that there will be any company that has access

3:22 to technology that no other company has access to.

3:26 That is mainly because researchers are frequently changing jobs and labs.

3:32 They rotate.

3:32 I don't think there will be a clear winner in terms of technology access.

3:36 However, I do think there will be,

3:39 The differentiating factor will be budget and hardware constraints.

3:43 I don't think the ideas will be proprietary,

3:46 but rather the resources needed to implement them.

3:52 I don't see currently a winner-take-all scenario.

3:55 I can't see that.

3:57 At the moment.

3:59 Nathan, what do you think?

4:00 You see the labs put different energy into what they're trying to do,

4:04 and I think to demarcate the point in time when we're recording this, the hype

4:08 over Anthropic's Claude Opus 4.5 model

4:11 has been absolutely insane, which is just...

4:13 I mean, I've used it and built stuff in the last few weeks, and it's...

4:17 it's almost gotten to the point where it feels like

4:19 a bit of a meme in terms of the hype.

4:21 And it's kind of funny because this is very organic,

4:24 and then if we go back a few months ago,

4:26 we can see the release date and the notes,

4:29 as Gemini 3 from Google got released, and it seemed like the marketing and just,

4:34 like, wow factor of that release was super high.

4:37 But then at the end of November,

4:39 Claude Opus 4.5 was released and the hype has been growing,

4:42 but Gemini 3 was before this.

4:43 And it kind of feels like people don't really talk about it as much,

4:46 even though when it came out, everybody was like,

4:48 this is Gemini's moment to retake Google's structural advantages in AI.

4:53 And Gemini 3 is a fantastic model, and I still use it.

4:56 It's just kind of differentiation is lower.

4:59 And I agree with Sebastian;

5:01 what you're saying with all these, the idea space is very fluid,

5:05 but culturally Anthropic is known for betting very hard on code,

5:09 which is the Claude Code thing, is working out for them right now.

5:12 So I think that even if the ideas flow pretty freely,

5:15 so much of this is bottlenecked

5:16 by human effort and the culture of organizations,

5:19 where Anthropic seems to at least be presenting as the least chaotic.

5:23 It's a bit of an advantage, if they can keep doing that for a while.

5:27 But on the other side of things, there's a lot of ominous technology from China

5:31 where there's way more labs than DeepSeek.

5:34 So DeepSeek kicked off a movement within China,

5:37 I say kind of similar to how ChatGPT kicked off

5:40 a movement in the US where everything had a chatbot.

5:43 There's now tons of tech companies in China

5:46 that are releasing very strong frontier open-weight models,

5:48 to the point where I would say that DeepSeek is kind

5:51 of losing its crown as the preeminent open model maker in China,

5:54 and the likes of Z.ai with their GLM models, Minimax's models,

6:00 Kimi Moonshot, especially in the last few months, has shown more brightly.

6:04 The new DeepSeek models are still very strong, but that's kind of a...

6:08 it could look back as a big narrative point

6:10 where in 2025 DeepSeek came and it provided this platform

6:14 for way more Chinese companies that are releasing these fantastic

6:17 models to kind of have this new type of operation.

6:20 So these models from these Chinese companies are open-weights,

6:23 and depending on this trajectory of business

6:25 models that these American companies are doing, they could be at risk.

6:29 But currently, a lot of people are paying for AI software in the US,

6:33 and historically in China and other parts of the world,

6:36 people don't pay a lot for software.

6:38 So some of these models like DeepSeek have

6:40 the love of the people because they are open-weight.

6:42 How long do you think the Chinese companies keep releasing open-weight models?

6:47 I would say for a few years.

6:49 I think that, like in the US, there's not a clear business model for it.

6:53 I have been writing about open models for a while,

6:55 and these Chinese companies have realized it.

6:57 So I get inbound from some of them.

6:59 And they're smart and realize the same constraints:

7:01 a lot of top US tech companies and other IT companies

7:04 won't pay for an API subscription to Chinese companies for security concerns.

7:08 This has been a long-standing habit in tech,

7:11 and the people at these companies then see open weight models as an ability

7:16 to influence and take part of a huge growing AI expenditure market in the US.

7:20 And they're very realistic about this, and it's working for them.

7:24 I think that the government will see that that is building

7:27 a lot of influence internationally in terms of uptake of the technology,

7:31 so there's going to be a lot of incentives to keep it going.

7:34 But building these models and doing the research is very expensive,

7:37 so at some point, I expect consolidation.

7:39 But I don't expect that to be a story of 2026,

7:42 where there will be more open model

7:45 builders throughout 2026 than there were in 2025.

7:47 And a lot of the notable ones will be in China.

7:50 You were going to say something?

7:51 Yes.

7:52 You mentioned DeepSeek losing its crown.

7:54 I do think to some extent, yes, but we also have to consider though,

8:00 they are still, I would say, slightly ahead.

8:02 And the other ones—it's not that DeepSeek got worse,

8:04 it's just that the other ones are using the ideas from DeepSeek.

8:08 For example, you mentioned Kimi—same architecture, they're training it.

8:11 And then again, we have this leapfrogging where they might be at some

8:14 point in time a bit better because they have the more recent model.

8:17 And I think this comes back to the fact that there won't be a clear winner.

8:22 It will just be like that: one person releases something,

8:25 the other one comes in, and the most

8:27 recent model is probably always the best model.

8:30 Yeah.

8:30 We'll also see the Chinese companies have different incentives.

8:33 Like, DeepSeek is very secretive,

8:35 whereas some of these startups are like the MiniMaxs and Z.ais of the world.

8:40 Those two literally have filed IPO paperwork,

8:42 and they're trying to get Western mindshare and do a lot of outreach there.

8:47 So I don't know if these incentives will change the model development,

8:50 because DeepSeek famously is built by a hedge fund,

8:53 Highflyer Capital, and we don't know exactly what they

8:56 use the models for or if they care about this.

8:59 They're secretive in terms of communication;

9:00 they're not secretive in terms of the technical

9:02 reports that describe how their models work.

9:04 They're still open on that front.

9:05 And we should also say, on the Claude Opus 4.5 hype,

9:10 there's the layer of something being the darling of the X echo chamber,

9:17 on the Twitter echo chamber,

9:18 and the actual amount of people that are using the model.

9:22 I think it's probably fair to say that ChatGPT and Gemini are focused

9:26 on the broad user base that just want to solve problems in their daily lives,

9:32 and that user base is gigantic.

9:34 So the hype about the coding may not be representative of the actual use.

9:39 I would say also a lot of the usage patterns are,

9:43 like you said, name recognition,

9:44 brand and stuff, but also muscle memory almost, where,

9:48 you know, ChatGPT has been around for a long time.

9:51 People just got used to using it, and it's almost like a flywheel:

9:54 they recommend it to other users and that stuff.

9:57 One interesting point is also the customization of LLMs.

10:00 For example, ChatGPT has a memory feature, right?

10:03 And so you may have a subscription and you use it for personal stuff,

10:07 but I don't know if you want to use that same thing at work.

10:10 Because it's a boundary between private and work.

10:12 If you're working at a company,

10:13 they might not allow that or you may not want that.

10:16 And I think that's also an interesting

10:18 point where you might have multiple subscriptions.

10:20 One is just clean code.

10:22 It has nothing of your personal images or hobby projects in there.

10:26 It's just like the work thing.

10:28 And then the other one is your personal thing.

10:30 So I think that's also something where there are two different use cases,

10:32 and it doesn't mean you only have to have one.

10:36 I think the future is also multiple ones.

10:39 What model do you think won 2025,

10:40 and what model do you think is going to win '26?

10:43 I think in the context of consumer chatbots,

10:45 it's a question of: are you willing to bet on Gemini over ChatGPT?

10:50 Which I would say, in my gut,

10:52 feels like a bit of a risky bet because OpenAI has been the incumbent,

10:55 and there are so many benefits to that in tech.

10:58 I think the momentum, if you look at 2025,

11:02 was on Gemini's side, but they were starting from such a low point.

11:07 And RIP Bard and these earlier attempts at getting started.

11:13 Huge credit to them for powering through

11:15 the organizational chaos to make that happen.

11:17 But also it's hard to bet against OpenAI

11:19 because they always come off as so chaotic,

11:23 but they're very good at landing things.

11:25 And I think, personally, I have very mixed reviews of GPT-5,

11:29 but it must have saved them so much money with the high-line feature being

11:33 a router where most users are no longer charging their GPU costs as much.

11:38 So I think it's very hard to dissociate the things that I like out

11:43 of models versus the things that are

11:45 going to actually be a general public differentiator.

11:50 What do you think about 2026?

11:51 Who's going to win?

11:53 I'll say something, even though it's risky.

11:54 I think Gemini will continue to make progress on ChatGPT.

11:56 I think Google's scale,

11:58 when both of these are operating at such extreme scales—and Google

12:02 has the ability to separate research and product a bit better,

12:06 whereas you hear so much about OpenAI

12:08 being chaotic operationally and chasing the high-impact thing,

12:11 which is a very startup culture.

12:13 And then on the software and enterprise side,

12:15 I think Anthropic will have continued success,

12:16 as they've again and again been set up for that.

12:19 And obviously Google Cloud has a lot of offerings,

12:23 but I think this kind of Gemini name brand is important for them to build.

12:27 Google Cloud will continue to do well,

12:30 but that's a more complex thing to explain in the ecosystem,

12:35 because that's competing with the likes of Azure

12:37 and AWS rather than on the model provider side.

12:41 So in infrastructure, you think TPU is giving an advantage?

12:46 Largely because the margin on NVIDIA chips is insane,

12:49 and Google can develop everything from top to bottom

12:51 to fit their stack and not have to pay this margin.

12:54 And they've had a head start in building data centers.

12:57 So all of these things that have both high

12:59 lead times and very hard margins on high costs,

13:02 Google has a just kind of historical advantage there.

13:05 And if there's going to be a new paradigm, it's most likely to come from OpenAI

13:09 where their research division again and again

13:12 has shown this ability to land a new research idea or a product.

13:17 Like Deep Research, Sora,

13:18 o1 thinking models—all these definitional things have come from OpenAI,

13:23 and that's got to be one of their top traits as an organization.

13:27 So it's kind of hard to bet against that, but I think a lot of this year

13:31 will be about scale and optimizing what

13:33 could be described as low-hanging fruit in models.

13:37 And clearly there's a trade-off between intelligence and speed.

13:41 This is what ChatGPT-5 was trying to solve behind the scenes.

13:46 It's like, do people actually want intelligence,

13:49 the broad public, or do they want speed?

13:52 I think it's a nice variety, or the option to have a toggle there.

13:56 I mean, for my personal usage, most of the time when I look something up,

14:00 I use ChatGPT to ask a quick question, get the information I wanted fast.

14:04 For most daily tasks, I use the quick model.

14:07 Nowadays, I think the auto mode is pretty good

14:09 where you don't have to specifically say thinking or non-thinking.

14:12 Then again, I also sometimes want the pro mode.

14:15 Very often what I do is, when I have something written,

14:18 I put it into ChatGPT and say, "Hey, do a very thorough check.

14:23 Are all my references correct?

14:24 Are all my thoughts correct?

14:26 Did I make any formatting mistakes and are

14:28 the figure numbers wrong?" Or something like that.

14:31 And I don't need that right away.

14:33 I finish my stuff, maybe have dinner, let it run, come back and go through this.

14:38 I think this is where it's important to have this option.

14:42 I would go crazy if for each query I

14:43 would have to wait 30 minutes or 10 minutes even.

14:46 That's me.

14:48 I'm sitting over here losing my mind

14:50 that you use the router and the non-thinking model.

14:52 I'm like, "How do you live with that?" That's like my reaction.

14:57 I've been heavily on ChatGPT for a while.

15:01 I never touched ChatGPT-5 non-thinking.

15:03 I find its tone and then its propensity

15:05 for errors—it has a higher likelihood of errors.

15:08 Some of this is from back when OpenAI released o3,

15:11 which was the first model to do this deep

15:14 search and find many sources and integrate them for you.

15:17 I became habituated with that.

15:18 So I will only use GPT-5.2 Thinking or Pro

15:21 when I'm finding any sort of information query for work,

15:25 whether that's a paper or some code reference that I found.

15:28 And I will regularly have like five Pro queries going simultaneously,

15:33 each looking for one specific paper or feedback on an equation or something.

15:38 I have a fun example where I needed the answer as fast

15:41 as possible for this podcast before I was going on the trip.

15:46 like a local GPU running at home and I wanted to run a long RL experiment.

15:50 And usually I also unplug things because you never know if you're not at home,

15:54 you don't want things plugged in.

15:56 And I accidentally unplugged the GPU.

15:57 My wife was already in the car and it's like,

16:00 "Oh dang." Then basically I wanted as fast as possible

16:04 a Bash script that runs my different experiments and the evaluation.

16:09 And it's something I know,

16:10 I learned how to use the Bash interface or Bash terminal,

16:14 but in that moment I just needed like 10 seconds, give me the command.

16:18 This is a hilarious situation but yeah, so what did you use?

16:21 So I did the non-thinking fastest model.

16:23 It gave me the Bash command to chain

16:26 different scripts to each other and then the thing

16:29 is like you have the tee thing where you want to route this to a log file.

16:34 Top of my head I was just like in a hurry, I could have thought about it myself.

16:37 By the way I don't know if there's a representative case,

16:39 wife waiting in the car-...

16:40 you have to run, you know, unplug the GPU.

16:42 You have to generate a Bash script.

16:43 This sounds like a movie, like- Mission Impossible.

16:46 I use Gemini for that.

16:47 So I use thinking for all the information stuff and then

16:50 Gemini for fast things or stuff that I could sometimes Google,

16:52 which is like it's good at explaining things and I trust

16:55 that it has this kind of background of knowledge and it's simple.

16:59 And the Gemini app has gotten a lot

17:00 better and- It's good for those sorts of things.

17:02 And then for code and any sort of philosophical discussion,

17:05 I use Claude Opus 4.5.

17:07 Also always with extended thinking.

17:09 Extended thinking and inference time scaling is just

17:11 a way to make the models marginally smarter.

17:14 And I will always err on that side when the progress is

17:18 very high because you don't know when that'll unlock a new use case.

17:21 And then sometimes use Grok for real-time

17:24 information or finding something on AI Twitter

17:26 that I knew I saw and I need to dig up and I just fixated on.

17:31 Although when Grok 4 came out, the Grok 4 SuperGrok Heavy,

17:35 which was like their pro variant was actually

17:37 very good and I was pretty impressed with it,

17:38 and then it just kind of like muscle memory

17:41 lost track of it with having the ChatGPT app open.

17:44 So I use many different things.

17:46 Yeah.

17:46 I actually do use Grok 4 Heavy for debugging.

17:50 For like hardcore debugging that the other ones can't solve,

17:54 I find that it's the best at.

17:56 And...

17:56 it's interesting 'cause you say ChatGPT is the best interface.

18:00 For me, for that same reason,

18:02 but this could be just momentum- Gemini is the better interface for me.

18:07 I think because I fell in love with their best needle in the haystack.

18:11 If I ever put something that has a lot of context but I'm looking

18:15 for very specific kinds of information to make sure it tracks all of it,

18:19 I find at least that Gemini for me has been the best.

18:24 So it's funny with some of these models,

18:26 if they win your heart over- for one particular feature on one particular day,

18:31 for that particular query, that prompt, you're like,

18:35 "This model's better." And so you'll just stick with it

18:38 for a bit until it does something really dumb.

18:41 There's like a threshold effect.

18:43 Some smart thing and then you fall in love with it

18:45 and then it does some dumb thing and you're like, "You know what?

18:47 I'm gonna switch and try Claude or ChatGPT." And all that kind of stuff.

18:51 This is exactly it: you use it until it breaks,

18:53 until you have a problem, and then you change the LLM.

18:57 And I think it's the same as how we use anything,

19:01 like our favorite text editor, operating systems, or the browser.

19:04 I mean, there are many options: Safari, Firefox, Chrome.

19:07 They're relatively similar, but then there are edge cases,

19:11 extensions you want, and then you switch.

19:14 But I don't think anyone types the same

19:18 thing into different browsers and compares them.

19:21 You only do that when something breaks.

19:23 So that's a good point.

19:25 You use it until it breaks, then you explore other options.

19:28 On the long context thing, I was also a Gemini user,

19:31 but the GPT-5.2 release blog had crazy long context scores.

19:34 People were like, "Did they just figure out some algorithmic change?"

19:38 It went from 30% to 70% in this minor model update.

19:42 It's very hard to keep track of all of these things,

19:46 but now I look more favorably at GPT-5.2's long context.

19:49 So it's just like, "How do I actually

19:53 get to testing this?" It's a never-ending battle.

19:57 Well, it's interesting that none of us talked

19:59 about the Chinese models from a usage perspective.

20:02 What does that say?

20:04 Does it mean the Chinese models are not as good,

20:08 or are we just very biased and US-focused?

20:11 I think currently there's a discrepancy between the model and the platform.

20:15 The open models are more known for the open weights, not the platform yet.

20:19 known for the open weights, not their platform yet.

20:21 Many companies will sell you open-model inference at a very low cost.

20:25 With OpenRouter, it's easy to look at multi-model things.

20:29 You can run DeepSeek on Perplexity.

20:31 Sitting here, we're like, "We use OpenAI GPT-5 Pro consistently." We're all

20:36 willing to pay for the marginal intelligence gain.

20:39 These models from the US are better in terms of the outputs.

20:45 I think the question is,

20:47 will they stay better for this year and for years to come?

20:51 As long as they're better, I'm gonna pay for them.

20:55 There's also analysis showing that the way

20:59 the Chinese models are served—you could argue

21:01 this is due to export controls— is that they use fewer GPUs per replica,

21:05 which makes them slower and have different errors.

21:07 If speed and intelligence are in your favor as a user,

21:10 in the US, a lot of users will go for this.

21:13 And I think that will spur these Chinese

21:15 companies to want to compete in other ways,

21:18 whether it's free or substantially lower costs,

21:21 or it'll breed creativity in terms of offerings,

21:24 which is good for the ecosystem.

21:26 But the simple thing is: the US models are currently better, and we use them.

21:30 I tried these other open models, and I'm like, "Fun,

21:32 but I don't go back." models, and I'm like, "Fun, but not gonna...

21:36 I don't go back to it."- We didn't really mention programming.

21:40 That's another use case that a lot of people deeply care about.

21:45 I use basically half-and-half Cursor and Claude Code, because they're...

21:49 I fundamentally different experiences and both are useful.

21:54 What do you guys...

21:55 You program quite a bit, so what do you use?

21:57 What's the current vibe?

21:59 So, I use the Codeium plugin for VS Code.

22:02 You know, it's very convenient.

22:03 It's just like a plugin,

22:04 and then it's a chat interface that has access to your repository.

22:06 I know that Claude Code is, I think, a bit different.

22:10 It is a bit more agentic.

22:11 It touches more things.

22:12 It does the whole project for you.

22:13 I'm not quite there yet where I'm comfortable

22:16 with that because maybe I'm a control freak,

22:18 but I still would like to see a bit what's going on.

22:21 And Codeium is kind of, right now, for me,

22:24 the sweet spot where it is helping me, but it is not taking completely over.

22:29 I should mention, one of the reasons I do use

22:31 Claude Code is to build the skill of programming with English.

22:34 I mean, the experience is fundamentally different.

22:37 You're...

22:38 As opposed to micromanaging the details

22:40 of the process of the generation of the code, and looking at the diff,

22:45 which you can in Cursor if that's the IDE you use, and in changing, altering.

22:52 Looking and reading the code and understanding the code deeply as you progress,

22:56 versus just thinking in this design space

23:00 and just guiding it at this macro level,

23:05 which I think is another way of thinking about the programming process.

23:10 Also, we should say that Claude Code just seems

23:14 to be somehow a better utilization of Claude Opus 4.5.

23:19 It's a good side-by-side for people to do.

23:20 You can have Claude Code open,

23:21 you can have Cursor open, you can have VS Code open,

23:24 and you can select the same models on all of them— ...and ask questions,

23:27 and it's very interesting.

23:28 Claude Code is way better in that domain.

23:32 It's remarkable.

23:33 All right, we should say that both of you are legit on multiple fronts:

23:36 researchers, programmers, educators, Tweeters.

23:43 And on the book front, too.

23:45 So Nathan, at some point soon, hopefully has an RLHF book coming out.

23:50 It's available for preorder, and there's a full digital preprint.

23:54 I'm just making it pretty and better organized for the physical thing,

23:56 which is a lot of why I do it,

23:58 because it's fun to create things that you think are excellent

24:01 in the physical form when so much of our life is digital.

24:05 I should say, going to Perplexity here,

24:07 Sebastian Raschka is a machine learning researcher

24:09 and author known for several influential books.

24:11 A couple of them that I wanted to mention—which is

24:14 a book I highly recommend—Build a Large Language Model from Scratch,

24:18 and the new one, Build a Reasoning Model from Scratch.

24:21 So, I'm really excited about that.

24:24 Building stuff from scratch is one of the most powerful ways of learning.

24:28 Honestly, building an LLM from scratch is a lot of fun.

24:30 It's also a lot to learn.

24:31 And like you said, it's probably the best

24:33 way to learn how something really works,

24:35 'cause you can look at figures, but figures can have mistakes.

24:38 You can look at concepts and explanations, but you might misunderstand them.

24:43 But if there is code, and the code works, you know it's correct.

24:48 I mean, there's no misunderstanding.

24:50 It's precise.

24:50 Otherwise, it wouldn't work.

24:52 And I think that's the beauty behind coding.

24:54 It doesn't lie.

24:56 It's math, basically.

24:57 So, even though with math,

24:59 I think you can have mistakes in a book you would never notice.

25:02 Because you are not running the math when you are reading the book,

25:05 you can't verify this.

25:06 And with code, what's nice is you can verify it.

25:09 Yeah, I agree with you about the Build an LLM from Scratch book.

25:12 It's nice to tune out everything else,

25:14 the internet and so on, and just focus on the book.

25:16 But, you know, I read several history books.

25:21 It's just less lonely somehow.

25:24 It's really more fun.

25:25 Like for example, on the programming front,

25:28 I think it's genuinely more fun to program with an LLM.

25:31 And I think it's genuinely more fun to read with an LLM.

25:36 But you're right.

25:37 That distraction should be minimized.

25:40 So you use the LLM to basically enrich the experience, maybe add more context.

25:48 I just find the rate of aha moments for me is really high with LLMs.

25:55 100%.

25:55 I also want to correct myself: I'm not suggesting not to use LLMs.

25:58 I suggest doing it in multiple passes.

26:01 Like, one pass just offline, focus mode, and then after that...

26:05 I mean, I also take notes, but I,

26:07 I try to resist the urge to immediately look things up.

26:11 I do a second pass.

26:13 It's just more structured this way.

26:15 Sometimes things are answered in the chapter,

26:18 but sometimes also it just helps to let it sink in and think about it.

26:22 Other people have different preferences.

26:24 I highly recommend using LLMs when reading books.

26:26 For me, it's not the first thing to do; it's the second pass.

26:30 My recommendation is the opposite.

26:31 I like to use the LLM at the beginning to lay out

26:36 the full context of what is this world that I'm now stepping into?

26:40 But I try to avoid clicking out of the LLM into the world of Twitter and blogs,

26:47 because then you're down this rabbit hole.

26:50 You're reading somebody's opinion.

26:51 There's a flame war about a particular topic and all of a sudden

26:55 you're in the realm of the internet and Reddit and so on.

27:00 But if you're purely letting the LLM give you the context of why this matters,

27:05 what are the big picture ideas...

27:07 sometimes books are good at doing that, but not always.

27:12 This is why I like the ChatGPT app,

27:14 because it gives the AI a home on your computer where you can focus on it,

27:17 rather than just being another tab in my mess of internet options.

27:21 And I think Claude Code does a good job of making that a joy,

27:26 where it seems very engaging as a product design to be

27:30 an interface that your AI will then go out into the world.

27:34 It's something that is intangible between it and Codex;

27:37 it just feels warm and engaging, where Codex can often be as good from OpenAI,

27:42 but it just, feel a little bit rough around the edges.

27:46 Whereas Claude Code makes it fun to build things from scratch,

27:50 where you just trust that it'll make something.

27:53 Obviously this is good for websites and kind of refreshing tooling

27:57 and stuff like this, which I use it for, or data analysis.

28:01 For my On my blog, we scrape Hugging Face

28:04 so we keep download numbers for every dataset and model.

28:06 over time, so we have them.

28:07 And Claude was just like, "Yeah,

28:09 I've made use of that data, no problem." And I was like,

28:12 "That would've taken me days." And then

28:14 I have enough situational awareness to be like,

28:16 "Okay, these trends obviously make sense." You can check things.

28:18 But that's just a wonderful interface where you

28:20 can have an intermediary and not have to do

28:23 the kind of awful low-level work that you

28:26 would have to do to maintain different web projects.

28:29 All right.

28:30 So we just talked about a bunch of the closed-weight models.

28:33 Let's talk about the open ones.

28:36 Tell me about the landscape of open LLM models.

28:39 Which are interesting?

28:40 Which stand out to you and why?

28:42 We already mentioned DeepSeek R1.

28:45 Do you wanna see how many we can name off the top of our head?

28:47 Yeah, without looking at notes.

28:49 DeepSeek, Kimi, MiniMax, Z.ai, Moonshot.

28:53 We're just going Chinese.

28:57 Let's throw in Mistral AI, Gemma...

29:01 ...gpt-oss, the open weight model by OpenAI.

29:04 Actually, NVIDIA had a really cool one, Nemotron 3.

29:09 There, there's a lot of stuff especially at the end of the year.

29:11 Qwen might be the one—- Oh, yeah.

29:13 Qwen was the obvious name I was gonna say.

29:15 You can get at least 10 Chinese and at least 10 Western.

29:18 I think that OpenAI released their first open model— ...since GPT-2.

29:23 When I was writing about OpenAI's open model release, they were like,

29:27 "Don't forget about GPT-2," which I thought was

29:29 really funny 'cause it's just such a different time.

29:32 But gpt-oss-120b is actually a very strong model and does

29:35 some things that other models don't do very well.

29:39 Selfishly, I'll promote a bunch of Western companies

29:43 in the US and Europe that have these fully open models.

29:46 I work at the Allen Institute for AI,

29:48 where we've been building OLMo, which releases data and code.

29:51 And now we have actual competition for people that are

29:55 trying to release everything so that others can train these models.

29:58 There's the Institute for Foundation Models/LM360,

30:00 which has had their K2 models of various types.

30:04 Apertus is a Swiss research consortium.

30:07 Hugging Face has SmolLM, which is very popular.

30:12 And NVIDIA's Nemotron 3 has started releasing data as well.

30:15 And then Stanford's Martini Community Project,

30:17 which is kind of making it so there's a pipeline for people to open a GitHub

30:21 issue and implement a new idea and then

30:23 have it run in a stable language modeling stack.

30:26 This space, that list was way smaller in 2024— ...so I think it was just AI2.

30:32 So it's a great thing for more people

30:34 to get involved and to understand language models,

30:36 which doesn't really have a Chinese analog.

30:39 While I'm talking, I'll say that the Chinese

30:44 open language models tend to be much bigger,

30:47 and that gives them higher peak performance as MoEs,

30:49 where a lot of these things that we like a lot,

30:52 whether it was Gemma and Nemotron, have tended to be smaller models from the US,

30:57 which is starting to change from the US and Europe.

30:59 Mistral Large 3 came out, which was a giant MoE model,

31:02 very similar to DeepSeek architecture in December.

31:05 And then a startup, RCAI,

31:08 and both Nemotron and NVIDIA have teased MoE models way bigger than 100

31:15 billion parameters- like this 400 billion parameter

31:17 range coming in this Q1 2026 timeline.

31:20 So I think this kind of balance is set

31:23 to change this year in terms of what people

31:25 are using the Chinese versus US open models for, which

31:28 I'm personally going to be very excited to watch.

31:32 First of all, huge props for being able to name so many of these.

31:36 Did you actually name LLaMA?

31:39 No.

31:39 I feel like...

31:41 RIP.

31:41 This was not on purpose.

31:43 RIP LLaMA.

31:45 All right.

31:45 Can you mention some interesting models that stand out?

31:48 You mentioned Qwen 3 is obviously a standout.

31:51 So I would say the year's almost bookended by both DeepSeek V3 and R1.

31:56 And then on the other hand, in December, DeepSeek-V3.2.

31:59 Because what I like about those is they always

32:01 have an interesting architecture tweak that others don't have.

32:05 But otherwise, if you want to go with the familiar but really good performance,

32:09 Qwen 3 and, like Nathan said, also gpt-oss-120b.

32:13 And I think what's interesting about it is it's kind of like the first

32:18 public or open weight model that was really trained with tool use in mind,

32:22 which I do think is kind of a paradigm

32:25 shift where the ecosystem was not quite ready for it.

32:27 By tool use, I mean that the LLM is able

32:30 to do a web search or to call a Python interpreter.

32:33 And I do think it's a standout because it's a huge unlock.

32:37 Because one of the most common complaints about LLMs are,

32:41 for example, hallucinations, right?

32:43 And so, in my opinion, one of the best ways to solve hallucinations is

32:46 to not try to always remember information or make things up.

32:51 For math, why not use a calculator app or Python?

32:54 If I ask the LLM, "Who won the soccer

32:57 World Cup in 1998?" instead of just trying to memorize, it could go do a search.

33:03 I think mostly it's still a Google search.

33:06 So ChatGPT and gpt-oss-120b, they would do a tool call to Google,

33:09 maybe find the FIFA website.

33:11 Find, okay, it was France.

33:13 It would get you that information reliably

33:15 instead of just trying to memorize it.

33:17 So I think it's a huge unlock which right

33:20 now is not fully utilized yet by the open-source, open-weight ecosystem.

33:24 A lot of people don't use tool call modes because I think,

33:28 first, it's a trust thing.

33:29 You don't want to run this on your computer where it has access to tools,

33:32 could wipe your hard drive or whatever.

33:34 So you want to maybe containerize that.

33:36 But I do think that is like a really

33:40 important step for the upcoming years to have this ability.

33:44 So a few quick things.

33:45 First of all, thank you for defining what you mean by tool use.

33:49 I think that's a great thing to do

33:50 in general for the concepts we're talking about.

33:53 Even things as sort of well-established as MoEs.

33:57 You have to say that means mixture of experts,

34:00 and you kind of have to build up an intuition for people what that means,

34:04 how it's actually utilized, what are the different flavors.

34:06 So what does it mean that there's just such an explosion of open models?

34:11 What's your intuition?

34:13 If you're releasing an open model,

34:14 you want people to use it, is the first and foremost thing.

34:17 And then after that comes things like transparency and trust.

34:20 I think when you look at China, the biggest reason is that they want

34:24 people around the world to use these models,

34:26 and I think a lot of people will not.

34:28 If you look outside of the US, a lot of people will not pay for software,

34:31 but they might have computing resources where you

34:32 can put a model on it and run it.

34:34 I think there can also be data that you don't want to send to the cloud.

34:37 So the number one thing is getting people to use models, use AI,

34:41 or use your AI that might not be able

34:43 to do it without having access to the model.

34:46 I guess we should state explicitly,

34:47 so we've been talking about these Chinese models and open weight models.

34:51 Oftentimes, the way they're run is locally.

34:54 So it's not like you're sending your data

34:58 to China or to whoever developed Silicon Valley, or whoever developed the model.

35:04 A lot of American startups make money

35:06 by hosting- ...these models from China and selling them.

35:09 It's called selling tokens,

35:11 which means somebody will call the model to do some piece of work.

35:15 I think the other reason is for US companies like OpenAI.

35:18 They are so GPU deprived.

35:20 They're at the limits of the GPUs.

35:22 Whenever they make a release, they're always talking about like,

35:25 "Our GPUs are hurting." And I think

35:27 during one of these gpt-oss-120b release sessions,

35:30 Sam Altman said, "Oh, we're releasing this because we can use your GPUs.

35:33 We don't have to use our GPUs, and OpenAI can still get distribution out

35:38 of this," which is another very real thing,

35:41 because it doesn't cost them anything.

35:44 And for the user, I think also,

35:45 there are users who just use the model locally how they would use ChatGPT.

35:49 But also for companies I think it's a huge

35:51 unlock to have these models because you can customize them,

35:53 you can train them, you can add post-training, add more data.

35:58 Like, specialize them into, let's say, law, medical models, whatever you have.

36:02 And the appeal, you mentioned Llama,

36:04 the appeal of the open-weight models from China

36:07 is that the open-weight models' licenses are even friendlier.

36:11 I think they are just unrestricted open source licenses

36:13 where if we use something like Llama or Gemma, there are some strings attached.

36:17 I think it's like an upper limit in terms of how many users you have.

36:20 And then if you exceed, I don't know, so and so many million users,

36:23 you have to report your financial situation to, let's say,

36:26 Meta or something like that.

36:28 And I think while it is a free model, there are strings attached,

36:33 and people do like things where strings are not attached.

36:36 So I think that's also one of the reasons, besides performance,

36:39 why the open-weight models from China are so popular,

36:42 because you can just use them.

36:43 There's no catch in that sense.

36:46 The ecosystem has gotten better on that front,

36:48 but mostly downstream of these new providers providing such open licenses.

36:51 That was funny when you pulled up Perplexity and said,

36:53 "Kimi K2 Thinking hosted in the US." Which is just like an exact...

36:56 I've never seen this, but it's an exact example

36:58 of what we're talking about where people are sensitive to this.

37:01 But Kimi K2 Thinking and Kimi K2 is a model that is very popular.

37:05 People say that has very good creative

37:07 writing and also in doing some software things.

37:09 So it's just these little quirks that people

37:11 pick up on with different models that they like.

37:14 What are some interesting ideas that some of these models have

37:18 explored that you can speak to, that are particularly interesting to you?

37:22 Maybe we can go chronologically.

37:23 I mean, there was, of course, DeepSeek.

37:25 DeepSeek R1 that came out in January of 2025, if we just focus on 2025.

37:29 However, this was based on DeepSeek-V3,

37:31 which came out the year before in December 2024.

37:34 There are multiple things on the architecture side.

37:37 What is fascinating is...

37:38 I mean, that's what I do with my from-scratch coding projects.

37:41 You can still start with GPT-2,

37:43 and you can add things to that model to make it into this other model.

37:47 So it's all still kind of like the same lineage.

37:50 It is a very close relationship between those.

37:53 But top of my head, DeepSeek—what was unique there is the Mixture of Experts.

37:57 Not that they were inventing Mixture of Experts—we

38:00 can maybe talk a bit more about what Mixture of Experts means—but just to list

38:04 these things first before we dive into detail.

38:07 Mixture of Experts, but then they also had Multi-head Latent Attention,

38:11 which is a tweak to the attention mechanism, where this was, I would say,

38:17 the main distinguishing factor between these open-weight models.

38:22 Different tweaks to make inference or KV cache size...

38:25 We can also define KV cache in a few moments,

38:29 but to kind of make it more economical to have long context,

38:32 to shrink the KV cache size.

38:34 So what are tweaks that we can do?

38:36 And most of them focused on the attention mechanism.

38:38 There is Multi-head Latent Attention in DeepSeek.

38:41 There is Group Query Attention, which is still very popular.

38:44 It's not invented by any of those models.

38:46 It goes back a few years.

38:47 But that would be the other option.

38:50 Sliding window attention—I think OLMo 3 uses it, if I remember correctly.

38:54 So there are these different tweaks that make the models different.

38:57 Otherwise, I put them all together

39:00 in an article once where I just compared them.

39:03 They are very, surprisingly similar.

39:05 It's just different numbers in terms of how many

39:08 repetitions of the transformer block you have in the center.

39:11 And, like, just little knobs that people tune.

39:14 But what's so nice about it is it works no matter what.

39:17 You can tweak things.

39:19 You can move the normalization layers around to get some performance gains.

39:22 And OLMo is always very good in ablation studies,

39:26 showing what it actually does to the model if you move something around.

39:30 Ablation studies: does it make it better or worse?

39:32 But there are so many, let's say,

39:33 ways you can implement a transformer and make it still work.

39:36 The big ideas that are still prevalent is Mixture of Experts,

39:40 multi-head latent attention, sliding window attention, group query attention.

39:44 And then at the end of the year, we saw a focus on making the attention

39:49 mechanism scale linearly with inference token prediction.

39:52 So there was Qwen2-VL, for example, which added a gated delta net.

39:57 It's kind of inspired by State space models,

40:00 where you have a fixed state that you keep updating.

40:02 But it makes essentially this attention cheaper,

40:06 or it replaces attention with a cheaper operation.

40:08 And it may be useful to step

40:11 back and talk about transformer architecture in general.

40:14 Yeah, so maybe we should start with the GPT-2 architecture.

40:17 The transformer that was derived from the "Attention Is All You Need" paper.

40:21 The "Attention Is All You Need" paper

40:23 had a transformer architecture that had two parts, an encoder and a decoder.

40:28 And GPT went just focusing in on the decoder part.

40:32 It is essentially still a neural network

40:35 and it has this attention mechanism inside.

40:37 And you predict one token at a time.

40:40 You pass it through an embedding layer.

40:43 There's the transformer block.

40:44 The transformer block has attention modules and a fully connected layer.

40:47 And there are some normalization layers in between.

40:50 But it's essentially neural network layers with this attention mechanism.

40:53 So coming from GPT-2 when we move on to gpt-oss-120b,

40:57 there is, for example, the Mixture of Experts layer.

41:00 It's not invented by gpt-oss-120b.

41:02 It's a few years old.

41:04 But it is essentially a tweak to make the model

41:09 larger without consuming more compute in each forward pass.

41:13 So there is this fully connected layer,

41:15 and if listeners are familiar with multi-layer perceptrons,

41:19 you can think of a mini multi-layer perceptron,

41:22 a fully connected neural network layer inside the transformer.

41:25 And it's very expensive, because it's fully connected.

41:27 If you have a thousand inputs and a thousand outputs,

41:30 that's like one million connections.

41:31 And it's a very expensive part in this transformer.

41:34 And the idea is to kind of expand that into multiple feedforward networks.

41:39 So instead of having one, let's say you have 256,

41:43 but it would make it way more expensive,

41:45 because now you have 256, but you don't use all of them at the same time.

41:49 So you now have a router that says, "Okay, based on this input token,

41:52 it would be useful to use this fully connected network." And in that context,

41:57 it's called an expert.

41:58 So a Mixture of Experts means you have multiple experts.

42:01 And depending on what your input is, let's say it's more math-heavy,

42:05 it would use different experts, compared to, let's say,

42:09 translating input text from English to Spanish.

42:11 It would maybe consult different experts.

42:13 It's not quite clear, I mean, not as clear-cut to say, "Okay,

42:16 this is only an expert for math and for Spanish." It's a bit more fuzzy.

42:20 But the idea is essentially that you pack more knowledge into the network,

42:25 but not all the knowledge is used all the time.

42:27 That would be very wasteful.

42:29 So, during the token generation, you are more selective.

42:32 There's a router that selects which tokens should go to which expert.

42:36 It adds more complexity.

42:38 It's harder to train.

42:39 There's a lot that can go wrong, like collapse and everything.

42:42 So I think that's why OLMo 3 still uses dense...

42:45 I mean, you have OLMo models with Mixture of Experts,

42:48 but dense models, where dense means...

42:50 So also, it's jargon.

42:52 There's a distinction between dense and sparse.

42:55 So Mixture of Experts is considered sparse, because we have a lot of experts,

42:59 but only a few of them are active.

43:01 So that's called sparse.

43:01 And then dense would be the opposite,

43:03 where you only have one fully connected module, and it's always utilized.

43:08 So maybe this is a good place to also talk about KV cache.

43:11 But actually, before that, even zooming out, like fundamentally,

43:14 how many new ideas have been implemented from GPT-2 to today?

43:22 Like, how different really are these architectures?

43:25 Take the Mixture of Experts.

43:27 The attention mechanism in gpt-oss-120b,

43:29 that would be the Group Query Attention mechanism.

43:31 So it's a slight tweak from Multi-Head Attention to Group Query Attention.

43:35 So that we have too...

43:36 I think they replaced LayerNorm by RMSNorm,

43:39 but it's just like a different normalization there and not a big change.

43:43 It's just like a tweak.

43:45 The nonlinear activation function— people familiar with deep neural networks,

43:49 I mean, it's the same as changing sigmoid with ReLU.

43:52 It's not changing the network fundamentally.

43:55 It's just a little tweak.

43:56 And that's about it, I would say.

43:59 It's not really fundamentally that different.

44:01 It's still the same architecture.

44:03 So you can go from one into the other by just adding these changes basically.

44:10 It fundamentally is still the same architecture.

44:12 Yep.

44:12 For example, you mentioned my book earlier.

44:14 That's a GPT-2 model in the book because it's simple and it's very small,

44:18 so 124 million parameters approximately.

44:20 But in the bonus materials, I do have OLMo from scratch,

44:25 Gemini 3 from scratch, and other types of from-scratch models.

44:28 And I always start it with my GPT-2 model and just tweak the—well,

44:31 add different components and you get from one to the other.

44:34 It's kind of like a lineage in a sense.

44:38 Can you build up an intuition for people?

44:40 Because when you zoom out, you look at it,

44:43 there's so much rapid advancement in the AI world.

44:46 And at the same time, fundamentally the architectures have not changed.

44:51 So where is all the turbulence, the turmoil of the advancement happening?

44:58 Where are the gains to be had?

45:01 So there are different stages where you

45:03 develop the network or train the network.

45:05 You have the pre-training.

45:06 Now back in the day, it was just pre-training with GPT-2.

45:09 Now you have pre-training, mid-training, and post-training.

45:12 So I think right now we are in the post-training focus stage.

45:17 Pre-training still gives you advantages if you scale it up with better,

45:23 higher quality data.

45:24 But then we have capability unlocks that were not there with GPT-2,

45:28 for For example, ChatGPT is basically a GPT-3 model.

45:32 And GPT-3 is the same as GPT-2 in terms of architecture.

45:36 What was new was adding supervised

45:39 fine-tuning and reinforcement learning with human feedback.

45:41 So it's more on the algorithmic side than the architecture.

45:45 I would say that the systems also change a lot.

45:47 If you listen to NVIDIA's announcements,

45:48 they talk about things like, "You now do FP8,

45:51 you can now do FP4." What's happening is these labs are figuring

45:55 out how to utilize more compute to put it into one model,

45:58 which lets them train faster and put more data in.

46:01 And then you can find better configurations faster by doing this.

46:05 So you can look at, essentially, tokens per second per GPU as a metric

46:09 that you look at when you're doing large-scale training.

46:12 You can go from 10k to 13k by turning on FP8 training,

46:16 which means you're using less memory per parameter in the model.

46:20 By saving less information, you do less communication and train faster.

46:24 So all of these system things underpin

46:27 way faster experimentation on data and algorithms.

46:35 It's a loop that keeps going where it's hard to describe

46:38 when you look at architectures and they're exactly the same,

46:40 but the code base used to train

46:42 these models is vastly different- -and you could probably...

46:45 the GPUs are different but you probably train gpt-oss-20b way

46:49 faster in wall-clock time than GPT-2 was trained at the time.

46:54 Yeah.

46:54 Like you said, they had, for example,

46:56 in Mixture of Experts this FP4 optimization where you get more throughput.

47:00 But I do think, for speed this is true,

47:04 but it doesn't give the model new capabilities.

47:07 It's just: how much can we make the computation

47:11 coarser without suffering in terms of model performance degradation?

47:15 But I do think- I mean, there are alternatives popping up to the transformer.

47:20 Text diffusion models, a completely different paradigm.

47:23 And there is also...

47:24 I mean, although text diffusion models might use transformer architectures,

47:27 it's not an autoregressive transformer.

47:30 And also Mamba models.

47:32 It's a state space model.

47:34 But they do have trade-offs, and nothing has yet replaced

47:39 the autoregressive transformer as the state-of-the-art model.

47:42 For state-of-the-art, you would still go with that, but there are now

47:46 alternatives for the cheaper end—alternatives

47:49 that are kind of making compromises.

47:52 It's not just one architecture anymore.

47:54 There are little ones coming up.

47:57 But if we talk about the state-of-the-art,

47:59 it's pretty much still the transformer architecture,

48:02 autoregressive, derived from GPT-2 essentially.

48:06 I guess the big question here is,

48:07 we talked quite a bit about the architecture behind the pre-training.

48:11 Are the scaling laws holding strong across pre-training,

48:16 post-training, inference, context size, data, and synthetic data?

48:21 I'd like to start with the technical definition

48:22 of a scaling law- -which informs all of this.

48:24 The scaling law is the power law relationship between...

48:27 You can think of the x-axis,

48:28 so kind of what you are scaling as a combination of compute and data,

48:32 which are kind of similar,

48:34 and then the y-axis is like the held-out prediction accuracy over next tokens.

48:38 We talked about models being autoregressive.

48:39 It's like if you keep a set of text that the model has not seen,

48:45 how accurate will it get when you train?

48:47 And the idea of scaling laws came when people

48:50 figured out that that was a very predictable relationship.

48:53 And I think that that technical term is continuing,

48:57 and then the question is, what do users get out of it?

49:01 Then there are more types of scaling where,

49:03 OpenAI's o1 was famous for introducing inference time scaling.

49:06 And I think less famously for also

49:08 showing that you can scale reinforcement learning training

49:11 and get kind of this log x-axis

49:14 and then a linear increase in performance on y-axis.

49:16 So there's kind of these three axes now where

49:19 the traditional scaling laws are talked about for pre-training,

49:21 which is how big your model is and how big your dataset is,

49:25 and then scaling reinforcement learning, which is like how long can you do

49:28 this trial and error learning that we'll talk about.

49:30 We'll define more of this, and then this inference time compute,

49:33 which is just letting the model generate more tokens on a specific problem.

49:36 So I'm kind of bullish, but they're all really still working,

49:40 but the low-hanging fruit has mostly been taken,

49:43 especially in the last year on reinforcement learning with verifiable rewards,

49:46 which is this RLVR, and then inference time scaling,

49:50 which is just why these models feel so different to use,

49:53 where previously you would get that first token immediately.

49:55 And now they'll go off for seconds, minutes, or even hours,

49:59 generating these hidden thoughts before giving

50:01 you the first word of your answer.

50:03 And that's all about this inference time scaling,

50:05 which is such a wonderful kind of step

50:08 function in terms of how the models change abilities.

50:11 They kind of enabled this tool use stuff and enabled

50:13 this much better software engineering that we were talking about.

50:17 And this, when we say enabled, is almost entirely downstream of the fact

50:21 that this reinforcement learning with verifiable

50:23 rewards training just kind of let the models pick up these skills very easily.

50:27 So let the models learn, so if you look at the reasoning process

50:32 when the models are generating a lot of tokens,

50:34 what it'll often be doing is: it tries a tool, it looks at what it gets back.

50:37 It tries another API, it sees what it gets back and if it solves the problem.

50:41 So the models, when you're training them, very quickly learn to do this.

50:45 And then at the end of the day,

50:47 that gives this kind of general foundation where the model

50:49 can use CLI commands very nicely in your repo

50:52 and handle Git for you and move things around

50:55 and organize things or search to find more information,

50:57 which if we were sitting in these chairs a year ago

51:00 is something that we didn't really think of the models doing.

51:03 So this is just kind of something that has

51:05 happened this year and has totally transformed how

51:07 has totally transformed how we think of using

51:10 AI which evolution and just unlocks so much value.

51:13 But it's like, just so- pr- unlocks so much value.

51:18 But it's- it's like, it's not clear what the next avenue will

51:21 be in terms of unlocking stuff like this.

51:23 I think there's...

51:24 we'll get to continual learning later,

51:25 but there's a lot of buzz around certain areas of AI,

51:28 but no one knows when the next step function will really come.

51:32 So you've actually said quite a lot of things there,

51:35 and said profound things quickly.

51:37 It would be nice to unpack them a little bit.

51:40 You say you're bullish basically on every version of scaling.

51:43 So can we just even start at the beginning?

51:47 Pre-training, are we kind of implying that the low-

51:51 hanging fruit on pre-training scaling has been picked?

51:55 Has pre-training hit a plateau,

51:58 or is even pre-training still something you're bullish on?

52:01 Pre-training has gotten extremely expensive.

52:03 I think to scale up pre-training,

52:05 it's also implying that you're gonna serve a very large model to the users.

52:10 So I think that it's been loosely established the likes of GPT-4

52:14 and similar models were around one trillion parameters at the biggest size.

52:18 There's a lot of rumors that they've actually

52:20 gotten smaller as training has gotten more efficient.

52:23 You want to make the model smaller because

52:25 then your costs of serving go down proportionately.

52:28 These models, the cost of training them is really low relative

52:31 to the cost of serving them to hundreds of millions of users.

52:34 I think DeepSeek had this famous number of about

52:36 five million dollars for pre-training at cloud market rates.

52:40 In OLMo 3, section 2.4 in the paper, we just detailed how long we had the GPU

52:46 clusters sitting around for training which includes engineering issues,

52:50 multiple seeds, and it was like about two million dollars to rent

52:53 the cluster to deal with all the headaches of training a model.

52:56 So these models are pretty— like,

52:59 a lot of people could get one to 10 million dollars to train a model,

53:02 but the recurring costs of serving millions

53:05 of users is really billions of dollars of compute.

53:08 I think that you can look at a thousand

53:11 GPU rental you can pay 100 grand a day for.

53:14 And these companies could have millions of GPUs.

53:16 Like you can look at how much these things cost to sit around.

53:19 So that's kind of a big thing, and then it's like,

53:23 if scaling is actually giving you a better model,

53:25 is it gonna be financially worth it?

53:27 And I think we'll slowly push it out as AI solves more compelling tasks,

53:31 so like the likes of Claude Opus 4.5, making Claude Code just work for things.

53:36 I— I launched this project called the ATOM project,

53:39 which is American Truly Open Models in July,

53:42 and that was like a true vibe coded website,

53:45 and like, I have a job to make plots and stuff.

53:49 And then I came back to refresh it in the last few weeks

53:51 and it's like Claude Opus 4.5 versus whatever model at the time was like,

53:55 just crushed all the issues that it had from building in June and July and like,

53:59 it might be a bigger model.

54:01 There's a lot of things that go into this, but there's still progress coming.

54:04 So what you're speaking to is the nuance of the y-axis

54:07 of the scaling laws—the way it's experienced versus on a benchmark,

54:11 the actual intelligence might be different.

54:13 But still, your intuition about pre-training,

54:16 if you scale the size of compute, will the models get better?

54:21 Not whether it's financially viable but just from the law aspect of it,

54:26 do you think the models will get smarter?

54:28 Yeah.

54:29 And I think that there's...

54:30 And this sometimes comes off as almost like disillusionment from people,

54:34 leadership at AI companies saying this, but they're like,

54:37 "It's held for 13 orders of magnitude of compute,

54:40 why would it ever end?" So I think fundamentally it is pretty unlikely to stop,

54:44 it's just eventually we're not even gonna be able to test

54:47 the bigger scales because of all the problems that come with more compute.

54:50 I think that there's a lot of talk on how

54:54 2026 is a year when very large Blackwell compute clusters,

54:58 like gigawatt-scale facilities at hyperscalers, are coming online.

55:03 These were all contracts for power and data centers

55:06 that were signed and sought out in 2022 and 2023.

55:11 So before or right after ChatGPT.

55:13 It took this two-to-three-year lead time to build

55:16 these bigger clusters to train the models.

55:18 While there's obviously immense interest in building

55:20 even more data centers than that.

55:21 So that is the crux that people are saying: these new clusters are coming.

55:25 The labs are gonna have more compute for training.

55:28 They're going to utilize this, but it's not a given.

55:31 I've seen so much progress that I expect it,

55:34 and I expect a little bit bigger models, and I expect...

55:39 I would say it's more like we'll see a $2,000 subscription this year.

55:42 We've seen $200 subscriptions.

55:43 That could 10X again, and these are the kind of things that could come,

55:47 and they're all downstream of this bigger model

55:50 that offers just a little bit more cutting edge.

55:53 So, you know, it's reported that xAI

55:55 is gonna hit that one-gigawatt scale early '26,

55:59 and a full two gigawatts by year end.

56:03 How do you think they'll utilize that in the context of scaling laws?

56:09 Is a lot of that inference?

56:10 Is a lot of that training?

56:13 It ends up being all of the above.

56:15 So I think that all of your decisions

56:17 when you're training a model come back to pre-training.

56:20 So if you're going to scale RL on a model,

56:22 you still need to decide on your architecture that enables this.

56:25 We were talking about other architectures

56:27 and using different types of attention, or a mixture of experts models.

56:31 The sparse nature of MoE models makes it much more efficient to do generation,

56:37 which becomes a big part of post-training,

56:40 and you need to have your architecture ready

56:42 so that you can actually scale up this compute.

56:45 I still think most of the compute is going in at pre-training.

56:48 Because you can still make a model better,

56:51 you still want to go and revisit this.

56:53 You still want the best base model you can.

56:55 And in a few years that'll saturate and the RL compute will just go longer.

57:00 Are there people who disagree with you and say pre-training is dead?

57:06 It's all about scaling inference, scaling post-training,

57:09 scaling context, continual learning, scaling data, synthetic data?

57:15 People vibe that way and describe it in that way,

57:17 but I think it's not the practice that is happening.

57:19 It's just the general vibe of people saying

57:21 this thing is dead-- The excitement is elsewhere.

57:23 So the low-hanging fruit- ...in RL is elsewhere.

57:26 For example, we released our model in November...

57:28 Every company has deadlines.

57:30 Our deadline was November 20th, and for that, our run was five days,

57:34 which compared to 2024 is a very long time to just

57:37 be doing post- training at a model of 30 billion parameters.

57:40 It's not a big model.

57:41 And then in December, we had another release,

57:43 where we let the RL run for another three and a half weeks,

57:47 and the model got notably better, so we released it.

57:50 And that's a to just allocate to something

57:53 that is going to be your peak- ...for the year.

57:56 So it's like-- The reasoning is-- There's these types

57:59 of decisions when training a model where they just...

58:01 They can't leave it forever.

58:03 You have to keep pulling in the improvements from researchers.

58:07 So you redo pre-training, you'll do this post-training for a month,

58:11 but then you need to give it to your users.

58:14 You need to do safety testing.

58:15 So it's just...

58:16 I think there's a lot in place

58:18 that reinforces this cycle of updating the models.

58:21 Things improve.

58:22 You get a new compute cluster that lets you do something more stably or faster.

58:27 It's like you hear a lot about Blackwell having rollout issues, where at AI2,

58:32 most of the models we're pre-training are on 1,000 to 2,000 GPUs.

58:35 But when pre-training on 10,000 or 100,000 GPUs,

58:38 you hit very different failures.

58:40 GPUs break in weird ways, and on a 100,000 GPU run,

58:44 you're pretty much guaranteed to have one GPU that is down.

58:48 Your training code must handle that redundancy,

58:50 which is a very different problem.

58:51 Whereas what we're doing, like playing with post-training on a cluster,

58:55 or for people learning ML, what they're battling to train these biggest

59:00 models is just- ...mass distributed scale, and it's very different.

59:05 But that's somewhat different than- That's a systems

59:10 problem- ...in order to enable scaling laws, especially at pre-training.

59:15 You need all these GPUs at once.

59:17 When we shift to RL, it actually lends itself to heterogeneous compute

59:21 because you have many copies of the model.

59:24 To do a primer for language model reinforcement learning,

59:28 what you're doing is having two sets of GPUs.

59:31 One you can call the actor, and one you call the learner.

59:34 The learner is where your actual reinforcement learning updates happen.

59:38 These are traditionally policy gradient algorithms.

59:42 Proximal Policy Optimization, PPO, and Group Relative Policy Optimization,

59:46 GRPO, are the two popular classes.

59:50 And on the other side you have actors which are generating completions,

59:54 and these completions are what you're going to grade.

59:57 Reinforcement learning is all about optimizing reward.

1:00:00 In practice, you can have a lot of different actors

1:00:03 in different parts of the world doing different types of problems,

1:00:06 and then you send it back to this highly networked compute

1:00:10 cluster to do this actual learning where you take the gradients.

1:00:14 You need to have a tightly meshed network to do different

1:00:18 types of parallelism and spread out your model for efficient training.

1:00:22 Every different type of training and serving has these considerations to scale.

1:00:28 We talked about pre-training and RL,

1:00:30 and then inference time scaling- how do you serve

1:00:33 a model that's thinking for an hour to 100 million users?

1:00:35 I don't know about that, but I know that's a hard problem.

1:00:39 In order to give people this intelligence, there's all these systems problems,

1:00:42 and we need more compute and you need more stable compute to do it."-

1:00:46 But you're bullish on all of these kinds of scaling is what I'm hearing.

1:00:49 On the inference, on the reasoning, even on the pre-training?

1:00:54 Yeah, so that's a big can of worms here, but there are basically two...

1:00:58 The knobs are the training and the inference scaling where you can get gains.

1:01:02 In a world where we had, let's say,

1:01:05 infinite compute resources, you want to do all of them.

1:01:08 So you have training, you have inference scaling,

1:01:10 and training is like a hierarchy:

1:01:12 it's pre-training, mid-training, post-training.

1:01:13 Changing the model size, more training data,

1:01:16 training a bigger model gives you more knowledge in the model.

1:01:20 Then the model, let's say, has a better base model.

1:01:23 Back in the day, or still, we call it a foundation model, and it unlocks...

1:01:28 But you don't, let's say, have the model be able to solve

1:01:32 your most complex tasks during pre-training or after pre-training.

1:01:36 You still have these other unlock phases

1:01:38 where you have mid-training or, for example,

1:01:41 post-training with RL that unlocks capabilities that the model

1:01:43 has in terms of knowledge in the pre-training.

1:01:45 And I think, sure, if you do more pre-training,

1:01:50 you get a better base model that you can unlock later.

1:01:53 But like Nathan said, it just becomes too expensive.

1:01:55 We don't have infinite compute, so you have to decide,

1:01:58 do I want to spend that compute more on making the model larger?

1:02:01 It's like a trade-off.

1:02:02 In an ideal world, you want to do all of them.

1:02:05 And I think in that sense, scaling is still pretty much alive.

1:02:08 You would still get a better model,

1:02:09 but like we saw with Claude Opus 4.5, it's just not worth it.

1:02:12 Because you can unlock more performance

1:02:16 with other techniques at that current moment,

1:02:19 especially if you look at inference scaling.

1:02:21 That's one of the biggest gains this year with o1,

1:02:24 where it took a smaller model further than

1:02:29 pre-training a larger model like Claude Opus 4.5.

1:02:31 So I wouldn't say pre-training scaling is dead,

1:02:33 it's just that there are other more attractive ways to scale right now.

1:02:37 But at some point, you will still

1:02:39 want to make some progress on the pre-training.

1:02:42 The thing also to consider is where you want to spend your money.

1:02:46 If you spend it more on the pre-training, it's like a fixed cost.

1:02:50 You train the model, and then it has this capability forever.

1:02:53 You can always use it.

1:02:56 With inference scaling, you don't spend money during training,

1:02:58 you spend money later per query, and then it's also like math.

1:03:02 How long is my model gonna be on the market if I replace it in half a year?

1:03:06 Maybe it's not worth spending $5 million,

1:03:08 $10 million, $100 million on training it longer.

1:03:12 Maybe I will just do more inference scaling and get performance there.

1:03:17 It maybe costs me $2 million in terms of user queries.

1:03:19 It becomes a question of how many users you have and doing the math,

1:03:23 and I think that's also where it's interesting where ChatGPT is in a position.

1:03:26 I think they have a lot of users where they need to go a bit cheaper,

1:03:29 where they have that GPT-5 model that is a bit smaller.

1:03:32 Other companies that have...

1:03:34 Let's say, if your customers have other trade-offs.

1:03:38 For example, there was also the Math Olympiad or some

1:03:41 of these math problems where ChatGPT or they had a proprietary model,

1:03:46 and I'm pretty sure it's just like a model

1:03:49 that has been fine-tuned a little bit more, but most of it was during inference

1:03:53 scaling to achieve peak performance in certain tasks.

1:03:56 need that all the time.

1:03:57 But yeah, long story short, I do think all of these pre-training,

1:04:02 mid-training, post-training, inference scaling,

1:04:04 they are all still things you want to do.

1:04:06 It's just finding—at the moment, in this year,

1:04:08 it's finding the right ratio that gives

1:04:10 you the best bang for the buck, basically.

1:04:13 I think this might be a good place to define pre-training,

1:04:16 mid-training, and post-training.

1:04:18 So, pre-training is the classic training one next token prediction at a time.

1:04:21 You have a big corpus of data.

1:04:23 And Nathan probably also has very interesting insights there because of OLMo 3.

1:04:27 A big portion of the paper focuses on the right data mix.

1:04:30 So, pre-training is essentially just, you know, training cross entropy loss,

1:04:34 training on next token prediction on a vast corpus of internet data,

1:04:39 books, papers and so forth.

1:04:41 It has changed a little bit over the years

1:04:43 in the sense people used to throw in everything they can.

1:04:46 Now, it's not just raw data.

1:04:48 It's also synthetic data where people, let's say, rephrase certain things.

1:04:54 So synthetic data doesn't necessarily mean purely AI-made data.

1:04:58 It's also taking something from an article, a Wikipedia article,

1:05:02 and then rephrasing it as a Q&A question or summarizing it,

1:05:07 rewording it, and making better data that way.

1:05:12 Because I think of it also like with humans.

1:05:15 If someone, let's say, reads a book compared to a messy—no offense,

1:05:19 but like—Reddit post or something like that, I do think

1:05:24 you learn—- There's going to be a post about this, Sebastian.

1:05:28 Some Reddit data is very coveted and excellent for training.

1:05:31 You just have to filter it.

1:05:33 And I think that's the idea.

1:05:35 I think it's like if someone took that and rephrased it in a, let's say,

1:05:40 more concise and structured way,

1:05:42 I think it's higher quality data that gets the LLM there faster.

1:05:46 You get the same LLM out of it at the end,

1:05:49 but it trains faster because if the grammar and the punctuation are correct,

1:05:54 it already learns the correct way,

1:05:57 versus getting information from a messy source

1:05:59 and then learning later how to correct that.

1:06:02 So, I think that is how pre-training evolved and why scaling still works.

1:06:09 It's not just about the amount of data,

1:06:13 it's also the tricks to make that data better for you, in a sense.

1:06:17 And then mid-training is...

1:06:18 I mean, it used to be called pre-training.

1:06:21 I think it's called mid-training because it was awkward

1:06:23 to have pre-training and post-training but nothing in the middle, right?

1:06:26 It sounds a bit weird.

1:06:27 You have pre-training and post-training, but what's the actual training?

1:06:29 So, the mid-training is usually similar to pre-training,

1:06:33 but it's a bit more specialized.

1:06:35 It's the same algorithm, but what you do is you focus,

1:06:39 for example, on long-context documents.

1:06:42 The reason you don't do that during pre-training is

1:06:46 because you don't have that many long context documents.

1:06:49 We have a specific phase.

1:06:50 And one problem of LLMs is still that it's a neural network.

1:06:54 It has the problem of catastrophic forgetting.

1:06:56 So, you teach it something, it forgets other things.

1:06:58 And you wanna...

1:06:59 I mean, it's not 100% forgetting,

1:07:01 but it's like "no free lunch." It's the same with humans.

1:07:04 If you ask me some math I learned 10 years ago,

1:07:07 I would have to look at it again.

1:07:09 Nathan was actually saying that he's consuming so

1:07:11 much content that there's a catastrophic forgetting issue.

1:07:14 Yeah, I'm trying to learn so much about AI,

1:07:16 and it's like I was learning about pre-training parallelism.

1:07:18 I'm like, "I lost something and I don't know

1:07:21 what it was."- I don't want to anthropomorphize LLMs,

1:07:23 but it's the same kind of thing in how humans learn.

1:07:27 I mean, quantity is not always better because you have to be selective.

1:07:32 And mid-training is being selective in terms of quality content at the end.

1:07:36 So the last thing the LLM has seen is the quality stuff.

1:07:39 And then post-training is all the fine-tuning, supervised fine-tuning, DPO,

1:07:45 Reinforcement Learning with Verifiable Rewards (RLVR),

1:07:49 with human feedback, and so forth.

1:07:51 So the refinement stages.

1:07:52 And it's also interesting, it's a cost thing.

1:07:54 You spend a lot of money on pre-training right now.

1:07:57 RL a bit less.

1:07:59 With RL, you don't really teach it knowledge.

1:08:02 It's more like unlocking the knowledge; it's more like a skill learning,

1:08:05 like how to solve problems with the knowledge that it has from pre-training.

1:08:09 There are actually three papers this year,

1:08:11 or last year, 2025, on RL for pre-training.

1:08:14 But I don't think anyone does that in production.

1:08:17 Toy examples for now.

1:08:18 Toy examples, right?

1:08:19 But to generalize, RL post-training is more like the skill unlock,

1:08:23 where pre-training is like soaking up the knowledge.

1:08:27 A few things that could be helpful.

1:08:29 A lot of people think of synthetic data as being bad for training the models.

1:08:34 You mentioned how DeepSeek got almost...

1:08:37 OCR, which is Optical Character Recognition.

1:08:39 A lot of labs did it.

1:08:41 Ai2 had one, Meta had multiple.

1:08:44 And the reason each of these labs has

1:08:47 these is because there are vast amounts of PDFs

1:08:49 and other digital documents on the web that aren't

1:08:52 in formats that are encoded with text easily.

1:08:54 So you use these Almost-OCR, DeepSeek OCR, or what we called our Almost-OCR,

1:08:59 to extract trillions of tokens of candidate data for pre-training.

1:09:04 Pre-training dataset size is measured in trillions of tokens.

1:09:08 Smaller models from researchers can be something like five to 10 trillion.

1:09:11 researchers can be something like five to 10 trillion.

1:09:15 Um, Qwen is documented going up to 50 trillion,

1:09:17 and there are rumors that these closed labs can go to 100 trillion tokens.

1:09:21 Getting this potential data to put in—they have a very big funnel,

1:09:24 and the data you actually train on is a small percentage of this.

1:09:29 This character recognition data would be described

1:09:32 as synthetic data for pre-training in a lab.

1:09:34 And then there's also the fact that ChatGPT now gives wonderful answers,

1:09:38 and you can train on those best answers, and that's synthetic data.

1:09:41 It's very different than early ChatGPT with lots of hallucination data.

1:09:46 when people became grounded in synthetic data.

1:09:49 One interesting question is, if I recall correctly,

1:09:51 OLMo 3 was trained with less data than

1:09:53 specifically some other open-weight models, maybe even OLMo 2.

1:09:56 But you still got better performance,

1:09:58 and that might be one example of how the data helped.

1:10:01 It's mostly down to data quality.

1:10:02 I think if we had more compute, we would train for longer.

1:10:05 I think we'd ultimately see that as something we would want to do.

1:10:09 And especially with big models, you need more compute,

1:10:11 because we talked about having more parameters and we talked about knowledge.

1:10:15 Essentially, there's a ratio where big models can absorb more from data,

1:10:19 and then you get more benefit out of this.

1:10:22 Any logarithmic graph in your mind is like a small

1:10:25 model will level off sooner if you're measuring tons of tokens,

1:10:28 and bigger models need more.

1:10:30 But mostly, we aren't training that big of models right now at AI2,

1:10:34 and getting the highest quality data we can is the natural starting point.

1:10:38 Is there something to be said about the topic of data quality?

1:10:41 Is there some low-hanging fruit there still where the quality could be improved?

1:10:46 It's like turning the crank.

1:10:47 Historically, in the open,

1:10:49 there's been a canonical best pre-training dataset that has moved around

1:10:54 between who has the most recent one or the best recent effort.

1:10:56 Like AI2's Dolma was very early with the first OLMo,

1:10:59 and Hugging Face had FineWeb.

1:11:01 And there's a DCLM project,

1:11:02 which has been kind of like a, which stands for Data Comp Language Model.

1:11:07 There's been Data Comp for other machine learning projects,

1:11:10 and they had a very strong dataset.

1:11:12 And a lot of it is the internet is becoming fairly closed off,

1:11:17 so we have Common Crawl,

1:11:18 which is hundreds of trillions of tokens, and you filter it.

1:11:21 It looks like scientific work where you're training

1:11:24 classifiers and making decisions based on how you

1:11:27 prune down this dataset into the highest quality

1:11:30 stuff and the stuff that suits your tasks.

1:11:33 Previously, language models were tested a lot

1:11:35 more on knowledge and conversational things,

1:11:37 but now they're expected to do math and code.

1:11:40 To train a reasoning model, you need to remix your whole dataset.

1:11:43 And there's a lot of wonderful scientific methods here where you can,

1:11:47 you can take your gigantic dataset,

1:11:49 sample really tiny things from different sources,

1:11:52 such as GitHub, Stack Exchange, Reddit, Wikipedia.

1:11:56 You can sample small things from them,

1:11:57 and train small models on each of these mixes

1:12:00 and measure their performance on your evaluations.

1:12:02 You can just do basic linear regression, and it's like,

1:12:04 "Here's your optimal dataset." But if your evaluations change,

1:12:07 your dataset changes a lot.

1:12:08 So a lot of OLMo 3 was new sources for reasoning to be better at math and code,

1:12:13 and then you do this mixing procedure and it gives you the answer.

1:12:17 I think that's happened at labs this year;

1:12:19 there's new hot things, whether it's coding environments or web navigation,

1:12:23 and you need to bring in new data,

1:12:24 change your whole pre-training so that your post-training can work better.

1:12:28 And that's like the constant evolution and the redetermining

1:12:31 of what they care about for their models.

1:12:35 Are there fun anecdotes of what sources of data

1:12:38 are particularly high quality that we wouldn't expect?

1:12:41 You mentioned Reddit sometimes can be a source.

1:12:45 Reddit was very useful.

1:12:47 I think PDFs is definitely one.

1:12:51 Oh, especially arXiv.

1:12:52 Yeah, so AI2 has run Semantic Scholar for a long time,

1:12:56 which is what you can say is a competitor

1:12:59 to Google Scholar with a lot more features.

1:13:01 And to do this, AI2 has found and scraped a lot of PDFs for openly

1:13:06 accessible papers that might not be behind

1:13:09 the closed walled garden of a certain publisher.

1:13:11 So, truly open scientific PDFs.

1:13:13 And if you sit on all of these and you process it, you can get value out of it.

1:13:17 And I think that like,

1:13:19 a lot of that style of work has been done by the frontier labs did much earlier.

1:13:24 You just need to have a pretty

1:13:26 skilled researcher that understands how things change models;

1:13:30 they bring it in, clean it, and it's a lot of labor.

1:13:33 When frontier labs scale researchers, a lot more goes into data.

1:13:38 If you join a frontier lab and you want to have impact,

1:13:41 the best way to do it is just find new data that's better.

1:13:45 And then, the fancy, glamorous algorithmic things like figuring out how to make

1:13:50 o1 is like the sexiest thought of a scientist.

1:13:52 It's like, "Oh, I figured out how to scale RL." There's a group

1:13:55 that did that, but most of the contribution is like—- On the dataset-

1:13:58 ..."I'm gonna make the data better,"

1:14:00 or, "I'm gonna make the infrastructure better

1:14:01 so everyone on my team can run experiments 5% faster."- At the same time,

1:14:05 I think it's also one of the closest guarded secrets,

1:14:07 what your training data is, for legal reasons.

1:14:09 And so there's also, I think,

1:14:10 a lot of work that goes into hiding what your training data was essentially.

1:14:14 Like training the model to not give

1:14:17 away the sources because you have legal reasons.

1:14:19 The other thing, to be complete,

1:14:20 is that some people are trying to train on only licensed data,

1:14:23 whereas Common Crawl is a scrape of the whole internet.

1:14:26 So if I host multiple websites, I'm happy to have them train language models,

1:14:32 but I'm not explicitly licensing what governs it.

1:14:35 And therefore, Common Crawl is largely unlicensed,

1:14:38 which means that your consent really hasn't

1:14:41 been provided for how to use the data.

1:14:43 There's another idea where you can train language

1:14:44 models only on data that has been licensed explicitly,

1:14:47 so that the kind of governing contract is provided,

1:14:50 and I'm not sure if Apertus is the copyright thing or the license thing.

1:14:53 I know that the reason that they did it was for an EU compliance thing,

1:14:56 where they wanted to make sure that their model fit one of those checks.

1:15:05 On that note, there's also the distinction in licensing.

1:15:09 Some people just purchase the license.

1:15:12 Let's say they buy an Amazon Kindle book,

1:15:15 or a Manning book, and then use that in training.

1:15:18 That is a gray zone 'cause you paid

1:15:19 for the content and you might want to train on it.

1:15:22 But then there are also restrictions where even that shouldn't be allowed.

1:15:25 And so that is where it gets a bit fuzzy.

1:15:28 And yeah, I think that is right now still a hot topic.

1:15:33 Big companies like OpenAI approached private companies

1:15:36 for their proprietary data and private companies,

1:15:39 they become more and more, let's say,

1:15:42 protective of their data because they know, "Okay,

1:15:44 this is going to be my moat in a few

1:15:47 years." And I do think that's like the interesting question,

1:15:50 where if LLMs become more commoditized,

1:15:53 and I think a lot of people learn about LLMs,

1:15:56 there will be a lot more people able to train LLMs.

1:15:58 Of course, there are infrastructure challenges.

1:16:00 But if you think of big industries like pharmaceutical industries,

1:16:04 law, finance industries, I do think they, at some point,

1:16:07 will hire people from other frontier labs

1:16:10 to build their in-house models on their proprietary data,

1:16:13 which will be then, again,

1:16:14 another unlock with pre-training that is currently not there.

1:16:17 Because even if you wanted to, you can't get that data.

1:16:21 You can't get access to clinical trials most

1:16:23 of the time and these types of things.

1:16:25 So, I do think scaling, in that sense, might be still pretty much alive

1:16:28 if you also look in domain-specific applications,

1:16:31 because we are still right now, in this year,

1:16:33 just looking at general purpose LLMs on, on ChatGPT, Anthropic and so forth.

1:16:37 They are just general purpose, they're not even, I think,

1:16:40 scratching the surface of what an LLM can do if

1:16:43 it is really specifically trained and designed for a specific task.

1:16:47 I think on the data thing—this is one of the things that happened in 2025,

1:16:50 and we totally forget it—is Anthropic lost

1:16:52 in court and owed $1.5 billion to authors.

1:16:55 Anthropic, I think, bought thousands of books and scanned them

1:16:59 and was cleared legally for that because they bought the books,

1:17:03 and that is kind of going through the system.

1:17:04 And then the other side, they also torrented some books,

1:17:07 and I think this torrenting was the path where the court said

1:17:10 that they were then culpable to pay these billions of dollars to authors,

1:17:13 which is just such a mind-boggling lawsuit that kind of just came and went.

1:17:17 That is so much money-...

1:17:20 from the VC ecosystem.

1:17:22 These are court cases that will define the future of human civilization,

1:17:25 because it's clear that data drives a lot

1:17:27 of this, and there's this very complicated human tension of...

1:17:30 I mean, you can empathize.

1:17:33 You're both authors.

1:17:34 And there's some degree to which, I mean,

1:17:36 you put your heart and soul and your sweat

1:17:39 and tears into the writing that you do.

1:17:42 It feels a little bit like theft for somebody

1:17:46 to train your data without giving you credit.

1:17:49 And there are, like Nathan said, also two layers to it.

1:17:51 Someone might buy the book and then train on it,

1:17:54 which could be argued fair or not fair,

1:17:56 but then there are the straight-up companies who use

1:18:00 pirated books where it's not even compensating the author.

1:18:03 That is, I think, where people got

1:18:04 a bit angry about it specifically, I would say.

1:18:06 Yeah, but there has to be some kind of compensation scheme.

1:18:09 This is like moving towards-...

1:18:11 towards something like Spotify streaming did-...

1:18:13 originally for music.

1:18:14 You know, what does that-...

1:18:15 compensation look like?

1:18:16 You have to define those kinds of models.

1:18:17 You have to think through all of that.

1:18:19 One other thing I think people are generally curious about,

1:18:22 I'd love to get your thoughts, as LLMs are used more and more.

1:18:26 If you look at even arXiv, but GitHub,

1:18:29 more and more of the data is generated by LLMs.

1:18:32 What do you do in that kind of world?

1:18:36 How big of a problem is that?

1:18:39 Largest problem's the infrastructure and systems,

1:18:41 but from an AI point of view, it's kind of inevitable.

1:18:45 So it's basically LLM-generated data

1:18:47 that's curated by humans essentially, right?

1:18:49 Yes, and I think that a lot

1:18:51 of open source contributors are legitimately burning out.

1:18:53 If you have a popular open source repo,

1:18:55 somebody's like, "Oh, I want to do open source AI.

1:18:57 It's good for my career," and they just

1:19:00 vibe- -code something and they throw it in.

1:19:02 You might get more of this-- I have a--- than I do.

1:19:05 Yeah, so I have actually a case study here.

1:19:09 I have a repository called MLxtend that I

1:19:11 developed as a student around 10 years ago,

1:19:14 and it is a reasonably popular library still for certain algorithms,

1:19:18 I think especially like frequent data mining stuff.

1:19:22 And there were recently two or three people who submitted

1:19:25 a lot of PRs in a very short amount of time.

1:19:28 I do think LLMs have been involved in submitting these PRs.

1:19:31 Me, as the maintainer, there are two things.

1:19:33 First, I'm a bit overwhelmed.

1:19:35 I don't have time to read through it because,

1:19:37 especially as an older library, that is not a priority for me.

1:19:40 At the same time, I kind of also appreciate it because

1:19:43 I think something people forget is it's not just using the LLM.

1:19:46 There's still a human layer that verifies something,

1:19:48 and that is in a sense also how data is labeled, right?

1:19:53 One of the most expensive things is getting

1:19:56 labeled data for RL from human feedback phases.

1:19:59 And this is kind of like that, where it goes through phases,

1:20:03 and then you actually get higher quality data out of it.

1:20:06 And so I don't mind it in a sense.

1:20:08 It can feel overwhelming, but I do think there is also value in it.

1:20:12 It feels like there's a fundamental

1:20:14 difference between raw LLM-generated data and LLM-generated

1:20:16 data with a human in the loop that does some kind of verification,

1:20:21 even if that verification is a small percent of the lines of code.

1:20:26 I think this goes with anything where people think also sometimes, "Oh, yeah.

1:20:31 I can just use an LLM to learn about XYZ," which is true.

1:20:34 You can, but there might be a person who is

1:20:37 an expert who might have used an LLM to write specific code.

1:20:41 There is kind of like this human work that went into it to make it nice,

1:20:45 throwing out the not-so-nice parts to kind of pre-digest it for you,

1:20:50 and that saves you time.

1:20:52 I think that's the value-add,

1:20:53 where you have someone filtering things or even using the LLMs correctly.

1:20:59 This is still labor that you get for free.

1:21:02 For example, if you read a Substack article,

1:21:04 I could maybe ask an LLM to give me opinions

1:21:07 on that, but I wouldn't even know what to ask.

1:21:10 And I think there is still value in reading that article

1:21:13 compared to me going to the LLM because you are the expert.

1:21:17 You select what knowledge is actually spot on, should be included,

1:21:20 and you give me this very...

1:21:23 this executive summary.

1:21:24 And this is a huge value-add because now I don't have

1:21:28 to waste three to five hours to go through this myself,

1:21:31 maybe get some incorrect information and so on.

1:21:34 And so I think that's also where the future

1:21:37 still is for writers even though there are LLMs that...

1:21:41 Can kind of save you time.

1:21:44 It's kinda fascinating to watch.

1:21:45 I'm sure you guys do this, but for me,

1:21:48 I look at the difference between a summary and the original content.

1:21:53 Even if it's a page-long summary of a page-long content,

1:21:57 it's interesting to see how the LLM-based summary takes the edge off.

1:22:03 What is the signal it removes from the thing?

1:22:07 The voice is what I talk about a lot.

1:22:09 Voice?

1:22:10 Well, voice...

1:22:10 I would love to hear what you mean by voice,

1:22:13 but sometimes there's literally insights.

1:22:16 By removing an insight, you're changing the meaning of the thing.

1:22:21 So I'm continuously disappointed how bad LLMs

1:22:24 are at really getting to the core insights, which is what a great summary does.

1:22:31 Yet even if I have these extremely elaborate prompts

1:22:36 where I'm really trying to dig for the insights,

1:22:40 it's still not quite there, which...

1:22:42 I mean, that's a whole deep philosophical question about what is human knowledge

1:22:46 and wisdom and what does it mean to be insightful and so on.

1:22:49 But when you talk about the voice, what do you mean?

1:22:52 When I write, I think a lot of what I'm trying to do

1:22:55 is take what you think as a researcher, which is very raw.

1:22:59 A researcher is trying to encapsulate an idea at the frontier

1:23:02 of their understanding and they're trying to put what is a feeling into words.

1:23:07 I try to do this in my writing,

1:23:11 which makes it come across as raw but also high-information

1:23:14 in a way that some people will get it and some won't.

1:23:17 And that's the nature of research.

1:23:19 And language models don't do this well.

1:23:21 They're all trained with Reinforcement Learning from Human Feedback,

1:23:25 which takes feedback from many people

1:23:27 and averages how the model behaves from this.

1:23:30 And I think it's going to be hard for a model

1:23:34 to be very incisive when there's that sort of filter.

1:23:37 This is a wonderful fundamental problem for researchers in RLHF.

1:23:43 This provides so much utility in making the models better,

1:23:47 but also the problem formulation is kind of...

1:23:51 there's this knot in it that you can't get past.

1:23:55 These language models don't have this prior

1:23:57 in their deep expression they're trying to get at.

1:24:00 I don't think it's impossible.

1:24:01 There are stories of models that really shock people.

1:24:04 Like, I think of...

1:24:05 I would love to have tried Bing Sydney.

1:24:08 Did that have more voice?

1:24:10 Because it would so often go off the rails,

1:24:13 which is historically obviously a scary way—like telling a reporter to leave

1:24:17 his wife—is a crazy model to potentially put in general adoption.

1:24:21 But that's kind of like a trade-off;

1:24:23 is this RLHF process in some ways adding limitations?

1:24:28 That's a terrifying place to be as one of these frontier labs and companies,

1:24:33 because millions of people are using them.

1:24:36 There was a lot of backlash last year with GPT-4o getting removed,

1:24:39 and I've personally never used the model,

1:24:41 but I've talked to people at OpenAI where they get emails from users

1:24:47 that might be detecting subtle differences

1:24:50 in the deployments in the middle of the night.

1:24:52 And they email them like, "My friend is different." And they find

1:24:56 these employees' emails and send them things because they

1:24:58 are so attached to this set of model

1:25:02 weights and configuration that is deployed to the users.

1:25:05 We see this with TikTok.

1:25:06 You open it...

1:25:07 I don't use TikTok, but supposedly in five minutes the algorithm gets you.

1:25:11 It's locked in.

1:25:12 And those are language models doing recommendations.

1:25:15 I think there are ways you can do this.

1:25:18 Within five minutes of chatting with it, the model just gets you.

1:25:21 And that is something that people aren't really ready for.

1:25:26 Like, don't give that to kids.

1:25:28 At least until we know what's happening.

1:25:30 But there's also going to be this mechanism...

1:25:32 What's going to happen with these LLMs as they're used more and more...

1:25:36 Unfortunately, the nature of the human

1:25:37 condition is such that people commit suicide.

1:25:39 And so what journalists will do is they

1:25:42 will report extensively on the people who commit suicide.

1:25:45 And they would very likely link it to the LLMs

1:25:48 because they have that data about the conversations.

1:25:50 If you're really struggling in your life, if you're depressed,

1:25:54 if you're thinking about suicide,

1:25:55 you're going to probably talk to LLMs about it.

1:25:58 And so what journalists will do is say, "Well,

1:26:01 the suicide was committed because of the LLM."

1:26:03 And that's going to lead to the companies,

1:26:06 because of legal issues and so on, more and more taking the edge off of the LLM.

1:26:13 So it's going to be as generic as possible.

1:26:15 It's so difficult to operate in this space because you don't

1:26:19 want an LLM to cause harm to humans at that level,

1:26:23 but also this is the nature of the human experience,

1:26:27 is to have a rich conversation,

1:26:29 a fulfilling conversation, one that challenges you from which you grow.

1:26:33 You need that edge.

1:26:35 And that's something extremely difficult for AI researchers

1:26:39 on the RLHF front to actually have to solve,

1:26:44 because you're dealing with the human condition.

1:26:47 A lot of researchers at these companies are so well-motivated,

1:26:50 and definitely Anthropic and OpenAI culturally want to do good for the world.

1:26:56 And it's such a—I'm like, "Ooh,

1:26:58 I don't wanna work on this," because on the one hand,

1:27:01 a lot of people see AI as a health ally,

1:27:04 as somebody they can talk to about their health confidentially,

1:27:08 but then it bleeds all the way into this, like talking about mental health,

1:27:14 where it's heartbreaking that this will be

1:27:17 the thing where somebody goes over the edge, but other people might be saved.

1:27:21 And I'm like, "I don't..." As a researcher, it's like,

1:27:25 I don't want to train image generation

1:27:26 models and release them openly because I don't

1:27:29 want to enable somebody to have a tool

1:27:31 on their laptop that can harm other people.

1:27:34 I don't have the infrastructure in my company to do that safely.

1:27:37 But there's a lot of areas like this where it

1:27:40 just needs people that will approach it with complexity and conviction.

1:27:44 It's just such a hard problem.

1:27:48 But also, we as a society, as users of these technologies,

1:27:50 need to make sure that we're having the complicated conversation about it versus

1:27:54 just fearmongering— that Big Tech is causing

1:27:57 harm to humans or stealing your data.

1:27:59 It's more complicated than that.

1:28:01 And you're right.

1:28:03 There's a very large number of people inside these companies,

1:28:05 many of whom I know, who deeply care about helping people.

1:28:09 They are considering the full human experience of people from across the world,

1:28:13 not just Silicon Valley.

1:28:14 People across the United States and the world, what their needs are.

1:28:18 It's really difficult to design this one system that is able

1:28:23 to help all these different kinds of people across different age groups,

1:28:26 cultures, and mental states.

1:28:31 I wish that the timing of AI was different relative

1:28:34 to the relationship of Big Tech to the average person.

1:28:37 Big Tech's reputation was so low, and with how AI is so expensive,

1:28:40 it's inevitably going to be a Big Tech thing.

1:28:42 It takes so many resources, and people say the US is,

1:28:46 "betting the economy on AI" with this build-out.

1:28:48 To have these be intertwined at the same

1:28:51 time makes for such a hard communication environment.

1:28:54 It would be good for me to go talk to more people

1:28:56 in the world who hate Big Tech and see AI as a continuation of this.

1:29:03 And one of the things you recommend,

1:29:05 one of the antidotes that you talk about, is to find agency in this system.

1:29:11 As opposed to sitting back in a powerless way and consuming

1:29:16 the AI slop as it rapidly takes over the internet.

1:29:20 Find agency by using AI to build stuff, build apps, build...

1:29:25 One, that actually helps you build intuition, but two,

1:29:29 it's empowering because you can understand how it works,

1:29:32 what the weaknesses are.

1:29:34 It gives your voice power to say, "This is a bad use of the technology,

1:29:38 and this is a good use." And you're more plugged into the system then,

1:29:44 so you can understand it better and you can steer it better as a consumer.

1:29:48 I think that's a good point you brought up about agency.

1:29:51 Instead of ignoring it and saying, "Okay,

1:29:52 I'm not going to use it," I think it's probably long-term healthier to say,

1:29:57 "Okay, it's out there.

1:29:58 I can't put it back." when they came out.

1:30:01 How do I make best use of it, and how does it help me to up-level myself?

1:30:06 The one thing I worry about here, though, is,

1:30:08 if you just fully use it for something you love to do,

1:30:10 the thing you love to do is no longer there.

1:30:13 And that could potentially, I feel, lead to burnout.

1:30:16 For example, if I use an LLM to do all my coding for me, now there's no coding.

1:30:21 I'm just managing something that is coding for me.

1:30:23 Two years later, let's say, if I just do that eight hours a day,

1:30:27 having something code for me, do I feel fulfilled still?

1:30:36 Is this hurting me in terms of being excited about my job,

1:30:39 excited about what I'm doing?

1:30:40 Am I still proud to build something?

1:30:43 On that topic of enjoyment, it's quite interesting.

1:30:45 We should just throw this in there,

1:30:47 that there's this recent survey of about 791 professional developers,

1:30:52 meaning 10-plus years of experience.

1:30:56 That's a long time.

1:30:58 As a junior developer?

1:31:01 Yeah, in this day and age.

1:31:03 So, there's also many fronts that are surprising.

1:31:06 They break it down by junior and senior developers.

1:31:10 But, I mean, it just shows that both junior

1:31:14 and senior developers use AI-generated code in code they ship.

1:31:22 So this is not just for fun or intermediate learning things.

1:31:26 This is code they ship.

1:31:28 25%—like, most of them use around 50% or more.

1:31:32 And what's interesting is,

1:31:33 for the category of over 50% of your code that you ship is AI-generated,

1:31:38 senior developers are much more likely to do so.

1:31:42 But you don't want AI to take away the thing you love.

1:31:46 I think this speaks to my experience, these results I'm about to say.

1:31:49 Together, about 80% of people find it either somewhat more enjoyable

1:31:54 or significantly more enjoyable to use AI as part of the work.

1:31:59 I think it depends on the task.

1:32:01 From my personal usage, for example,

1:32:03 I have a website where I sometimes tweak things.

1:32:07 I personally don't enjoy this.

1:32:09 So in that sense, if the AI can help me

1:32:13 to implement something on my website, I'm all for it.

1:32:15 It's great.

1:32:16 But then, at the same time,

1:32:17 when I solve a complex problem— well, if there's a bug,

1:32:21 and I hunt this bug, and I find it, it's the best feeling in the world.

1:32:26 You feel great.

1:32:27 But now, if you don't even think about the bug,

1:32:31 you just go directly to the LLM, well,

1:32:33 you never have this kind of feeling, right?

1:32:35 But then there could be the middle ground where, well,

1:32:39 you try yourself, you can't find it, you use the LLM,

1:32:42 and then you don't get frustrated because it helps

1:32:44 you and you move on to something that you enjoy.

1:32:46 And so, looking at these statistics, I think what is not factored in is

1:32:51 that it's averaging over all the different scenarios.

1:32:54 We don't know if it's for the core task or if

1:32:59 it's for something mundane that people would not have enjoyed otherwise.

1:33:02 So, in a sense, AI is really great

1:33:04 for doing mundane things that take a lot of work.

1:33:06 For example, my wife the other day—she has

1:33:09 a podcast for book discussions, a book club,

1:33:13 and she was transferring the show notes from Spotify to YouTube,

1:33:19 and then the links somehow broke.

1:33:21 And she had in some episodes, because it is so many books, like 100 links,

1:33:25 and it would have been really painful to go in there and fix each link manually.

1:33:29 So I suggested, "Hey,

1:33:30 let's try ChatGPT." We copied the text into ChatGPT, and it fixed them.

1:33:35 Instead of two hours going from link to link fixing

1:33:39 that, it made that type of work much more seamless.

1:33:43 I think everyone has a use case where AI is

1:33:47 useful for something that would be really boring, really mundane.

1:33:51 For me personally, since we're talking about coding,

1:33:56 and you mentioned debugging...

1:33:58 the source of enjoyment for me, more on the Cursor side than Claude Code,

1:34:02 is that I have a friend, I have a pair programmer.

1:34:08 It's less lonely.

1:34:10 You made debugging sound like this great joy.

1:34:15 No, I would say debugging is like a drink

1:34:18 of water after you've been going through a desert for days.

1:34:22 You skip the whole desert part where you're suffering.

1:34:27 Sometimes it's nice to have a friend who can't really find the bug,

1:34:31 but can give you some intuition about the code,

1:34:35 and together you're going through the desert and finding that drink of water.

1:34:40 At least for me, maybe it speaks

1:34:43 to the loneliness of the programming experience.

1:34:45 That is a source of joy.

1:34:48 It's maybe also related to delayed gratification.

1:34:51 I'm a person who, even as a kid,

1:34:54 I liked the idea of Christmas presents—having them,

1:34:57 getting them—better than actually receiving the presents.

1:35:01 I would look forward to the day I get the presents,

1:35:04 but then it's over and I'm disappointed.

1:35:06 And maybe it's the same with food.

1:35:08 I think food tastes better when you're really hungry.

1:35:12 You're right, with debugging, it is not always great.

1:35:17 It's often frustrating, but then if you can solve it, then it's great.

1:35:23 But there's a sweet Goldilocks zone;

1:35:24 if it's too hard, then it's just wasting your time.

1:35:28 But I think another challenge is how will people learn?

1:35:33 We looked at the chart and saw that more senior

1:35:37 developers are shipping more AI-generated code than the junior ones.

1:35:42 It's very interesting,

1:35:42 because intuitively you would think it's the junior developers

1:35:45 because they don't know how to do the thing yet,

1:35:49 and so they use AI to do that thing.

1:35:51 It could mean the AI is not good enough yet to solve that task,

1:35:55 but it could also mean experts are more effective at using it.

1:35:59 They know how to use it better,

1:36:02 review the code, and then they trust the code more.

1:36:06 One issue for society in the future will be:

1:36:08 how do you become an expert if you never try to do the thing yourself?

1:36:14 One way I always learned is by trying things myself.

1:36:18 If you look at math textbooks and the solutions,

1:36:21 you learn something, but you learn actually better if you try first.

1:36:27 Then you appreciate the solution differently because you

1:36:29 know how to put it into your mental framework.

1:36:32 If LLMs are here all the time,

1:36:35 would you actually go to the length of struggling?

1:36:38 Would you be willing to struggle?

1:36:40 Struggle is not nice, right?

1:36:42 But if you use the LLM to do everything,

1:36:44 at some point you will never really take the next step,

1:36:47 and then you will maybe not get that unlock

1:36:50 that you would get as an expert using an LLM.

1:36:52 So, I think there's like a Goldilocks sweet spot where maybe the trick

1:36:58 here is you make dedicated offline time where you study two hours a day,

1:37:01 and the rest of the day use LLMs.

1:37:03 But I think it's important also for people to still invest in themselves,

1:37:07 in my opinion, to not just LLM everything.

1:37:11 Yeah, as a civilization, we each individually have to find that Goldilocks zone.

1:37:16 And in the programming context as developers.

1:37:18 Now, we've had this fascinating conversation

1:37:20 that started with pre-training and mid-training.

1:37:24 Let's get to post-training.

1:37:25 A lot of fun stuff in post-training.

1:37:28 So, what are some of the interesting ideas in post-training?

1:37:32 The biggest one from 2025 is

1:37:34 learning this reinforcement learning with verifiable rewards.

1:37:37 You can scale up the training there,

1:37:39 which means doing a lot of this kind of iterative generate-grade loop,

1:37:43 and that lets the models learn both interesting

1:37:47 behaviors on the tool use and software side.

1:37:49 This could be searching, running commands on their own and seeing the outputs,

1:37:53 and then also that training enables this inference time scaling very nicely.

1:37:57 And it just turned out that this paradigm was very nicely linked,

1:38:02 where this kind of RL training enables inference time scaling.

1:38:05 But inference time scaling could have been found in different ways.

1:38:07 So, it was kind of this perfect storm where the models change a lot,

1:38:10 and the way that they're trained is a major factor in doing so.

1:38:15 And this has changed how people approach post-training dramatically.

1:38:20 Can you describe RLVR, popularized by DeepSeek R1?

1:38:23 Can you describe how it works?

1:38:26 Yeah.

1:38:26 Fun fact: I was on the team that came up with the term RLVR,

1:38:29 which is from our Tulu 3 work before DeepSeek.

1:38:33 We don't take a lot of credit for being the people to popularize the scaling RL,

1:38:37 but as fun as what academics get,

1:38:40 as an aside, is the ability to name and influence the discourse,

1:38:44 because the closed labs can only say so much.

1:38:47 That one of the things you can do as an academic

1:38:49 is you might not have the compute to train the model,

1:38:51 but you can frame things in a way that ends up being described

1:38:55 as a community coming together around this RLVR term, which is very fun.

1:39:00 And then DeepSeek are the people that did the training breakthrough,

1:39:03 which is, they scaled the reinforcement learning.

1:39:06 You have the model generate answers and then

1:39:09 grade the completion if it was right,

1:39:11 and then that accuracy is your reward for reinforcement learning.

1:39:16 So reinforcement learning is classically an agent that acts in an environment,

1:39:20 and the environment gives it a state and a reward back,

1:39:24 and you try to maximize this reward.

1:39:26 In the case of language models,

1:39:28 the reward is normally accuracy on a set of verifiable tasks,

1:39:31 whether it's math problems or coding tasks.

1:39:34 And it starts to get blurry with things like factual domains.

1:39:38 That is also, in some ways, verifiable, or constraints on your instruction,

1:39:44 like respond only with words that start with A." All

1:39:47 of these things are verifiable in some way,

1:39:50 and the core idea of this is you find a lot more of these problems

1:39:55 that are verifiable and you let the model

1:39:56 try it many times while taking these RL steps, these RL gradient updates.

1:40:02 The infrastructure evolved from reinforcement learning from human feedback,

1:40:06 where in that era the score they were trying

1:40:10 to optimize was a learned reward model of human preferences.

1:40:13 So you kind of changed the problem domains

1:40:15 and that let the optimization go on to much bigger scales,

1:40:19 which kind of kickstarted a major change in what

1:40:22 the models can do and how people use them.

1:40:25 What kind of domains is RLVR amenable to?

1:40:28 Math and code are the famous ones,

1:40:30 and then there's a lot of work on what is called the rubrics,

1:40:34 which is related to a word people might have heard: LLM-as-a-judge.

1:40:38 For each problem, I'll have a set of problems in my dataset.

1:40:42 I will then have an LLM and ask it,

1:40:45 "What would a good answer to this problem look like?" And then you could

1:40:49 try the problem over and over again and assign a score based on this rubric.

1:40:54 That's not necessarily verifiable like math and code domains,

1:40:57 but this rubrics idea and other scientific problems that might

1:41:00 be a little bit more vague is where the attention is,

1:41:04 where they're trying to push this set of methods into these kind

1:41:07 of more open-ended domains so the models can learn a lot more.

1:41:11 I think that's called reinforcement learning with AI feedback, right?

1:41:14 That's the older term for it coined in Anthropic's Constitutional AI paper.

1:41:18 It's like a lot of these things come in cycles.

1:41:21 Also, just one step back for RLVR.

1:41:24 I think the interesting thing here is that you

1:41:27 ask the LLM a, let's say, math question, and then you know the correct answer,

1:41:31 and you let the LLM, as you said, figure it out.

1:41:35 How it does it—you don't constrain it much.

1:41:37 There are some constraints like "use the same language,

1:41:40 don't switch between Spanish and English."

1:41:42 But let's say you're pretty much hands-off.

1:41:45 You only give the question and the answer,

1:41:47 and then the LLM has the task to arrive at the right answer,

1:41:51 but the beautiful thing here is what happens in practice:

1:41:55 the LLM will do a step-by-step description,

1:41:57 like as a student or as a mathematician would derive the solution.

1:42:02 It will use those steps, and that helps the model to improve its own accuracy.

1:42:07 And then, like you said, the inference scaling.

1:42:11 Inference scaling loosely means spending more compute during inference,

1:42:16 and here the inference scaling is that the model would use more tokens.

1:42:22 In the DeepSeek R1 paper,

1:42:24 they showed the longer they train the model, the longer the responses are.

1:42:28 They grow over time.

1:42:29 They use more tokens, so it becomes more expensive.

1:42:32 It becomes expensive for simple tasks, but these explanations help accuracy.

1:42:36 There are also papers showing what the model explains does not

1:42:40 necessarily have to be correct or maybe it's unrelated to the answer,

1:42:44 but for some reason, it still helps the model that it is explaining.

1:42:48 And I think it's also—again, I don't want to anthropomorphize these LLMs,

1:42:52 but it's kind of like how we humans operate.

1:42:55 If there's a complex math problem in a math class, class,

1:42:59 you usually have a note paper and you do it step by step.

1:43:02 You cross out things.

1:43:03 And the model also self-corrects, and that was,

1:43:05 I think, the aha moment in the DeepSeek R1 paper.

1:43:08 They called it the aha moment because the model

1:43:10 itself recognized it made a mistake and then said, "Ah, I did something wrong,

1:43:13 let me try again." And I think that's just so cool that this falls out

1:43:18 of just giving it the correct answer and having it figure out how to do it,

1:43:22 that it kind of does in a sense what a human would do.

1:43:26 Although LLMs don't think like humans,

1:43:28 it's kind of like an interesting coincidence and it...

1:43:31 And the other nice side effect is it's

1:43:33 great for us humans often to see these steps.

1:43:36 It builds trust, but also we us humans to see these steps.

1:43:38 It builds trust, but also we learn and can double check things.

1:43:40 There's a lot in here.

1:43:40 I think some of the debate...

1:43:42 There's been a lot of debate this year on if the language models like these...

1:43:45 I think the aha moments are kind of fake

1:43:48 because in pre-training you essentially have seen the whole internet.

1:43:51 so you have definitely seen people explaining their work,

1:43:54 even verbally, like a transcript of a math lecture.

1:43:57 "You try this, oh, I messed this up."

1:43:58 And what RLVR is very good at doing is amplifying

1:44:02 these behaviors because they're very useful in enabling

1:44:04 the model to think longer and to check its work.

1:44:07 And I agree that it is very beautiful that this training kind of...

1:44:11 The model learns to amplify this in a way

1:44:13 that is so useful for the final answers being better.

1:44:17 I can give you also a hands-on example.

1:44:18 I was training the Qwen 3 base model with RLVR on MATH-500.

1:44:22 The base model had an accuracy of about 15%.

1:44:26 Just 50 steps, like in a few minutes with RLVR,

1:44:30 the model went from 15% to 50% accuracy.

1:44:33 And the model...

1:44:34 You can't tell me it's learning anything fundamentally about math

1:44:38 in- The Qwen example is weird because there've been two papers this year,

1:44:41 one of which I was on, about data contamination in Qwen

1:44:44 and specifically that they train on a lot of this special

1:44:47 mid-training phase that we should take a minute on, because it's

1:44:50 weird because they train on problems that are almost identical to MATH.

1:44:53 Exactly.

1:44:54 And so you can see that basically the RL,

1:44:57 it's not teaching the model any new knowledge about math.

1:44:59 You can't do that in 50 steps.

1:45:01 So the knowledge is already there,

1:45:02 in the pre-training, you're just unlocking it.

1:45:04 I still disagree with the premise because there's a lot

1:45:06 of weird complexities that you can't prove because one

1:45:10 of the things that points to weirdness is that if

1:45:12 you take the Qwen 3 so-called base model and you...

1:45:15 You could Google like "math dataset, Hugging Face",

1:45:18 and you could take a problem and what you do if you put it into Qwen 3 base...

1:45:23 All these math problems have words,

1:45:24 so it'd be like "Alice has five apples and takes one...

1:45:27 and gives three to whoever," and there are these word problems.

1:45:30 With these Qwen-based models, why people are suspicious of them is if you

1:45:34 change the numbers but keep the words- Qwen will produce,

1:45:38 without tools, will produce a very

1:45:41 high accuracy decimal representation of the answer, which means there's some...

1:45:45 At some time, it was shown problems that were almost identical to the test set,

1:45:50 and it was using tools to get a very high precision answer,

1:45:53 but a language model without tools will never actually have this.

1:45:57 So it's kind of been this big debate in the research community:

1:46:01 how much of these reinforcement learning papers that are

1:46:04 training on Qwen and measuring specifically on this math benchmark,

1:46:07 where there's been multiple papers talking about contamination,

1:46:10 is like, how much can you believe them?

1:46:12 And I think this is what caused the reputation of RLVR being about formatting,

1:46:15 because you can get these gains so quickly,

1:46:18 therefore it must already be in the model.

1:46:20 But there's a lot of complexity here that we...

1:46:22 It's not really like controlled experimentation, so we don't really know.

1:46:27 But if it weren't true, I would say distillation wouldn't work, right?

1:46:30 I mean, distillation can work to some extent, but the thing is that is, I think,

1:46:35 the biggest problem,

1:46:36 and I research this contamination because we don't know what's in the data.

1:46:38 Unless you have a new dataset, it is really impossible.

1:46:42 And the same, you mentioned the math dataset,

1:46:44 where you have a question and then answer and an explanation is given,

1:46:48 but then also even something simpler like MMLU,

1:46:51 which is a multiple-choice benchmark.

1:46:53 If you just change the format slightly, like, I don't know,

1:46:57 if you use a dot instead of a parenthesis

1:47:00 or something like that, the model accuracy will vastly differ.

1:47:05 I think that that could be like a model issue rather than a general issue.

1:47:09 It's not even malicious by the developers of the LLM, like, "Hey,

1:47:11 we want to cheat at that benchmark." It has seen something at some point.

1:47:14 I think the only fair way to evaluate an LLM is to have

1:47:18 a new benchmark that is after the cutoff date when the LLM was deployed.

1:47:22 Can we lay out what would be the recipe

1:47:25 of all the things that go into post-training?

1:47:28 And you mentioned RLVR was a really exciting, effective thing.

1:47:32 Maybe we should elaborate.

1:47:34 RLHF still has a really important component to play.

1:47:37 What kind of other ideas are there on post-training?

1:47:40 I think you can kind of take this in order.

1:47:41 I think you could view it as what made o1,

1:47:45 which is this first reasoning model, possible, or what will the latest model be?

1:47:49 And they actually...

1:47:51 You're going to have similar interventions

1:47:53 at these, where you start with mid-training,

1:47:56 and the thing that is rumored to enable

1:47:59 o1 and similar models is really careful data curation,

1:48:02 where you're providing a broad set of what is called reasoning traces,

1:48:07 which is just the model generating words

1:48:10 in a forward process that is reflecting,

1:48:13 like breaking down a problem into intermediate steps and trying to solve them.

1:48:16 So at mid-training, you need to have data that is similar

1:48:20 to this to make it so that when you move into post-training,

1:48:23 primarily with these verifiable rewards, it can learn.

1:48:27 And then what is happening today is you're figuring out which problems

1:48:32 to give the model and how out which problems to give the model

1:48:35 and how long you can train it for and how much inference

1:48:37 you can enable the model to use when solving these verifiable problems.

1:48:41 So as models get better,

1:48:43 certain problems models get better, certain problems are no longer...

1:48:47 The model will solve them 100% of the time,

1:48:48 and therefore there's very little signal in this.

1:48:51 If we look at the GRPO equation, this one is famous for this because essentially

1:48:55 the reward given to the agent is based on how

1:48:59 good a given action—an action is a completion—is

1:49:02 relative to the other answers to that same problem.

1:49:05 So if all the problems get the same answer,

1:49:07 there's no signal in these types of algorithms.

1:49:08 So what they're doing is they're finding harder problems,

1:49:11 which is why you hear about things like scientific domains,

1:49:14 where it's so hard to get anything right.

1:49:17 If you have a lab or something,

1:49:19 it just generates so many tokens or much harder software problems.

1:49:22 So the frontier models are all pushing into these harder domains when they

1:49:26 can train on more problems and the model will learn more skills at once.

1:49:30 The RLHF link to this is that RLHF has been

1:49:33 and still is kind of like the finishing touch on the models,

1:49:36 where it makes the models more useful

1:49:38 by improving the organization or style or tone.

1:49:41 There are different things that resonate with different audiences,

1:49:43 like some people like a really quirky model

1:49:45 and RLHF could be good at enabling that personality,

1:49:49 and some people hate this markdown bulleted list thing that the models do,

1:49:54 but it's actually really good for quickly parsing information.

1:49:57 In RLHF, this human feedback stage is really great for putting

1:50:02 this into the model at the end of the day.

1:50:05 It's what made ChatGPT so magical for people.

1:50:07 And that use has actually remained fairly stable.

1:50:10 This formatting can also help the models

1:50:14 get better at math problems, for example.

1:50:16 So it's like the border between style and formatting,

1:50:20 and like the method that you use to answer a problem is

1:50:24 actually all very closely linked in terms of when you're training these models,

1:50:29 which is why RLHF can still make a model better at math,

1:50:32 but these verifiable domains are a much more direct process

1:50:35 to doing this because it makes more sense with the problem formulation,

1:50:39 which is why it ends up all forming together.

1:50:42 But to summarize, it's like mid-training is give

1:50:44 the model the skills it needs to then learn.

1:50:47 RL with verifiable rewards is letting the model try a lot of times,

1:50:52 so put a lot of compute into trial-and-error learning across hard problems.

1:50:55 And then RLHF would be like finishing the model,

1:50:57 making it easy to use and kind of just rounding the model out.

1:51:02 Can you comment on the amount of compute required for RLVR?

1:51:06 It's only gone up and up.

1:51:08 I think Ilya Sutskever was famous for saying they

1:51:10 use a similar amount of compute for pre-training and post-training.

1:51:12 Back to the scaling discussion,

1:51:15 they involve very different hardware for scaling.

1:51:17 Pre-training is very compute-bound, which is like this FLOPs discussion,

1:51:20 which is just how many matrix multiplications can you get through at once.

1:51:24 And because with RL you're generating these answers,

1:51:26 you're trying the model in real-world environments,

1:51:29 it ends up being much more memory-bound because

1:51:31 you're generating long sequences and the attention mechanisms

1:51:35 have this behavior where you get a quadratic

1:51:38 increase in memory as you're getting to longer sequences.

1:51:42 So the compute becomes very different.

1:51:43 In pre-training we would talk about a model—if we

1:51:46 go back to like the Biden administration executive order,

1:51:48 it's like 10 to the 25th FLOPs to train a model.

1:51:51 If you're using FLOPs in post-training,

1:51:53 it's a lot weirder because the reality is just like:

1:51:56 how many hours are you allocating?

1:51:58 How many GPUs for?

1:51:59 And I think in terms of time, the RL compute is getting much closer because

1:52:04 you just can't put it all into one system.

1:52:06 Pre-training is so computationally dense where all the GPUs

1:52:09 are talking to each other and it's extremely efficient,

1:52:11 whereas RL has all these moving parts and can take

1:52:13 a long time to generate a sequence of 100,000 tokens.

1:52:17 If you think about GPT-5.2 Pro taking an hour, it's like,

1:52:20 what if your training run has a sample for an hour

1:52:23 and you have to make sure that's handled efficiently?

1:52:25 So I think in GPU hours or just wall-clock hours,

1:52:29 the RL runs are probably approaching the same number of days as pre-training,

1:52:32 but they probably aren't using as many GPUs at the same time.

1:52:36 There are rules of thumb where in labs you don't want

1:52:40 your pre-training runs to last more

1:52:41 than a month because they fail catastrophically.

1:52:43 And if you are planning a huge cluster to be

1:52:46 held for two months and then it fails on day 50,

1:52:49 the opportunity costs are just so big.

1:52:52 So people don't want to put all their eggs in one basket.

1:52:56 GPT-4 was the ultimate YOLO run, and nobody ever wanted to do it before,

1:53:01 where it took three months to train and everybody was shocked that it worked.

1:53:04 I think people are a little bit more cautious and incremental now.

1:53:07 So RLVR is more, let's say,

1:53:10 unlimited in how much you can train or still get benefit, where RLHF,

1:53:14 because it's a preference tuning,

1:53:15 you reach a certain point where it doesn't really

1:53:17 make sense to spend more RL budget on that.

1:53:20 So just a step back with preference tuning:

1:53:22 there are multiple people that can give multiple, let's say,

1:53:26 explanations for the same thing and they can both be correct,

1:53:29 but at some point you learn a certain style

1:53:31 and it doesn't make sense to iterate on it.

1:53:34 My favorite example is: if relatives ask me what laptop they should buy,

1:53:38 I give them an explanation or ask,

1:53:41 "What is your use case?" They, for example, prioritize battery life and storage.

1:53:46 Other people like us, for example, we would prioritize RAM and compute.

1:53:50 Both answers are correct, but different people require different answers.

1:53:55 With preference tuning, you're trying to average somehow.

1:53:58 You are asking the data labelers to give you, not the right,

1:54:02 but the preferred answer and then you train on that.

1:54:04 But at some point you learn that average preferred answer.

1:54:07 And there's no reason to keep training longer on it because it's just a style,

1:54:13 whereas with RLVR, you let the model

1:54:16 solve more and more complex, difficult problems.

1:54:19 So I think it makes more sense to allocate more budget long-term to RLVR.

1:54:25 Also, right now we are in an RLVR 1.0 blend where

1:54:31 it's still that simple thing where we have a question and answer,

1:54:34 but we don't do anything with the stuff in between.

1:54:38 There were multiple research papers, also by Google for example,

1:54:42 on process reward models that also give

1:54:44 scores for the explanation—how correct is the explanation.

1:54:47 And I think that will be the next thing, let's say RLVR 2.0 for this year,

1:54:53 focusing in between question and answer, like how to leverage that information,

1:54:58 the explanation, to help it get better accuracy.

1:55:01 So that's one angle.

1:55:03 And there was a DeepSeek Math-V2 paper where

1:55:07 they also had interesting inference scaling there where,

1:55:11 first, they had developed models that grade themselves, a separate model.

1:55:16 And I think that will be one aspect.

1:55:18 And the other, like Nathan mentioned,

1:55:20 it will be for RLVR branching into other domains.

1:55:24 The place where people are excited are value functions, which is pretty similar.

1:55:28 So process reward models are kind of like...

1:55:30 Process reward models assign how good something is

1:55:34 at each intermediate step in a reasoning process,

1:55:37 where value functions apply value to every token the language model generates.

1:55:41 Both of these have been largely unproven

1:55:44 in the language modeling and reasoning model era.

1:55:48 People are more optimistic about value functions for whatever reason now.

1:55:52 I think process reward models were tried a lot more in this pre-o1,

1:55:57 pre-reasoning model era, and a lot of people had a lot of headaches with them.

1:56:00 So I think a lot of it is human nature...

1:56:03 Value models have a very deep history in reinforcement learning.

1:56:06 They're one of the first things core to deep reinforcement learning existing,

1:56:10 is training value models.

1:56:12 So right now people are excited about trying value models,

1:56:16 but there's very little proof.

1:56:18 And there are negative examples in trying to scale up process reward models.

1:56:22 These things don't always hold in the future.

1:56:24 We came to this discussion by talking about scaling.

1:56:27 The simple way to summarize what you're saying

1:56:29 is you don't want to do too much RLHF, where the signal doesn't scale.

1:56:33 People have worked on RLHF for language models for years,

1:56:36 especially with intense interest after ChatGPT.

1:56:39 And the first release of a reasoning model trained with RLVR, OpenAI's o1,

1:56:44 had a scaling plot where if you increase training compute logarithmically,

1:56:47 you get a linear increase in evaluations.

1:56:50 This has been reproduced multiple times.

1:56:52 DeepSeek had a plot like this.

1:56:54 But there's no scaling law for RLHF where if you log-increase the compute,

1:56:58 you get performance.

1:56:59 In fact, the seminal scaling paper for RLHF

1:57:02 is scaling laws for reward model over-optimization.

1:57:05 So that's a big line to draw with RLVR and the methods we have now.

1:57:10 In the future, they will follow this scaling paradigm:

1:57:13 where you can let the best runs run for an extra 10x and you get performance,

1:57:18 but you can't do this with RLHF.

1:57:20 And that is just going to be field-defining in how people approach them.

1:57:24 While I'm a shill for people to academically do RLHF,

1:57:28 to do the best RLHF you might not need the extra 10 or 100x of compute,

1:57:34 but to do the best RLVR you do.

1:57:38 I think there's a seminal paper from a Meta internship.

1:57:42 It's called something like "The Art of Scaling Reinforcement Learning

1:57:46 with Language Models." What they describe as a framework is Scale-RL.

1:57:50 Their incremental experiment was like 10,000 V100 hours,

1:57:54 which is like thousands or tens of thousands of dollars per experiment.

1:57:58 They do a lot of them, and This cost is not accessible to the average academic,

1:58:04 which is a hard equilibrium where it's trying

1:58:08 to figure out how to learn from each community.

1:58:11 I was wondering if we could take a bit

1:58:13 of a tangent and talk about education and learning.

1:58:16 If you're someone listening to this who's

1:58:20 a smart person interested in programming and AI,

1:58:23 I presume building something from scratch is a good beginning.

1:58:28 So can you take me through what you would recommend people do?

1:58:32 I would personally start, like you said,

1:58:34 implementing a simple model from scratch that you can run on your computer.

1:58:38 The goal is not, when you build a model from scratch,

1:58:41 to have something for every day use.

1:58:43 It's not going to be your personal

1:58:46 assistant replacing an existing open-weight model or ChatGPT.

1:58:49 It's to see what exactly goes into the LLM, what comes out,

1:58:53 and how the pre-training works on your own computer, preferably.

1:58:59 Then you learn about pre-training,

1:59:01 supervised fine-tuning, and the attention mechanism.

1:59:03 You get a solid understanding of how things work,

1:59:06 but at some point you reach a limit, because small models can only do so much.

1:59:11 The problem with learning about LLMs at scale is

1:59:14 that it's exponentially more complex to make a larger model,

1:59:18 because the model isn't just larger—you have

1:59:21 to shard your parameters across multiple GPUs.

1:59:24 Even for the KV cache, there are multiple ways to implement it.

1:59:27 One is just to understand how it works, just to grow the cache.

1:59:31 You grow it step-by-step by, let's say, concatenating lists,

1:59:35 but then that wouldn't be optimal on GPUs.

1:59:39 You would pre-allocate a tensor and then fill it in.

1:59:42 But that adds another 20 or 30 lines of code.

1:59:45 And for each thing, you add so much code.

1:59:47 The goal with the book is basically to understand how the LLM works.

1:59:51 It's not going to be a production-level LLM,

1:59:53 but once you have that, you can understand the production-level LLM.

1:59:56 So you're trying to always build an LLM that's going to fit on one GPU?

2:00:00 Yes.

2:00:00 Most of them do.

2:00:02 I have some bonus materials on some MoE models.

2:00:04 One or two of them may require multiple GPUs,

2:00:08 but the goal is to have it on one GPU.

2:00:10 And the beautiful thing is, you can self-verify.

2:00:13 It's almost like RLVR.

2:00:14 When you code these from scratch,

2:00:17 you can take an existing model from the Hugging Face Transformers library.

2:00:21 The library is great, but if you want to learn about LLMs,

2:00:26 it's not the best place to start because the code

2:00:28 is so complex to fit so many use cases.

2:00:32 Because people use it in production,

2:00:34 it has to be really sophisticated, really intertwined, and hard to read.

2:00:38 It's not linear.

2:00:39 It started as a fine-tuning library,

2:00:40 and then it grew to be the standard representation of every model architecture.

2:00:45 Hugging Face is the default place to get a model,

2:00:48 and Transformers is the software.

2:00:49 It enables it so people can easily load a model and do something basic with it.

2:00:57 And all frontier labs that have open-weight

2:00:59 models have a Transformers version of it, like from DeepSeek to gpt-oss-120b.

2:01:03 That's the canonical weight format you can load.

2:01:06 But even even Transformers, the library, is not used in production.

2:01:10 People use SGLang or vLLM, and it adds another layer of complexity.

2:01:15 We should say that the Transformers library has like 400 models.

2:01:19 So it's the one library that tries to implement a lot of LLMs,

2:01:22 and so you have a huge codebase, basically.

2:01:25 It's huge.

2:01:26 It's like, I don't know, maybe millions—- That's crazy.

2:01:30 hundreds of thousands of lines of code.

2:01:33 Understanding the part you want to understand

2:01:35 is finding the needle in the haystack.

2:01:36 But what's beautiful is you have a working implementation,

2:01:39 so you can work backwards.

2:01:40 What I would recommend doing, or what I also do,

2:01:44 is if I want to understand, for example, how OLMo is implemented,

2:01:47 I would look at the weights in the model hub, the config file,

2:01:50 and then you can see, "Oh, they used so many layers.

2:01:53 They use, let's say,

2:01:55 Group Query Attention or Multi-Head Attention in that case."

2:01:58 Then you see all the components in a human-readable, 100-line config file.

2:02:01 And then you start, let's say, with your GPT-2 model and add these things.

2:02:05 The cool thing here is you can then load

2:02:08 the pretrained weights and see if they work in your model.

2:02:12 You want to match the same output that you get with a Transformer model,

2:02:15 and then you can use that, basically

2:02:18 as a verifiable reward to make your architecture correct.

2:02:21 Sometimes it takes me a day.

2:02:23 With OLMo 3, the challenge was RoPE for the position embeddings.

2:02:26 They had a YaRN extension and there was some custom scaling there,

2:02:32 and I couldn't quite match these things.

2:02:35 In this struggle, you kind of understand things.

2:02:38 At the end, you know you have it correct because you can unit test it.

2:02:42 You can check against the reference implementation.

2:02:44 I think that's one of the best ways to learn, really.

2:02:48 To basically reverse-engineer something.

2:02:51 I think that is something everyone interested

2:02:53 in getting into AI today should do.

2:02:56 That's why I liked your book.

2:02:58 I came to language models from the RL and robotics field.

2:03:01 I had never taken the time to just learn all the fundamentals.

2:03:06 This transformer architecture is so fundamental,

2:03:09 just as deep learning was in the past, and people need to do this.

2:03:14 I think where a lot of people get overwhelmed is,

2:03:19 "How do I apply this to have impact or find

2:03:21 a career path?" Because language models

2:03:23 make this fundamental stuff so accessible,

2:03:26 and people with motivation will learn it.

2:03:29 Then it's like, "How do I get cycles on goal to contribute to research?"

2:03:34 I'm actually fairly optimistic because the field moves so fast that a lot

2:03:38 of times the best people don't fully solve a problem because there's

2:03:42 a bigger problem to solve that's very low-hanging fruit, so they move on.

2:03:46 I think that a lot of what I was trying to do in this RLHF book is

2:03:51 take post-training techniques and describe how people think about

2:03:54 them influencing the model and what people are doing.

2:03:57 Then it's remarkable how many things I

2:04:01 just think people stop studying or don't pursue.

2:04:05 I think people trying to go narrow after doing the fundamentals is good,

2:04:08 and then reading the relevant papers and being engaged in the ecosystem.

2:04:14 It's like you actually...

2:04:16 actually...

2:04:16 The proximity that random people online have

2:04:19 to the leading researchers—no one knows who all the...

2:04:23 The anonymous accounts on X and ML are very popular,

2:04:26 and no one knows who all these people are.

2:04:28 It could just be random people that study this stuff deeply,

2:04:31 especially with the AI tools.

2:04:32 To just be like, "I don't understand this, keep

2:04:34 digging into it," is a very useful thing.

2:04:36 But there's a lot of research areas that maybe

2:04:39 have three papers that you need to read,

2:04:42 and then one of the authors will probably email you back.

2:04:45 But you have to put in a lot

2:04:47 of effort into these emails to understand the field.

2:04:50 I think it would be for a newcomer easily weeks of work

2:04:53 to feel like they can truly grasp what is a very narrow area,

2:04:57 but I think going narrow after you have the fundamentals will be

2:05:00 very useful to people because I've become very interested in character training,

2:05:05 which is how you make the model funny or sarcastic or serious,

2:05:11 and what do you do to the data to do this?

2:05:14 A student at Oxford reached out to me and was like,

2:05:16 "Hey, I'm interested in this," and I advised him.

2:05:18 And that paper now exists.

2:05:20 There's like two or three people in the world that were very interested in this.

2:05:25 He's a PhD student, which gives him an advantage, but for me,

2:05:28 that was a topic I was waiting for someone to be like,

2:05:31 "Hey, I have time to spend cycles on this." I'm sure

2:05:33 there's a lot more very narrow things where you're just like,

2:05:36 "It doesn't make sense that there was no answer to this." I

2:05:38 think it's just there's so much information coming that people are like,

2:05:42 "I can't grab onto any of these," but if you just stick in an area,

2:05:46 I think there's a lot of interesting things to learn.

2:05:48 Yeah, I think you can't try to do it all

2:05:51 because it would be very overwhelming and you would burn out.

2:05:53 For me, for example, I haven't kept up with computer vision in a long time;

2:05:57 I just focused on LLMs.

2:05:58 But coming back to your book,

2:06:00 I think this is a really great book and a really good

2:06:03 bang for the buck because if you want to learn about RLHF,

2:06:06 I wouldn't go out there and read RLHF papers because

2:06:08 you would be spending two years—- Some of them contradict.

2:06:11 I've just edited the book, and there's no chapter where I had to be like,

2:06:15 "X papers say one thing and Y papers say another,

2:06:18 and we'll see what comes out to be

2:06:21 true."- Just to go through the table of contents,

2:06:23 what are some ideas we might have missed in the bigger picture of post-training?

2:06:26 First of all, you did the problem setup,

2:06:28 training overview, what are preferences, preference data,

2:06:31 and the optimization tools, reward modeling, regularization,

2:06:35 instruction tuning, rejection sampling, and reinforcement learning.

2:06:42 Then, Constitutional AI and AI feedback, reasoning,

2:06:46 and inference-time scaling to use in function calling,

2:06:49 synthetic data and distillation,

2:06:51 evaluation, and then an open questions section, over-optimization,

2:06:54 style and information, and then product UX, character and post-training.

2:06:59 What are some ideas worth mentioning that connect

2:07:03 both the educational and the research components?

2:07:06 You mentioned character training, which is pretty interesting.

2:07:08 Character training is interesting because there's so little on it.

2:07:10 We talked about how people engage with these models.

2:07:13 We feel good using them because they're positive,

2:07:16 but that can go too far; it can be too positive.

2:07:19 And it's like, essentially, it's:

2:07:20 How do you change your data and decision-making

2:07:23 to make it exactly what you want?

2:07:26 And like, OpenAI has this thing called a model spec,

2:07:29 which is essentially their internal guideline

2:07:31 for what they want the model to do, and they publish this to developers.

2:07:35 So, essentially, you can know what is a failure

2:07:38 of OpenAI's training—where they have the intentions and they

2:07:41 haven't met them yet— versus what is something that they

2:07:43 actually wanted to do and that you don't like.

2:07:46 And that transparency is very nice,

2:07:47 but all the methods for curating these documents and how

2:07:50 easy it is to follow them is not very well known.

2:07:53 I think the way the book is designed is that the RL

2:07:56 chapter is obviously what people want

2:07:57 because everybody hears about it with RLVR,

2:07:59 and it's the same algorithms and the same math,

2:08:01 but you can use it in very different documents.

2:08:05 So I think the core of RLHF is like how messy preferences are.

2:08:09 It's essentially a rehash of a paper I wrote years ago,

2:08:13 but this is essentially the chapter that'll tell

2:08:15 you why RLHF is never ever fully solvable because,

2:08:21 the way that even RL is set up, it assumes that preferences can be quantified

2:08:29 and that multiple preferences can be reduced to single values.

2:08:33 And I think it relates in the economics

2:08:35 literature to the Von Neumann-Morgenstern utility theorem,

2:08:38 and that is the chapter where all of that philosophical, economic,

2:08:43 and psychological context tells you what gets compressed into doing RLHF.

2:08:47 So it's like you have all of this and then later in the book it's like:

2:08:50 You use this RL map to make the number go up.

2:08:52 And I think that's why it'll be very

2:08:54 rewarding for people to do research on, because

2:08:57 quantifying preferences is something that humans have designed

2:09:01 a problem in order to make preferences studyable.

2:09:04 But there's kind of fundamental debates, like, an example is in a language model

2:09:09 response you have different things you care about, like accuracy or style.

2:09:12 And when you're collecting the data, they all get compressed into: "I like

2:09:16 this more than another." And that is happening,

2:09:19 and there's a lot of research in other areas

2:09:22 of the world that go into how you should actually do this.

2:09:26 I think social choice theory is the subfield

2:09:30 of economics around how you should aggregate preferences.

2:09:33 And I went to a workshop that published a white paper

2:09:37 on: "How can you think about using social choice theory for RLHF?" So

2:09:41 I mostly would want people that get excited about the math

2:09:44 to come and find things where they could stumble into this broader context.

2:09:48 I think there's a fun thing:

2:09:49 I just keep a list of all the tech reports of reasoning models I like.

2:09:54 So in Chapter 14, where there's a short summary of RLVR,

2:09:57 there's just a gigantic table where I

2:09:59 list every single reasoning model that I like.

2:10:03 I think in education, a lot of it needs to be like, at this point,

2:10:07 what I like, because the language models are so good at the math.

2:10:11 For example, the famous paper, Direct Preference Optimization,

2:10:13 which is a much simpler way of solving the problem than RL.

2:10:17 The derivations in the appendix skip steps of math.

2:10:21 And for this book, I redid the derivations and I'm like,

2:10:24 "What the heck is this log trick that they use

2:10:26 to change the math?" But doing it with language models, they're like,

2:10:29 "This is the log trick." And I'm like,

2:10:31 "I don't know if I like this, that the math is so commoditized." I think

2:10:35 some of the struggle in reading this appendix-

2:10:38 ...and following the math is good for learning.

2:10:44 Yeah, we're returning to this often on the topic of education.

2:10:47 You both have brought up the word "struggle" quite a bit.

2:10:51 So there is value.

2:10:52 If you're not struggling as part of this process,

2:10:55 you're not fully following the proper process for learning.

2:10:59 proper process for learning, I suppose.

2:11:02 Some providers are working on models

2:11:04 for education designed to not give- actually, I haven't used them,

2:11:07 but I'd guess they're designed to not give all the information at once.

2:11:12 And make people work for it.

2:11:13 Training models to do this would be a wonderful contribution.

2:11:16 Where, like all of the stuff in the book, you had to reevaluate every decision.

2:11:19 decision for it- It's a great example.

2:11:21 There's a chance we work on it at Ai2, which I thought would be so fun.

2:11:26 It makes sense.

2:11:27 I did something like that the other day for video games.

2:11:30 Sometimes for pastime I play video games,

2:11:32 like I like- Video games with puzzles, like Zelda and Metroid.

2:11:36 And there's this new game where I really got stuck and was okay with it.

2:11:41 I don't want to struggle for two days, so I used an LLM.

2:11:45 But then you say, "Hey, please don't add spoilers.

2:11:47 Just, you know, I'm here and there.

2:11:49 What do I have to do next?" You can do the same thing for math where you say,

2:11:53 "Okay, I'm stuck at this point.

2:11:55 Don't give me the full solution,

2:11:57 but what is something I could try?" Where you carefully probe it.

2:12:01 But the problem here is I think it requires discipline.

2:12:05 Many people enjoy math,

2:12:06 but there are also a lot of people who need to do it for their homework,

2:12:11 and then it's like a shortcut.

2:12:13 We could develop an educational LLM, but other LLMs are still there,

2:12:17 and there's still a temptation to use the other LLMs.

2:12:20 I think many people in college understand the stuff they're passionate

2:12:23 about- about- ...they're self-aware and they understand it shouldn't be easy.

2:12:27 I think we just have to develop a good taste- ...talk about research taste,

2:12:33 school taste about stuff that you should

2:12:36 be struggling on- ...and stuff you shouldn't be.

2:12:38 It's tricky, because you don't have

2:12:40 good long-term vision sometimes you don't have

2:12:43 good long-term vision about what would be actually useful to you in your career.

2:12:48 But you have to develop that taste, yeah.

2:12:52 I was talking to my fiancee or friends about this, there's this brief

2:12:56 10-year window where all of the homework and all the exams could be digital.

2:13:00 Before that, everybody had to do all the exams

2:13:02 in blue books because there was no other way.

2:13:04 And now after AI, everyone's going to need to be

2:13:06 in blue books and oral exams because everyone could cheat so easily.

2:13:09 It's like this brief generation that had

2:13:11 a different education system where everything could be digital,

2:13:15 but you still couldn't cheat.

2:13:16 And now it's just going back.

2:13:18 It's just very funny.

2:13:21 You mention character training.

2:13:22 Just zooming out on a more general topic,

2:13:24 for that project how much compute was required?

2:13:28 And in general, to contribute as a researcher,

2:13:31 are there places where not too much compute is

2:13:35 required where you can actually contribute as an individual researcher?

2:13:39 For the character training thing, I think this research is built on fine-tuning

2:13:43 about 7 billion parameter models with LoRA,

2:13:46 which is essentially only fine-tuning a small

2:13:48 subset of the weights of the model.

2:13:51 I don't know exactly how many GPU hours that would take.

2:13:55 But it's doable.

2:13:56 Not doable for every academic.

2:13:57 The situation for some academics is so dire that the only

2:14:00 work you can do is doing inference where you have

2:14:02 closed models or open models and you get completions from them

2:14:05 and you can look at them and understand the models.

2:14:07 And that's very well-suited to evaluation,

2:14:09 where you want to be the best at creating representative

2:14:14 problems that the models fail on or show certain abilities,

2:14:17 which I think that you can break through with this.

2:14:21 I think that the top-end goal for a researcher working on evaluation,

2:14:25 if you want to have career momentum,

2:14:27 is that Frontier Labs pick up your evaluation.

2:14:30 You don't need to have every project do this.

2:14:32 But if you go from a small university with no compute and find something

2:14:36 that Claude struggles with, and then the next

2:14:39 Claude model has it in the blog post, there's your career rocket ship.

2:14:42 I think that's hard,

2:14:44 but if you want to scope the maximum possible impact with minimum compute,

2:14:48 it's something like that, which is just get very narrow

2:14:51 and it takes learning of where the models are going.

2:14:54 So you need to build a tool that tests where Claude 4.5 will fail.

2:14:59 If I'm going to start a research project,

2:15:02 I need to think where the models in eight months are going to be struggling.

2:15:06 But what about developing totally novel ideas?

2:15:08 This is a trade-off.

2:15:09 I think that if you're doing a PhD,

2:15:11 you could also be like, "It's too risky to work in language models.

2:15:15 I'm going way longer term," which is like what is— what is

2:15:19 the thing that's going to define language model development in 10 years?

2:15:22 I end up being a person that's pretty practical.

2:15:25 I mean, I went to my PhD where it was like, "I got into Berkeley.

2:15:28 Worst case, I get a master's,

2:15:29 and then I go work in tech." I'm very practical about it,

2:15:32 so I'm like the life afforded to people

2:15:36 to work at these AI companies, the amount of...

2:15:38 OpenAI's average compensation is over a million

2:15:40 dollars in stock a year per employee.

2:15:43 For any normal person in the US,

2:15:46 to get into this AI lab is transformative for your life.

2:15:49 So I'm pretty practical about it.

2:15:50 there's still a lot of upward mobility

2:15:52 working in language models if you're focused.

2:15:54 And look at these jobs.

2:15:55 But from a research perspective,

2:15:57 the transformative impact in these academic awards...

2:16:01 to be the next Yann LeCun is from not

2:16:04 working on— not caring about language model development very much.

2:16:07 It's a big financial sacrifice in that case.

2:16:09 So I work with some awesome students, and they're like,

2:16:12 "Should I go work at an AI lab?" And I'm like,

2:16:14 "You're getting a PhD at a top school.

2:16:16 Are you gonna leave to go to a lab?" I don't know.

2:16:19 If you go work at a top lab, I don't blame you.

2:16:22 Don't go work at some random startup that might go to zero.

2:16:24 But if you're going to OpenAI, I'm like, "It could be worth leaving a PhD

2:16:29 for."- Let's more rigorously think through this.

2:16:32 So where would you give a recommendation

2:16:34 for people to do a research contribution?

2:16:36 So the options are academia: get a PhD.

2:16:40 Spend five years publishing.

2:16:44 Compute resources are constrained.

2:16:46 There's— there's research labs that are more

2:16:50 focused on open-weight models, and working there.

2:16:57 Or closed frontier research labs.

2:17:01 So OpenAI, Anthropic, xAI, and so on.

2:17:04 The two gradients are: the more closed,

2:17:06 the more money you tend to get, but you also get less credit.

2:17:10 In terms of building a portfolio of things that you've done,

2:17:17 it's very clear what you have done as an academic.

2:17:20 Versus if you are going to trade this fairly

2:17:25 reasonable progression for being a cog in the machine,

2:17:28 which could also be very fun.

2:17:30 So I think it's very different career paths.

2:17:33 But the opportunity cost for being a researcher is

2:17:36 very high because PhD students are paid essentially nothing.

2:17:38 So it ends up rewarding people that have a fairly stable safety net,

2:17:42 and they realize that they can operate in the long term,

2:17:45 wanting to do very interesting work and get a very interesting job.

2:17:49 So it is a privileged position to be like,

2:17:53 "I'm gonna see out my PhD and figure it out

2:17:56 after because I want to do this." At the same time,

2:18:00 the academic ecosystem is getting bombarded by funding getting cut and stuff.

2:18:04 So there's just so many different trade-offs where I understand

2:18:06 plenty of people that are like, "I don't enjoy it.

2:18:08 I can't deal with this funding search.

2:18:10 My grant got cut for no reason by the government," or, "I don't know what's

2:18:15 gonna happen." So I think there's a lot

2:18:17 of uncertainty and trade-offs that, in my opinion,

2:18:20 favor just taking the well-paying job with meaningful impact.

2:18:24 It's not like you're getting paid to sit around at OpenAI.

2:18:27 You're building the cutting edge of things

2:18:29 that are— changing millions of people's relationship to tech.

2:18:35 But publication-wise, they're being more secretive, increasingly so.

2:18:38 So you're publishing less and less.

2:18:40 And so you are having a positive impact at scale,

2:18:44 but you're a cog in the machine.

2:18:48 I think it honestly hasn't changed that much.

2:18:51 I have been in academia.

2:18:53 I'm not in academia anymore.

2:18:55 wouldn't want to miss my time in academia.

2:18:57 But what I wanted to say before I get

2:18:59 to that is that I think it hasn't changed that much.

2:19:02 I was working in computational biology,

2:19:04 using AI or machine learning methods with collaborators,

2:19:08 and a lot of people went from academia directly to Google.

2:19:15 And I think it's the same.

2:19:17 Back then, professors were sad that their students went

2:19:21 into industry because they couldn't carry on their legacy.

2:19:25 I think it's the same.

2:19:26 It hasn't changed that much.

2:19:28 The only thing that has changed is the scale.

2:19:32 Cool stuff was always developed in industry that was closed.

2:19:36 You couldn't talk about it.

2:19:38 And I think the difference now is your preference.

2:19:42 Do you like to publish your work, or are you more in a closed lab?

2:19:47 That's one difference.

2:19:48 The compensation, of course, is another, but it's always been like that.

2:19:54 It depends on where you feel comfortable.

2:19:56 And nothing is forever.

2:19:58 Right now, there's a third option, which is launching a startup.

2:20:02 A lot of people are doing that.

2:20:05 It's a very risky move, but it can be a high-risk,

2:20:10 high-reward situation, whereas joining an industry lab is pretty safe.

2:20:14 You also have upward mobility.

2:20:17 I think once you've been at an industry lab, it's easier to find future jobs.

2:20:22 But then again, how much do you enjoy the team

2:20:28 and working on proprietary things versus how much you like publishing work?

2:20:33 I mean, publishing is stressful.

2:20:35 Acceptance rates at conferences can be arbitrary and very frustrating,

2:20:40 but it's high reward if you have a paper published.

2:20:43 You feel good because your name is on there.

2:20:46 It's a high accomplishment.

2:20:48 I feel like my friends who are professors seem happier

2:20:51 than those who work at a frontier lab, to be honest.

2:20:55 There's a grounding there.

2:20:57 The frontier labs definitely do this 9-9-6,

2:21:00 which is shorthand for working all the time.

2:21:03 Can you describe 9-9-6?

2:21:05 It's a culture invented, I believe, in China and adopted in Silicon Valley.

2:21:10 What is 9-9-6?

2:21:11 It's 9:00 AM to 9:00 PM,- Six days a week.

2:21:15 six days a week.

2:21:16 What is that, 72 hours?

2:21:18 Okay.

2:21:18 So, is this basically the standard in AI companies in Silicon Valley?

2:21:24 This kind of grind mindset.

2:21:27 Yeah, I mean, maybe not exactly like

2:21:28 that, but I think there is a trend towards it.

2:21:30 And it's interesting.

2:21:31 I think it almost flipped because when I was in in academia, I felt like that.

2:21:36 As a professor, you write grants, you teach, and you do research.

2:21:39 It's like three jobs in one,

2:21:41 and it's more than a full-time job if you want to be successful.

2:21:45 successful.

2:21:46 And I feel like now, like Nathan just said,

2:21:49 the professors, in comparison to a lab,

2:21:52 I think they have less pressure or workload than

2:21:55 at a frontier lab because—- I think they work a lot.

2:21:58 They're just so fulfilled.

2:21:59 By working with students— and having a constant runway

2:22:02 of mentorship and a mission that is very people-oriented,

2:22:05 I think in a era when things are moving very fast and are very chaotic,

2:22:09 it's very rewarding to people.

2:22:11 Yeah, and I think at a startup, it's this pressure.

2:22:14 It's like you have to make it.

2:22:16 And it's really important that people put in the time,

2:22:19 but it is really hard because you have to deliver constantly,

2:22:22 and I've been at a startup.

2:22:24 I had a good time, but I don't know if I could do it forever.

2:22:28 It's an interesting pace and it's exactly like we talked about in the beginning.

2:22:33 These models are leapfrogging each other,

2:22:35 and they are just constantly trying to take

2:22:38 the next step compared to their competitors.

2:22:40 It's just ruthless right now.

2:22:42 I think this leapfrogging nature and having

2:22:44 multiple players is actually an underrated

2:22:46 driver of language modeling progress where

2:22:48 competition is so deeply ingrained in people,

2:22:52 and these companies have intentionally created very strong cultures.

2:22:56 Like, Anthropic is known to be so culturally,

2:23:00 like, deeply committed and organized.

2:23:02 I mean, we hear so little from them,

2:23:05 and everybody at Anthropic seems very aligned.

2:23:07 And it's like being in a culture that is super tight and having this competitive

2:23:13 dynamic is a thing that's gonna make you

2:23:16 work hard and create things that are better.

2:23:20 But that comes at the cost of human capital,

2:23:22 which is like you can only do this for so long,

2:23:26 and people are definitely burning out.

2:23:28 I wrote a post on burnout as I've tread in and out of this myself,

2:23:33 especially trying to be a manager, full-mode training.

2:23:36 It's a crazy job doing this.

2:23:37 The book Apple in China by Patrick McGee,

2:23:40 he talked about how hard the Apple engineers

2:23:42 worked to set up the supply chains in China,

2:23:44 and he was like, they had "saving marriage" programs, and he told in a podcast,

2:23:49 he was like, "People died from this level of working hard." So I think

2:23:53 it's just like it's a perfect environment

2:23:56 for creating progress based on human expense, and there's gonna be a lot of...

2:24:02 the human expense is the 996 that we started this with, which is like—...

2:24:07 people do really grind.

2:24:08 I also read this book.

2:24:09 I think they had a code word for if someone had to go

2:24:12 home to spend time with their family to save the marriage, and it's crazy.

2:24:15 Then the colleagues said, "Okay, this is like red alert for this situation.

2:24:19 We have to let that person go home this weekend." But at the same time,

2:24:23 I don't think they were forced to work.

2:24:25 They were so passionate about the product,

2:24:27 I guess, that you get into that mindset.

2:24:29 And I had that sometimes as an academic,

2:24:32 but also as an independent person, I have that sometimes.

2:24:35 I overwork, and it's unhealthy.

2:24:37 I had back issues, I had neck issues,

2:24:39 because I did not take the breaks that I maybe should have taken.

2:24:43 But no one forced me to; it's because I wanted to work,

2:24:45 because it's exciting stuff.

2:24:46 That's what OpenAI and Anthropic are like.

2:24:47 They want to do this work.

2:24:49 Yeah, but there's also a feeling of fervor that's building,

2:24:53 especially in Silicon Valley, aligned with the scaling laws idea,

2:24:56 where there's this hype where the world will be transformed in a scale

2:25:00 of weeks and you want to be at the center of it.

2:25:03 And then, you know, I have this great fortune

2:25:07 of having conversations with a wide variety of human beings,

2:25:12 and from there I get to see all

2:25:14 these bubbles and echo chambers across the world.

2:25:17 It's fascinating to see how we humans form them.

2:25:19 And I think it's fair to say that Silicon Valley is a kind of echo chamber,

2:25:25 a kind of silo and bubble.

2:25:27 I think bubbles are actually really useful and effective.

2:25:31 It's not necessarily a negative thing because you could be ultra-productive.

2:25:34 It could be the Steve Jobs reality distortion field,

2:25:39 because you just convince each other that breakthroughs are imminent,

2:25:42 and by convincing each other of that, you make the breakthroughs imminent.

2:25:49 Byrne Hobart wrote a book classifying bubbles.

2:25:51 One of them is financial bubbles, which is like speculation, which is bad,

2:25:54 and the other one is for build-outs,

2:25:56 because it pushes people to build these things.

2:25:58 And I do think AI is in this, but I

2:26:01 worry about it transitioning to a financial bubble, which is- Yeah,

2:26:05 but also in the space of ideas,

2:26:07 that bubble—you are doing a reality distortion field,

2:26:12 and that means you are deviating from reality.

2:26:14 And if you go too far from reality while also working, you know, 996,

2:26:22 you might miss some fundamental aspects of the human experience,

2:26:26 including beyond Silicon Valley.

2:26:27 This is a common problem in Silicon Valley:

2:26:30 it's a very specific geographic area.

2:26:32 You might not understand the Midwest perspective,

2:26:34 the full experience of all the other humans

2:26:38 in the United States and across the world,

2:26:40 and you speak a certain way to each other,

2:26:42 you convince each other of a certain thing,

2:26:44 and that can get you into real trouble.

2:26:47 Whether AI is a big success and becomes a powerful technology or it's not,

2:26:53 in either trajectory you can get yourself into trouble.

2:26:56 So you have to consider all of that.

2:26:58 Here you are, a young person trying to decide

2:27:00 what you want to do with your life.

2:27:02 The thing that is...

2:27:03 I don't even really understand this, but the SF AI memes

2:27:07 have gotten to the point where "permanent underclass" was one of them,

2:27:11 which was the idea that the last six months of 2025 was

2:27:14 the only time to build durable value in an AI startup or model.

2:27:18 Otherwise, all the value will be captured by existing

2:27:21 companies and you will therefore be poor, which...

2:27:24 that's an example of the SF thing that goes so far.

2:27:28 I still think for young people going to be able to tap into it,

2:27:31 if you're really passionate about wanting to have an impact in AI,

2:27:35 being physically in SF is the most likely place where you're going to do this.

2:27:39 But it has has trade-offs.

2:27:42 I think SF is an incredible place, but there is a bit of a bubble.

2:27:46 And if you go into that bubble, which is extremely valuable, just get out also.

2:27:52 Read history books, read literature, visit other places in the world.

2:27:57 Twitter and Substack are not the entire world.

2:28:01 I would say, one of the people I worked with is moving to SF,

2:28:04 and it's like, I need to get him a copy of Season of the Witch,

2:28:07 which is a history of SF from 1960 to 1985,

2:28:10 which goes through the hippie revolution,

2:28:14 like all the gays taking over the city and that culture emerging,

2:28:19 and then the HIV/AIDS crisis and other things.

2:28:22 And it's just like, that is so recent,

2:28:24 and so much turmoil and hurt, but also love in SF.

2:28:28 And it's like, no one knows about this.

2:28:30 It's a great book, Season of the Witch.

2:28:31 I recommend it.

2:28:32 A bunch of my SF friends who get out recommended it to me.

2:28:37 And I think that's just like living there...

2:28:39 I lived there and I didn't appreciate this context, and it's just so recent.

2:28:46 Yeah.

2:28:47 Okay, let's...

2:28:48 We talked a lot about a lot of things.

2:28:52 Certainly about the things that were exciting last year.

2:28:56 But this year, One of the things you guys

2:28:59 mentioned that's exciting is the scaling of text diffusion models,

2:29:02 and just a different exploration of text diffusion.

2:29:04 Can you talk about what that is and what the possibility it holds?

2:29:09 So, different kinds of approaches than the current LMs?

2:29:13 Yeah, so we talked a lot about the transformer

2:29:16 architecture and the autoregressive

2:29:17 transformer architecture specifically, like GPT.

2:29:19 And it doesn't mean no one else is working on anything else.

2:29:23 So, people are always on the, let's say, lookout for the next big thing.

2:29:27 Because I think it would be almost stupid not to.

2:29:30 Because sure, right now, the transformer architecture is the thing,

2:29:33 and it works best, and there's, right now, nothing else out there.

2:29:37 But, you know, it's always a good idea to not put all your eggs into one basket.

2:29:41 So, people are developing other alternatives to the autoregressive transformer.

2:29:45 One of them would be, for example, text diffusion models.

2:29:49 And listeners may know diffusion models from image generation,

2:29:52 like Stable Diffusion popularized it.

2:29:54 There was a paper on generating images.

2:29:57 Back then, people used GANs, Generative Adversarial Networks.

2:30:00 And then there was this diffusion

2:30:02 process where you iteratively denoise an image,

2:30:04 and that resulted in really good quality images over time.

2:30:08 Stable Diffusion was a company.

2:30:09 Other companies build their own diffusion models.

2:30:11 And then people are now like, "Okay,

2:30:13 can we try this also for text?" Doesn't, you know,

2:30:16 make intuitive sense yet, because it feels like, okay,

2:30:18 it's not something continuous like a pixel that we can differentiate.

2:30:21 It's discrete text, so how do we implement that denoising process?

2:30:26 It's kind of similar to the BERT models by Google.

2:30:31 Like, when you go back to the original transformer,

2:30:33 they were the encoder and the decoder.

2:30:35 The decoder is what we are using right now in GPT and so forth.

2:30:39 The encoder is more like a parallel technique where

2:30:43 you have multiple tokens that you fill in in parallel.

2:30:47 GPT models, they do autoregressive generation,

2:30:49 completing the sentence one token at a time.

2:30:52 And in BERT models, you have a sentence that has gaps.

2:30:57 You mask them out, and then one iteration is filling in these gaps.

2:31:02 Text diffusion is kind of like that, where

2:31:04 you are starting with some random text, and then you are filling in the missing

2:31:10 parts or refining them iteratively over multiple iterations.

2:31:12 And the cool thing here is that this can do multiple tokens at the same time.

2:31:18 It's like the promise of having it more efficient.

2:31:21 Now, the trade-off is, of course, how good is the quality?

2:31:25 It might be faster, and now you have this dimension of the denoising process.

2:31:29 The more steps you do, the better the text becomes.

2:31:32 And people...

2:31:34 I mean, you can scale in different ways.

2:31:37 They try to see if that is maybe a valid alternative to the autoregressive

2:31:41 model in terms of giving you the same quality for less compute.

2:31:46 Right now, there are papers that suggest if you want to get the same quality,

2:31:51 you have to crank up the denoising steps, and then you end up spending the same

2:31:56 compute you would spend on an autoregressive model.

2:31:58 The other downside is, while being parallel sounds appealing,

2:32:01 some tasks are not parallel.

2:32:03 Like reasoning tasks or tool use, maybe where you have to ask a code

2:32:08 interpreter to give you an intermediate result.

2:32:10 That is tricky with diffusion models.

2:32:12 So, there are some hybrids, but the main idea is: how can we parallelize it?

2:32:16 It's an interesting avenue.

2:32:18 I think right now, there are mostly research models out there,

2:32:22 like LaMDA and some other ones.

2:32:24 I saw some by startups, some deployed models.

2:32:27 There is no big diffusion model at scale yet,

2:32:30 like on the Gemini or ChatGPT level.

2:32:32 But there was an announcement by Google,

2:32:36 a site where they said they are launching Gemini Diffusion,

2:32:39 and they put it into context of their Gemini Nano 2 model,

2:32:44 and they said basically:

2:32:45 for the same quality on most benchmarks, we can generate things much faster.

2:32:50 You mentioned what's next.

2:32:52 I don't think the text diffusion model is going to replace autoregressive LLMs,

2:32:55 but it will be something maybe for quick, cheap, at-scale tasks.

2:33:00 Maybe the free tier in the future will be something like that.

2:33:08 I think there are examples where it's already being used.

2:33:10 To paint an example of why this is better, for example,

2:33:13 when GPT-5 is taking 30 minutes to respond, it's generating one token at a time.

2:33:17 And this diffusion idea is essentially to generate all

2:33:20 of those tokens and the completion in one batch,

2:33:23 which is why it could be way faster.

2:33:25 And I think it could be suited for...

2:33:27 the startups I'm hearing about are code startups where you have a code base,

2:33:31 and you have somebody that's effectively "vibe coding," and they say,

2:33:34 "Make this change." And a code diff is essentially a huge reply from the model,

2:33:39 but it doesn't have to have that much external context,

2:33:42 and you can get it really fast by using these diffusion models.

2:33:45 One example I've heard is that they

2:33:47 use text diffusion to generate really long diffs,

2:33:50 because doing it with an autoregressive model would take minutes,

2:33:53 and that time for a user-facing product causes a lot of churn.

2:33:57 Every second, you lose a lot of users.

2:33:59 So, I think it's going to be this thing

2:34:01 where it's going to— ...grow and have some applications,

2:34:03 but I actually thought that different types of models were going

2:34:06 to be used for different things much sooner than they have been,

2:34:10 so I kind of trade off.

2:34:11 I think the tool-use point is the one

2:34:13 that's stopping them from being most general purpose because,

2:34:18 for Claude Code and ChatGPT search,

2:34:22 the autoregressive chain is interrupted with some external tool,

2:34:25 and I don't know how to do that with the diffusion setup.

2:34:29 So what's the future of tool use this year and then in the coming years?

2:34:32 Do you think there's going to be a lot of developments there,

2:34:35 and how that's integrated into the entire stack?

2:34:37 I do think right now, it's mostly on the proprietary LLM side,

2:34:41 but I think we will see more of that in the open-source tooling.

2:34:44 And I think it is a huge unlock because then you

2:34:48 can really outsource certain tasks from just memorization to actual— you know,

2:34:54 instead of having the LLM memorize what is 23 plus 5, just use a calculator.

2:34:59 So do you think that can help solve hallucination?

2:35:02 Not solve it, but reduce it.

2:35:03 So the LLM still needs to know when to ask for a tool call.

2:35:09 And the second one is, well, it doesn't mean the internet is always correct.

2:35:13 You can do a web search,

2:35:14 but let's say I asked who won the World Cup in, let's say,

2:35:18 1998; it still needs to find the right website and get the right information.

2:35:21 You can still go to the incorrect website and give me incorrect information.

2:35:25 So I don't think it will fully solve that, but it is improving it in that sense.

2:35:31 And so another cool paper earlier this year—I think it was December 31st,

2:35:36 so it's not technically 2026, but close—the recursive language model.

2:35:43 That's a cool idea to kind of take this even a bit further.

2:35:47 Just to explain, Nathan, you also mentioned earlier,

2:35:51 it's harder to do cool research in academia because of the compute budget.

2:35:54 If I recall correctly, they did everything with GPT-5,

2:35:57 so they didn't even use local models,

2:35:59 but the idea is, let's say you have a long-context task;

2:36:01 instead of having the LLM solve all of it in one shot or even in a chain,

2:36:06 you break it down into sub-tasks.

2:36:08 You have the LLM decide what is a good sub-task,

2:36:12 and then recursively call an LLM to solve that.

2:36:16 And I think something like that, adding tools—you know,

2:36:20 each one maybe you have a huge Q&A task,

2:36:23 so each one goes to the web and gathers information,

2:36:26 and then you pull it together at the end and stitch it back together.

2:36:29 I think there's going to be a lot of unlock using

2:36:33 things like that where you don't necessarily improve the LLM itself;

2:36:37 you improve how the LLM is used and what the LLM can use.

2:36:41 One downside right now with tool use is you

2:36:43 have to give the LLM permission to use tools.

2:36:46 And that will take some trust,

2:36:49 especially if you want to unlock things like having

2:36:51 an LLM answer emails for you—or not even answer,

2:36:54 but just sort them for you or select them for you or something like that.

2:36:57 I don't know if I would today give an LLM access to my emails, right?

2:37:01 I mean, this is a huge risk.

2:37:03 I think there's a cool...

2:37:04 one last point on the tool use thing.

2:37:06 I think that you hinted at this, and we've both come at this in our own ways,

2:37:10 is that the open versus closed models use

2:37:12 tools in very different ways, where open models,

2:37:15 people go to Hugging Face and download the model,

2:37:17 and then the person's going to be like,

2:37:18 "What tool do I want?" I don't know, Exa is my preferred search provider,

2:37:22 but somebody else might care for a different search startup.

2:37:25 Where you release a model,

2:37:26 it needs to be useful for multiple tools, for multiple use cases,

2:37:29 which is really hard because you're making a general reasoning engine model,

2:37:33 which is actually what gpt-oss-120b is good for.

2:37:36 But on the closed models,

2:37:38 you're deeply integrating the specific tool into your experience,

2:37:41 and I think that open models will struggle to replicate some

2:37:45 of the things that I like to do with closed models,

2:37:47 which will be like, you can reference a mix of public and private information.

2:37:51 And something that I keep trying every three to six months,

2:37:55 I try Claude Code on the web, which is just prompting a model to make

2:37:59 an update to some GitHub repository that I have.

2:38:02 And it's just like that set of secure cloud environments is just so nice

2:38:06 for just sending it off to do this thing and then come back to me,

2:38:10 and these will probably help define some of the local open and closed niches.

2:38:18 But I think initially, because there was such a rush to get tool use working,

2:38:22 the open models were on the back foot, which is kind of inevitable.

2:38:25 I think there's so much research, so many resources in these frontier labs,

2:38:28 but it will be fun when the open models solve

2:38:31 this because it's going to necessitate a bit more flexible

2:38:34 and potentially interesting model that might work with this recursive

2:38:37 idea to be an orchestrator and a tool use model,

2:38:41 so hopefully the necessity drives some interesting innovation there.

2:38:45 So, continual learning—this is a longstanding topic, important problem.

2:38:51 I think that increases in importance as the cost of training the models goes up.

2:38:56 So can you explain what continual learning is and how important it

2:38:59 might be this year and in the coming years to make progress?

2:39:03 This relates a lot to this kind of SF zeitgeist of, what is AGI,

2:39:07 which is Artificial General Intelligence,

2:39:08 and what is ASI, Artificial Superintelligence,

2:39:11 and what are the language models that we have today capable of doing?

2:39:15 I think the language models can solve a lot of tasks,

2:39:18 but a key milestone among the AI community

2:39:21 is essentially when AI could replace any remote worker,

2:39:25 taking in information and solving digital tasks and doing them.

2:39:29 And the limitation that's highlighted by people is that a language model

2:39:33 will not learn from feedback the same way that an employee does.

2:39:36 So if you hire an editor, the editor will mess up, but you will tell them.

2:39:41 And if you hired a good editor, they don't do it again.

2:39:43 But language models don't have this ability

2:39:45 to modify themselves and learn very quickly.

2:39:47 So the idea is, if we are going to actually get to something that is a true,

2:39:52 general adaptable intelligence that can go into any remote work scenario,

2:39:55 it needs to be able to learn quickly from feedback and on-the-job learning.

2:40:00 I'm personally more bullish on language models being

2:40:02 able to just provide them with very good context.

2:40:06 You said, maybe offline,

2:40:07 that you can write extensive documents to models where you say,

2:40:11 "I have all this information.

2:40:13 Here are all the blog posts I've ever written.

2:40:15 I like this type of writing.

2:40:17 My voice is based on this." But many people don't provide this to models,

2:40:20 and the models weren't designed to take this amount of context previously.

2:40:24 Agentic models are just starting.

2:40:26 So it's this kind of trade-off: do we need to update the weights of this model

2:40:31 with this continual learning thing to make them learn fast?

2:40:34 Or the counterargument is we just need

2:40:36 to provide them with more context and information,

2:40:38 and they will have the appearance of learning fast

2:40:40 by having a lot of context and being smart.

2:40:43 So we should mention the terminology here.

2:40:45 Continual learning refers to changing the weights continuously so

2:40:49 that the model adapts and adjusts based on the new incoming information,

2:40:57 doing so continually, rapidly, and frequently.

2:41:00 And then the thing you mentioned on the other side

2:41:03 of it generally will be referred to as in-context learning.

2:41:07 As you learn stuff, there's a huge context window.

2:41:11 You can just keep loading it with extra

2:41:13 information every time you prompt the system,

2:41:15 which I think both legitimately can be seen as learning.

2:41:21 It's just a different place where you're doing the learning.

2:41:24 I think, to be honest with you,

2:41:26 continual learning— updating weights— we already have that in different flavors.

2:41:30 If you think about how...

2:41:32 I think the distinction here is:

2:41:34 do you do that on a personalized custom model for each person,

2:41:39 or do you do it on a global model scale?

2:41:42 I think we have that already, going from GPT-5 to 5.1 and 5.2.

2:41:47 It's maybe not immediate, but it is a curated update,

2:41:50 a quick curated update where there was feedback about things they couldn't do,

2:41:54 feedback by the community.

2:41:55 They updated the weights, next model, and so forth.

2:41:58 So it is a flavor of that.

2:42:02 Another even finer-grained example is like RLVR; you run it, it updates.

2:42:08 The problem is you can't just do that for each person because

2:42:12 it would be too expensive to update the weights for each person,

2:42:15 and I think that's the problem.

2:42:16 Unless you get...

2:42:17 Even at OpenAI scale, building the data centers, it would be too expensive.

2:42:22 I think that is only feasible once you have something

2:42:25 on the device where the cost is on the consumer.

2:42:27 Like what Apple tried to do with the Apple Foundation models,

2:42:30 putting them on the phone, where they learn from experience.

2:42:34 A bit of a related topic,

2:42:37 but this kind of, maybe anthropomorphized term: memory.

2:42:42 What are different ideas for the mechanism of how

2:42:44 to add memory to these systems as we're increasingly seeing?

2:42:47 Personalized memory especially?

2:42:50 Right now, it's mostly basically stuffing things

2:42:53 into the context and then just recalling that.

2:42:56 But again, I think it's expensive because you have to—you can cache it,

2:43:03 but still you spend tokens on that.

2:43:06 And the second one is you can only do so much.

2:43:09 I think it's more like a preference or style.

2:43:11 I mean, a lot of people do that when they solve math problems.

2:43:14 You say it's way so you can add previous knowledge and stuff,

2:43:17 but you also give it certain preference prompts:

2:43:20 "do what I preferred last time," or something like that.

2:43:23 But it doesn't unlock new capabilities.

2:43:26 So for that, one thing people still use is LoRA adapters.

2:43:31 These are basically, instead of updating the whole weight matrix,

2:43:35 there are two smaller weight matrices that you kind

2:43:38 of have in parallel or overlays like the delta.

2:43:41 But yeah, you can do that to some extent, but then again, it is economics.

2:43:47 There were also papers, for example, LoRA learns less but forgets less.

2:43:53 It's like, there's no free lunch.

2:43:54 If you want to learn more,

2:43:56 you need to use more weights, but it gets more expensive.

2:43:58 And then again, if you learn more, you forget more,

2:44:01 and you have to find that Goldilocks zone basically.

2:44:05 We haven't really mentioned it much,

2:44:06 but implied in this discussion is context length also.

2:44:09 Is there a lot of innovation that's possible there?

2:44:13 I think the colloquially accepted thing is that it's

2:44:16 a compute and data problem where you can...

2:44:19 and sometimes small architecture things like attention variants.

2:44:23 We talked about hybrid attention models,

2:44:26 which is essentially if you have what looks

2:44:29 like a state space model within your transformer.

2:44:31 And those are better suited because you have

2:44:34 to spend less compute to model the furthest along token.

2:44:38 I think that, those aren't free because they have to be

2:44:43 accompanied by a lot of compute or the right data.

2:44:47 How many sequences of 100,000 tokens do you have in the world,

2:44:51 and where do you get these?

2:44:53 It just ends up being pretty expensive to scale them.

2:44:56 We've gotten pretty quickly to a million tokens of input context length.

2:45:00 I would expect it to keep increasing and get

2:45:03 to 2 million or 5 million this year, but I don't expect it to go to 100 million.

2:45:07 That would be like a true breakthrough,

2:45:09 and I think those breakthroughs are possible.

2:45:11 I think of the continual learning thing as a research problem where there could

2:45:15 be a breakthrough that just makes transformers

2:45:17 work way better at this and it's cheap.

2:45:20 These things could happen with so much scientific attention.

2:45:22 But turning the crank, it'll be consistent increases over time.

2:45:28 Looking at the extremes, I think there's, again, no free lunch.

2:45:30 So, the one extreme to make it cheap: you have, let's say,

2:45:33 an RNN that has a single state

2:45:35 where you save everything from the previous stuff.

2:45:37 It's like a specific fixed-size thing, so you never really grow the memory

2:45:43 because you are stuffing everything into one state,

2:45:46 but then the longer the context gets, the more information you forget because

2:45:50 you can't compress everything into one state.

2:45:53 Then on the other hand, you have the transformers,

2:45:56 which try to remember every token,

2:45:57 which is great sometimes if you want to look up specific information,

2:46:00 but very expensive because you have the KV cache that grows,

2:46:04 the dot product that grows.

2:46:05 But then, like you said, the Mamba layers—I mean,

2:46:08 they kind of have the same problem.

2:46:10 Like an RNN, you try to compress everything into one state;

2:46:12 you're a bit more selective there.

2:46:14 But then I think it's like this Goldilocks zone again.

2:46:17 With Nemotron 3, they found a good ratio of how many attention layers do you

2:46:22 need for the global information where everything

2:46:24 is accessible compared to having these compressed states.

2:46:27 And I think that's how we will scale more—by finding better,

2:46:32 let's say, ratios in the Goldilocks zone,

2:46:35 like between making computing cheap enough to run,

2:46:39 but then also making it powerful enough to be useful.

2:46:43 And one more plug here, the Recursive Language Model paper,

2:46:47 that is one of the papers that tries to kind of address the long context thing.

2:46:51 So what they found is essentially instead of stuffing everything

2:46:55 into this long context if you break it up into multiple smaller tasks,

2:46:59 so you save memory by having multiple smaller cores,

2:47:03 you can actually get better accuracy than

2:47:05 having the LLM try everything all at once.

2:47:08 I mean, it's a new paradigm.

2:47:10 We will see, you know, there might be other flavors of that.

2:47:13 So I think with that, we will still make improvement on long context,

2:47:17 but then also, like Nathan said, I think the problem is for pre-training itself,

2:47:20 we don't have as many long context documents as other documents.

2:47:24 So it's harder to study basically how LLMs

2:47:28 behave and stuff like that on that level.

2:47:31 There are some rules of thumb where

2:47:33 essentially you pre-train a language model, like OLMo.

2:47:35 we pre-trained at like 8K context length and then extended to 32K with training.

2:47:39 And there are some rules of thumb

2:47:41 where you're essentially doubling the training context length,

2:47:44 it takes like 2X compute,

2:47:45 and then you can normally like 2 to 4X the context length again.

2:47:50 So I think a lot of it ends up being

2:47:52 kind of compute bound at pre-training, which is in this...

2:47:55 Like we talked about,

2:47:56 everyone talks about this big increase in compute for the top labs this year,

2:47:59 and that should reflect in some longer context windows.

2:48:02 But I think on the post-training side, there are some more interesting things.

2:48:04 As we have agents, the agents are gonna manage this context on their own,

2:48:08 where now agents, people that use Claude Code a lot dread the compaction,

2:48:12 which is when Claude takes its entire full 100,000

2:48:14 tokens of work and compacts it into a bulleted list.

2:48:17 But what the next models will do—and I'm sure people are already

2:48:22 working on this—is essentially the model can control when it compacts and how.

2:48:26 So you can essentially train your RL algorithm where compaction is an action-

2:48:30 ...where it shortens the history and then the problem formulation will be,

2:48:34 "I want to keep the maximum evaluation scores that I

2:48:38 have gotten while the model compacts its history to the minimum

2:48:42 length." Because then you have the minimum amount of tokens

2:48:44 that you need to do this kind of compounding autoregressive prediction.

2:48:47 So there are actually pretty nice problem setups in this, where the...

2:48:51 Like these agentic models learn to use their context

2:48:54 in a different way than just plow forward.

2:48:57 One interesting recent example would be DeepSeek-V3.2,

2:49:00 where they had a sparse attention mechanism

2:49:03 where they have essentially a very efficient, small, lightweight indexer.

2:49:07 And instead of attending to all tokens, it selects:

2:49:10 "What tokens do I actually need?" I mean,

2:49:13 it almost comes back to the original idea of attention where you are selective,

2:49:17 but attention is always on, you have maybe zero weight on some of them,

2:49:21 but you use them all.

2:49:22 But they are even more like, "Let's just mask that out or not even do

2:49:27 that." And even with sliding window attention in OLMo,

2:49:30 that is also kind of that idea.

2:49:32 You have a rolling window where you keep it fixed,

2:49:34 because you don't need everything.

2:49:35 Occasionally, some layers you might, but it's wasteful.

2:49:38 But right now, I think, if you use everything,

2:49:40 you're on the safe side—it gives you the best

2:49:42 bang for the buck because you never miss information.

2:49:44 And I think this year will be more about figuring out,

2:49:48 like you said, how to be smarter about that.

2:49:51 Right now, people want to have the next state-of-the-art,

2:49:54 and the state-of-the-art happens to be the brute-force, expensive thing.

2:49:59 And then once you have that, as you said, keep that accuracy,

2:50:03 but let's see how we can do that cheaper now, with tricks.

2:50:07 Yeah.

2:50:07 All this scaling thing.

2:50:08 The reason we get the Claude 4.5 Sonnet model first is because you

2:50:13 can train it faster and you're not hitting these compute walls as soon.

2:50:16 They can just try a lot more things and get the model faster,

2:50:18 even though the bigger model is actually better.

2:50:22 I think we should say that there's a lot

2:50:23 of exciting stuff going on in the AI space.

2:50:25 My mind has recently been really focused on robotics.

2:50:29 Today, we almost entirely didn't talk about robotics.

2:50:33 There's a lot of stuff on image generation, video generation.

2:50:38 I think it's fair to say that the most

2:50:42 exciting research work in terms of the amount, intensity,

2:50:45 and fervor is in the LLM space, which is why I think it's justified for us

2:50:50 to focus on the LLMs that we're discussing.

2:50:53 But it'd be nice to bring in certain things that might be useful.

2:50:57 For example, world models— there's growing excitement on that.

2:51:00 Do you think there will be any use

2:51:03 in this coming year for world models in the LLM space?

2:51:06 Yes, I do think so.

2:51:09 Also with LLMs, what's interesting here is

2:51:12 that if we unlock more LLM capabilities,

2:51:14 it also automatically unlocks all the other

2:51:17 fields because it makes progress faster.

2:51:20 A lot of researchers and engineers use LLMs for coding.

2:51:24 So even if they work on robotics,

2:51:27 if you optimize these LLMs that help with coding, it pays off.

2:51:31 But then, yes, world models are interesting.

2:51:34 It's basically where you have the model run

2:51:37 a simulation of the world in a sense,

2:51:39 like a little toy thing of the real thing, which can,

2:51:43 again, unlock capabilities regarding data the LLM is not aware of.

2:51:49 It can simulate things.

2:51:50 And I think LLMs happen to work

2:51:55 well by pre-training and doing next-token prediction.

2:51:59 But we could do this even more sophisticatedly in a sense.

2:52:03 I think there was a paper by Meta, a paper called World Models.

2:52:09 So where they basically apply the concept of world models to LLMs again,

2:52:14 where instead of just having next-token prediction and verifiable rewards,

2:52:18 checking the answer correctness,

2:52:20 they also make sure the intermediate variables are correct.

2:52:23 You know, it's kind of like the model

2:52:25 is learning basically a code environment in a sense.

2:52:28 And I think this makes a lot of sense.

2:52:30 It's just expensive to do, but it is making things more sophisticated,

2:52:37 like modeling the whole thing, not just the result.

2:52:42 And so it can add more value.

2:52:45 I remember when I was a grad student, there is a...

2:52:51 competition called CASP, I think, where they do protein structure prediction.

2:52:56 They predict the structure of a protein that is not solved yet at that point.

2:53:02 So in a sense, this is actually great,

2:53:04 and I think we need something like that for LLMs also,

2:53:07 where you do the benchmark, but no one does.

2:53:09 You hand in the results, but no one knows the solution.

2:53:12 And then after the fact, someone reveals that.

2:53:14 But, AlphaFold, when it came out, it crushed this benchmark.

2:53:20 I mean, there were also multiple iterations, but I remember the first one.

2:53:25 I'm not an expert in that subject,

2:53:28 but the first one explicitly modeled the physical interactions of the...

2:53:32 You know, the physics of the molecule.

2:53:34 Also the angles, impossible angles.

2:53:35 And then in the next version,

2:53:37 I think they got rid of this, and just with brute force, scaling it up.

2:53:40 And I think with LLMs,

2:53:42 we are currently in this brute force scaling because it just happens to work.

2:53:45 But I do think at some point it might make sense to bring back this thing.

2:53:50 And I think with world models,

2:53:52 I think that is where I think that might be actually quite cool.

2:53:56 I mean, yeah.

2:53:57 And of course, also for robotics, which is completely unrelated to LLMs.

2:54:03 Yeah.

2:54:03 And robotics is very explicit.

2:54:05 So there's the problem of locomotion or manipulation.

2:54:08 Locomotion is much more solved, especially in the learning domain.

2:54:11 But there's a lot of value, just like with the initial protein folding systems,

2:54:14 bringing in the traditional model-based methods.

2:54:17 So you don't...

2:54:19 it's unlikely that you can just learn the manipulation or the whole body,

2:54:25 local manipulation problem end to end.

2:54:28 That's the dream.

2:54:28 But then you realize when you look at the magic

2:54:31 of the human hand and the complexity of the real world,

2:54:35 it's really hard to learn this all the way through,

2:54:37 the way I guess AlphaFold 2 didn't.

2:54:41 I'm excited about the robotic learning space.

2:54:43 I think it's collectively getting supercharged by all

2:54:46 the excitement and investment in language models generally,

2:54:49 where the infrastructure for training transformers,

2:54:52 which is a general modeling thing, is becoming world-class industrial tooling,

2:54:59 where wherever there was a limitation for robotics, it's just way better.

2:55:03 There's way more compute.

2:55:04 And then on top of that, they take these language models as kind

2:55:07 of central units where you can do

2:55:09 interesting explorative work around something that already works.

2:55:12 And then I see it emerging as, kind of like we talked about,

2:55:17 Hugging Face transformers and Hugging Face.

2:55:18 I think when I was at Hugging Face,

2:55:20 I was trying to get this to happen, but it was too early.

2:55:22 It's like these open robotic models on Hugging Face,

2:55:26 and having people be able to contribute data and fine-tune them.

2:55:29 I think we're much closer now that the investment in robotics and self-driving

2:55:33 cars is related and it enables this, where once you get to the point

2:55:37 where you can have this sort of ecosystem where somebody can download a robotics

2:55:41 model and maybe fine-tune it to their robot or share datasets across the world.

2:55:45 There's some work in this area like RTX,

2:55:48 I think it was a few years ago, where people are starting to do that.

2:55:52 But once they have this ecosystem, it'll look very different.

2:55:54 And then this whole post-ChatGPT boom is putting more resources

2:55:58 into that, which I think is a very good area for doing research.

2:56:02 This is also resulting in much better, more accurate,

2:56:05 and more realistic simulators being built,

2:56:07 closing the sim-to-real gap in the robotic space.

2:56:10 But, you know, you mentioned a lot of excitement

2:56:13 in the robotics space and a lot of investment.

2:56:16 The downside of that, which happens in hype cycles,

2:56:20 I personally believe, and most robotics people believe, that it's not...

2:56:24 Robotics is not going to be solved

2:56:27 at the time scale as being implicitly or explicitly promised.

2:56:32 And so what happens when there's all these robotics companies

2:56:36 that spring up and then they don't have a product that works?

2:56:41 Then there's going to be this kind of crash of excitement,

2:56:44 which is nerve-wracking.

2:56:45 Hopefully something else will come in and keep swooping in so

2:56:50 that the continued development of some of these ideas keeps going.

2:56:54 I think it's also related to the continual learning issue,

2:56:57 essentially, where the real world is so complex.

2:57:00 With LLMs, you don't really need to have something learn for the user,

2:57:05 because there are a lot of things everyone has to do.

2:57:08 Everyone maybe wants to, I don't know,

2:57:10 fix their grammar in their email or code or something like that.

2:57:14 It's more constrained, so you can kind of prepare the model for that.

2:57:17 But preparing the robot for the real world is harder.

2:57:20 I mean, you have the robotic foundation models,

2:57:22 and you can learn certain things like grasping things.

2:57:27 But then again, everyone's house is different.

2:57:30 It's so different, and that is, I think,

2:57:33 where the robot would have to learn on the job, essentially.

2:57:36 And that, I guess, is the bottleneck right now:

2:57:39 how to, customize it on the fly, essentially.

2:57:42 I don't think I can possibly understate the importance of the thing that doesn't

2:57:48 get talked about almost at all by robotics folks or anyone, which is safety.

2:57:53 All the interesting complexities we talk about learning,

2:57:55 all the failure modes and failure cases, everything we've been talking about

2:57:59 with LLMs—sometimes they fail in interesting ways.

2:58:01 All of that is fun and games in the LLM space.

2:58:06 In the robotic space, in people's homes,

2:58:09 across millions of minutes and billions of interactions,

2:58:14 you really are almost allowed to fail never.

2:58:17 When you have embodied systems that are put out there in the real world,

2:58:23 you just have to solve so many problems you never thought you'd

2:58:28 have to solve when just thinking about the general robot learning problem.

2:58:33 I'm so bearish on in-home learned robots for consumer purchase.

2:58:38 I'm very bullish on self-driving cars,

2:58:39 and I'm very bullish for robotic automation, e.g.,

2:58:43 like Amazon distribution where Amazon has built whole new

2:58:46 distribution centers designed for robots first rather than humans.

2:58:49 There's a lot of excitement in AI

2:58:51 circles about AI enabling automation and mass-scale manufacturing,

2:58:54 and I do think that the path to robots doing that is more reasonable,

2:58:57 where it's a thing that is designed and optimized to do

2:59:03 a repetitive task that a human could conceivably do but doesn't want to.

2:59:07 And then I'm much, but it's also going

2:59:10 to take a lot longer than people probably predict.

2:59:14 I think the leap from AI singularity to we

2:59:18 can now scale up mass manufacturing in the US because

2:59:21 we have a massive AI advantage is one that is

2:59:25 troubled by a lot of political and other challenging problems.

2:59:32 Let's talk about timelines, specifically timelines to AGI or ASI.

2:59:38 Is it fair, as a starting point,

2:59:40 to say that nobody really agrees on the definitions of AGI and ASI?

2:59:46 I kind of think there's a lot of disagreement,

2:59:48 but I've been getting pushback where a lot of people kind of say the same thing,

2:59:53 which is like a thing that could reproduce most digital economic work.

2:59:57 So, the remote worker is a fairly reasonable example.

3:00:01 And I think OpenAI's definition is somewhat related to that, which is like an AI

3:00:06 that can do a lot of economically valuable

3:00:08 tasks—which I don't really love as a definition,

3:00:11 but I think it could be a grounding point,

3:00:15 because language models today, while immensely powerful,

3:00:19 are not this remote worker drop-in.

3:00:21 And there are things that could be done

3:00:24 by an AI that are way harder than remote work,

3:00:27 which are like finding an unexpected

3:00:30 scientific discovery that you couldn't even posit,

3:00:32 which would be an example of something

3:00:34 that somebody says is an artificial superintelligence problem.

3:00:37 Or, taking in all medical records and finding

3:00:43 linkages across certain illnesses that people didn't know,

3:00:46 or figuring out that some common drug can treat some niche cancer.

3:00:50 They would say that that is a superintelligence thing.

3:00:52 So these are kind of natural tiers.

3:00:54 My problem with it is that it becomes deeply entwined

3:00:58 with the quest for meaning of AI and these religious aspects to it.

3:01:03 So there's different paths you can take it.

3:01:06 And I don't even know if the remote worker

3:01:09 is a good definition because what exactly is that?

3:01:12 I actually, I mean, I like...

3:01:14 I don't know if you like the originally titled AI27 report.

3:01:18 They focus more on code and research taste,

3:01:22 so the target there is the superhuman coder.

3:01:25 So they have several milestone systems:

3:01:28 Superhuman coders, superhuman AI researcher,

3:01:31 then superintelligent AI researcher,

3:01:33 and then the full ASI, artificial superintelligence.

3:01:37 But after you develop the superhuman coder, everything else follows quickly.

3:01:45 There, the task is to have fully autonomous, automated coding.

3:01:52 So any kind of coding you need to do

3:01:54 in order to perform research is fully automated.

3:01:57 And from there, humans would be doing AI research together with that system,

3:02:02 and they will quickly be able to develop

3:02:04 a system that can actually do the research for you.

3:02:07 That's the idea.

3:02:09 And initially their prediction was 2027, 2028,

3:02:12 and now they've pushed it back by three to four years to 2031 (mean prediction).

3:02:19 Probably my prediction is even beyond 2031, but at least you can,

3:02:24 in a concrete way, think about how

3:02:28 difficult it is to fully automate programming.

3:02:31 Yeah, I disagree with some of their presumptions

3:02:34 and dynamics on how it would play out,

3:02:36 but I think they did good work in the scenario-defining

3:02:39 milestones that are concrete and tell a useful story,

3:02:42 which is why the reach for this AI 2027 document transcended Silicon Valley.

3:02:47 It's because they told a good story and they

3:02:50 did a lot of rigorous work to do this.

3:02:53 I think the camp that I fall into is that AI is so-called

3:02:56 "jagged," which will be excellent at some things and really bad at some things.

3:03:00 I think that when they're close to this automated software engineer,

3:03:05 what it will be good at is traditional ML systems and frontend,

3:03:09 the model is excellent at; but distributed ML,

3:03:12 the models are actually quite bad at because there's

3:03:14 so little training data on doing large-scale distributed learning.

3:03:17 This is something we already see, and I think this will just get amplified.

3:03:21 And then it's kind of messier in these trade-offs,

3:03:24 like how you think AI research works and so on.

3:03:28 So you think basically a superhuman coder is almost unachievable,

3:03:31 because of the jagged nature of the thing,

3:03:33 you're just always going to have gaps in capabilities?

3:03:38 I think it's assigning completeness to something where

3:03:41 the models are already superhuman at some types of code.

3:03:44 I think that will continue.

3:03:46 And people are creative, so they'll utilize these incredible abilities to fill

3:03:50 in the weaknesses of the models and move really fast.

3:03:53 There'll always be this dance for a long time between

3:03:57 the humans enabling the thing that the model can't do.

3:04:00 And the best AI researchers are the ones that can enable this superpower.

3:04:04 And I think those lines lead to what we already see.

3:04:06 I think like Claude Code for building a website,

3:04:08 you can stand up a beautiful website in a few

3:04:10 hours or do data going to keep getting better,

3:04:12 and we'll pick up some new coding skills along the way.

3:04:17 Linking to what's happening in big tech,

3:04:23 this AI 2027 report leans into the singularity idea,

3:04:27 whereas I think research is messy,

3:04:30 social and largely in the data in ways that AI models can't process.

3:04:34 But what we do have today is really powerful and these tech companies

3:04:39 are all collectively buying into this with tens

3:04:41 of billions of dollars of investment.

3:04:43 So we are going to get some much better version of ChatGPT,

3:04:47 a much better version of Claude Code than we already have.

3:04:50 I think it's just hard to predict where that is going,

3:04:53 but the bright clarity of that future is why some of the most

3:04:57 powerful people in the world are putting so much money into this.

3:05:00 And I think it's just kind of small differences between

3:05:04 like—we don't actually know what a better version of ChatGPT is,

3:05:07 but also, can it automate AI research?

3:05:10 I would say probably not, at least in this timeframe.

3:05:14 Big tech is going to spend $100 billion much faster than

3:05:17 we get an automated AI researcher that enables an AI research singularity.

3:05:23 So you think your prediction would be— if this is

3:05:26 even a useful milestone or more than 10 years out?

3:05:31 I would say less than that on the software side,

3:05:33 but I think longer than that on things like research.

3:05:37 Well, let's just for fun try to imagine

3:05:40 a world where all software writing is fully automated.

3:05:43 Can you imagine that world?

3:05:46 By the end of this year,

3:05:48 the amount of software that'll be automated will be so high.

3:05:50 But it'll be things like trying to train a model with RL

3:05:54 and you need to have multiple bunches of GPUs communicating with each other.

3:05:59 That'll still be hard, but it'll be much easier.

3:06:02 One way to think about this— the full automation

3:06:04 of programming—is just thinking of lines of useful code written,

3:06:10 the fraction of that to the number of humans in the loop.

3:06:15 So presumably there'll be for a long

3:06:17 time humans in the loop of software writing.

3:06:19 It'll just be fewer and fewer relative to the amount of code written.

3:06:23 Right?

3:06:24 And the superhuman coder—I think the presumption there is it goes to zero,

3:06:29 the number of humans in the loop.

3:06:31 What does that world look like when the number

3:06:33 of humans in the loop is in the hundreds, not in the hundreds of thousands?

3:06:39 I think software engineering will be driven

3:06:41 more to system design and goals of outcomes,

3:06:44 where I do think software is largely going to be.

3:06:47 I think this has been happening over the last few weeks,

3:06:50 where people have gone from a month ago saying, "Oh yeah,

3:06:53 agents are kind of slop," which is a famous Karpathy quote,

3:06:56 to what is a little bit of a meme—the industrialization

3:07:01 of software when anyone can just create software with their fingerprints.

3:07:04 I do think we are closer to that side of things,

3:07:07 and it takes direction and understanding how systems

3:07:11 work to extract the best from the language models.

3:07:14 I think it's hard to accept the gravity of how much is going to change

3:07:18 with software development and how many more people

3:07:20 can do things without ever looking at the code.

3:07:22 I think what's interesting is to think about whether

3:07:25 these systems will be independent—completely independent in the sense

3:07:27 that, while I have no doubt that LLMs will

3:07:30 kind of at some point solve coding in a sense,

3:07:33 like calculators solve calculating, right?

3:07:35 So at some point, humans developed a tool where

3:07:38 you never need a human to calculate that number.

3:07:41 You just type it in, and it's an algorithm.

3:07:43 You can do it in that sense.

3:07:45 And I think that's the same probably for coding.

3:07:48 But the question isn't...

3:07:50 I think what will happen is, you will just say,

3:07:52 "Build that website." It will make a really good website,

3:07:55 and then you maybe refine it.

3:07:56 But will it do things independently where...

3:07:59 Will you still be having humans asking the AI to do something?

3:08:05 Like will there be a person to say, "Build that website?" Or will there be

3:08:09 AI that just builds websites or something?

3:08:12 I think talking about building websites is—- Too simple.

3:08:17 The problem with websites and the problem with the web, you know,

3:08:21 HTML and all that kind of stuff, it's very resilient to just...

3:08:25 slop.

3:08:25 It will show you slop; it's good at showing slop.

3:08:28 I would rather think of safety-critical systems,

3:08:32 like asking AI to end-to-end generate something that manages logistics,

3:08:40 or manages cars, a fleet of cars—all that kind of stuff.

3:08:43 So it end-to-end generates that for you.

3:08:46 I think a more intermediate example is

3:08:47 take something like Slack or Microsoft Word.

3:08:50 I think if organizations allow it,

3:08:52 AI could very easily implement features end-to-end and do a fairly

3:08:57 good job for like things that you want to try.

3:09:00 You want to add a new tab in Slack that you want to use,

3:09:03 and I think AI will be able to do that pretty well.

3:09:06 Actually, that's a really great example.

3:09:07 How far away are we from that?

3:09:09 Like this year.

3:09:12 See, I don't know.

3:09:13 I don't know.

3:09:15 I guess I don't know how bad production codebases are,

3:09:17 but I think that within...

3:09:18 on the order of a few years, a lot of people are going to be pushed

3:09:22 to be more of a designer and product manager,

3:09:24 where you have multiple of these agents that can try things for you and they

3:09:28 might take one to two days to implement a feature or attempt to fix a bug.

3:09:32 And you have these dashboards,

3:09:34 which I think Slack is actually a good dashboard where

3:09:36 your agents will talk to you and you'll then give feedback.

3:09:40 But things like, if I make a website, like,

3:09:43 "Do you want a passable logo?" I think

3:09:45 these cohesive design things and the style is

3:09:48 going to be very hard for models and deciding on what to add the next time.

3:09:54 I just...

3:09:54 Okay.

3:09:55 I hang out with a lot of programmers and some

3:09:57 of them are a little bit on the skeptical side in general.

3:10:03 That's just their vibe.

3:10:05 I just think there's a lot of complexity

3:10:07 involved in adding features to complex systems.

3:10:10 Like, if you look at the browser, Chrome.

3:10:13 If I wanted to add a feature,

3:10:15 if I wanted to have tabs as opposed to up top, I want them on the left side.

3:10:21 Interface-wise, right?

3:10:22 I think we're not...

3:10:23 This is not a next-year thing.

3:10:26 One of the Claude releases this year, one of their tests was:

3:10:28 we give it a piece of software and leave Claude to run to recreate it entirely,

3:10:32 and it could already almost rebuild Slack from scratch,

3:10:36 just given the parameters of the software

3:10:38 and left in a sandbox environment to do that.

3:10:41 So the "from scratch" part, I like almost better.

3:10:44 So it might be that smaller and newer companies are advantaged,

3:10:47 and they're like, "We don't have the bloat and complexity,

3:10:51 and therefore this feature exists."- And I

3:10:54 think this gets to the point you mentioned,

3:10:56 that some people you talk to are skeptical.

3:10:58 I think that's not because the LLM can't do X, Y, Z.

3:11:02 It's because people don't want it to do it this way.

3:11:05 Some of that could be a skill issue on the human side.

3:11:08 We have to be honest with ourselves.

3:11:10 And some of that could be an underspecification issue.

3:11:13 So, programming, it's like you're just assuming...

3:11:17 This is like an issue with communication in relationships and friendships.

3:11:22 You're assuming the LLM is supposed to read your mind.

3:11:26 This is where spec-driven design is really important.

3:11:28 Using natural language to specify what you want.

3:11:32 If you talk to people at the labs,

3:11:34 they use these in their training and production code.

3:11:37 Claude Code is built with Claude Code,

3:11:39 and they all use these things extensively.

3:11:41 Dario talks about how much of Claude's code...

3:11:44 It's like these people are slightly ahead in terms

3:11:48 of the capabilities they have and what they probably spend on inference.

3:11:53 They could spend 10 to 100x as much as we're

3:11:56 spending on a lowly $100 or $200 a month plan.

3:11:59 They truly let it rip.

3:12:01 And I think that, with the pace of progress that we have,

3:12:06 it seems like- a year ago we didn't have

3:12:09 Claude Code and we didn't really have reasoning models.

3:12:11 The difference between sitting here today and what

3:12:13 we can do with these models is significant,

3:12:16 and there's a lot of low-hanging fruit to improve them.

3:12:21 The failure modes are pretty dumb.

3:12:23 Like- "Claude, you tried to use a CLI command I don't have installed 14 times,

3:12:27 and then I sent you the command to run."

3:12:30 That, from a modeling perspective, is pretty fixable.

3:12:33 So I don't know.

3:12:34 I agree with you.

3:12:36 I've been becoming more and more bullish in general.

3:12:38 Speaking to what you're articulating, I think it is a human skill issue.

3:12:44 Anthropic is leading the way, along with other companies,

3:12:49 in understanding how to best use the models for programming;

3:12:53 therefore, they're effectively using them.

3:12:54 There are a lot of programmers on the outskirts who don't...

3:12:58 I mean, there's not a really good guide on how to use them.

3:13:03 People are trying to figure it out, but-- It might be very expensive.

3:13:06 The entry point might be $2,000 a month,

3:13:09 which is only for tech companies and rich people.

3:13:12 That could be it.

3:13:14 But it might be worth it.

3:13:15 If the final result is a working software system, it might be worth it.

3:13:19 By the way, it's funny how we converged from the discussion

3:13:22 of timeline to AGI to something more pragmatic and useful.

3:13:25 Is there anything concrete, interesting, useful,

3:13:29 and profound to be said about the timeline to AGI and ASI?

3:13:33 Or are these discussions a bit too detached from the day to day?

3:13:39 There are interesting bets.

3:13:40 A lot of people are trying to do

3:13:42 RLVR— Reinforcement Learning with Verifiable Rewards—in real scientific domains,

3:13:46 where startups with hundreds of millions of funding have wet labs where

3:13:50 they're having language models propose hypotheses

3:13:52 that are tested in the real world.

3:13:54 I would say that they're early, but with the pace of progress,

3:13:59 it's like- ...maybe they're early by six months

3:14:02 and they make it because they were there first,

3:14:05 or maybe they're early by eight years; you don't know.

3:14:08 That type of moonshot to branch this momentum

3:14:13 into other sciences would be very transformative.

3:14:18 If, AlphaFold moments happen in all sorts

3:14:21 of other scientific domains by a startup solving this.

3:14:24 I think there are startups—maybe Harmonic is one—where they're

3:14:27 going all in on language models plus Lean for math.

3:14:31 You had another guest where you talked about this recently,

3:14:34 and we don't know exactly what's going to fall

3:14:37 out of spending $100 million on that model.

3:14:40 Most of them will fail, but a couple might be big breakthroughs that are

3:14:45 very different than ChatGPT or Claude Code type software experiences.

3:14:50 A tool that's only good for a PhD mathematician but makes them 100X effective...

3:14:58 I agree.

3:14:58 I think this will happen in a lot of domains,

3:15:01 especially domains that have a lot of resources,

3:15:05 like finance, legal, and pharmaceutical companies.

3:15:09 But then again, is it really AGI?

3:15:12 Because we are now specializing it again.

3:15:14 Is it really that much different from back

3:15:17 in the day when we had specialized algorithms?

3:15:19 It's just the same thing, way more sophisticated,

3:15:23 but I don't know, is there a threshold for AGI?

3:15:27 I think the real cool thing here is

3:15:29 that we have foundation models we can specialize.

3:15:32 That's like the breakthrough.

3:15:33 Right now, I think we are not there yet because, first, it's too expensive,

3:15:38 but also, ChatGPT doesn't just give away their model to customize it.

3:15:42 I think once that's true...

3:15:44 And I can imagine this as a business model, where OpenAI says at some point,

3:15:51 "Hey, Bank of America,

3:15:52 for $100 million we will do your custom model," something like that.

3:15:56 I think that will be the huge economic value-add.

3:15:59 The other thing, though, is also...

3:16:02 Companies, I mean, what is the differentiating factor?

3:16:06 If everyone uses the same LLM, if everyone uses ChatGPT,

3:16:10 they will all do the same thing.

3:16:12 Well, if everyone is moving in lockstep,

3:16:15 but companies want to have a competitive advantage,

3:16:18 there is no way around using some of their private data and specializing.

3:16:24 It's gonna be interesting.

3:16:26 Seeing the pace of progress, it does feel like things are coming.

3:16:30 I don't think the AGI and ASI thresholds are particularly useful.

3:16:36 I think the real question, and this relates to the remote worker thing,

3:16:40 is: when are we going to see a big, obvious leap in economic impact?

3:16:47 Because currently there's not been an obvious leap

3:16:51 in the economic impact of LLM models, for example.

3:16:55 And that's, you know, aside from AGI or ASI, all that stuff,

3:16:59 there's a real question of, "When are

3:17:03 we going to see a GDP..." "...jump?"- Yeah, what is the GDP made up of?

3:17:08 A lot of it is financial services, so I don't know what this is.

3:17:13 Right, GDP is a-- It's just hard for me to think about the GDP bump,

3:17:17 but I would say that software development becomes valuable in a different way,

3:17:22 when you no longer have to look at the code anymore.

3:17:25 When Claude Code will make you a small business.

3:17:29 Which is essentially, Claude can set up your website,

3:17:31 your bank account, your email, and your whatever else.

3:17:34 And you just have to express what you're trying to put into the world.

3:17:39 That's not just an enterprise market, but it is hard.

3:17:42 I don't know how you get people to try doing that.

3:17:45 I guess if ChatGPT can do it—people are trying ChatGPT.

3:17:49 I think it boils down to the scientific question of, "How hard

3:17:52 is tool use to solve?" Because a lot of the stuff you're implying,

3:17:57 the remote work stuff, is tool use.

3:17:59 It's like...

3:18:00 computer use, like how you have an LLM that goes out there, this agentic system,

3:18:06 and does something in the world, and only screws up 1% of the time.

3:18:12 Computer use-- Or less.

3:18:12 ...is a good example of what labs care about

3:18:14 and we haven't seen a lot of progress on.

3:18:16 We saw multiple demos in 2025 of, like,

3:18:20 Claude can use your computer, or OpenAI had operator, and they all suck.

3:18:24 They're investing money in this, and I think that'll be a good example.

3:18:29 Whereas actually, something where it just seems

3:18:32 like taking over the whole screen seems

3:18:34 a lot harder than having an API that they can call in the back end.

3:18:39 For some of that, you have to set up

3:18:41 a different environment for them all to work in.

3:18:43 They're not working on your MacBook;

3:18:45 they are individually interfacing with Google and Amazon and Slack,

3:18:49 and they handle all these things in a very different way than humans do.

3:18:53 So some of this might be structural blockers.

3:18:56 Also, specification-wise, I think the problem is for arbitrary tasks, well,

3:19:01 you still have to specify what you want your LLM to do.

3:19:04 And how do you do that?

3:19:06 What is the environment?

3:19:07 How do you specify?

3:19:08 You can say what the end goal is, but if it can't solve the end goal...

3:19:13 with LLMs, if you ask it for text, it can always clarify or do sub-steps.

3:19:17 How do you put that information into a system that, let's say,

3:19:20 books a travel trip for you?

3:19:22 You can say, "You screwed up my credit card

3:19:24 information," but even to get it to that point,

3:19:27 even to get it to that point, how do you,

3:19:29 as a user, guide the model before it can even attempt that?

3:19:33 I think the interface is really hard.

3:19:36 Yeah, it has to learn a lot about you specifically.

3:19:39 And this goes to continual learning,

3:19:41 about the general mistakes that are made throughout,

3:19:45 and then mistakes that are made through you.

3:19:48 All the AI interfaces are getting set up to ask humans for input.

3:19:51 I think Claude Code we talked about a lot.

3:19:54 It asks feedback and questions.

3:19:55 If it doesn't have enough specification on your plan or your desired goal,

3:19:59 it starts to ask questions,

3:20:01 "Would you rather?" We talked about Memory, which saves across chats.

3:20:06 Its first implementation is kind of odd,

3:20:08 where it'll mention my dog's name or something in a chat.

3:20:11 I'm like, "You don't need to be subtle about this.

3:20:14 I don't care." But things that are emerging, ChatGPT has the Pulse feature.

3:20:19 Which is like a curated couple of paragraphs with links to something to look

3:20:24 at, and people talk about how models are going to ask you questions.

3:20:28 Which I think is a very...

3:20:31 It's probably going to work.

3:20:33 The language model knows you had a doctor appointment and asks, "Hey,

3:20:36 how are you feeling after that?" Which

3:20:38 again goes into the territory where humans

3:20:40 are very susceptible to this, and there's a lot of social change to come.

3:20:45 But also, they're experimenting with having the models engage.

3:20:48 Some people like this Pulse feature,

3:20:50 which processes your chats and automatically searches

3:20:53 for information and puts it in the app.

3:20:56 So there are a lot of things coming.

3:20:59 I used that feature before,

3:21:00 and I always feel bad because it does that every day, and I rarely check it out.

3:21:05 It's like, how much compute is burned

3:21:07 on something I don't even look at, you know?

3:21:10 It's kind of like, "Oh..."- There's also a lot of idle compute in the world,

3:21:14 so don't feel too bad.

3:21:16 Okay.

3:21:17 Do you think new ideas might be needed?

3:21:20 Is it possible that the path to AGI,

3:21:22 however we define that, to solve computer use more generally,

3:21:26 to solve biology and chemistry and physics—sort

3:21:31 of the Dario Amodei definition of AGI?

3:21:35 Do you think it's possible that totally new ideas are needed?

3:21:42 Non-LLM, non-RL ideas?

3:21:45 What might they look like?

3:21:47 We're going into philosophy land a bit.

3:21:51 For something like a singularity to happen, I would say yes.

3:21:54 The new ideas could be architectures or training algorithms,

3:21:58 fundamental deep learning things.

3:22:00 But in that nature, they're pretty hard to predict.

3:22:04 I think we won't get very far even without those advances.

3:22:08 We might get the software solution,

3:22:10 but it might stop at software and not do computer use without more innovation.

3:22:15 So I think that a lot of progress will be coming, but if you're gonna zoom out,

3:22:20 there's still ideas in the next 30 years that are gonna look like

3:22:24 that was a major scientific innovation that enabled the next chapter of this.

3:22:29 And I don't know if it comes in one year or in 15 years.

3:22:33 Yeah.

3:22:33 I wonder if the bitter lesson holds true for the next 100 years,

3:22:36 what that looks like.

3:22:38 If scaling laws are fundamental in deep learning,

3:22:40 I think the bitter lesson will always apply,

3:22:42 which is compute will become more abundant, but even within abundant compute,

3:22:48 the ones that have a steeper scaling law slope or a better offset— like,

3:22:53 this is a 2D plot of performance

3:22:55 and compute—and like even if there's more compute available,

3:22:58 the ones that get 100x out of it will win.

3:23:01 It might be something like literally

3:23:04 computer clusters orbiting Earth with solar panels.

3:23:09 The problem with that is heat dissipation.

3:23:11 You get all the radiation from the sun and don't have any air to dissipate heat.

3:23:15 But there is a lot of space to put clusters.

3:23:17 There's a lot of solar energy there

3:23:19 and you could figure out the heat dissipation,

3:23:21 as there is a lot of energy and there probably could

3:23:24 be engineering will to solve the heat problem— so there could be.

3:23:27 Is it possible—and we should say that it definitely is

3:23:30 possible— that we're basically going to be plateauing this year?

3:23:36 Not in terms of— the system capabilities,

3:23:40 but what they actually mean for human civilization.

3:23:44 So on the coding front, really nice websites will be built.

3:23:50 Very nice auto-complete.

3:23:53 Very nice way to understand code bases and maybe help debug,

3:23:59 but really just a very nice helper on the coding front.

3:24:03 It can help research mathematicians do some math.

3:24:06 It can help you with shopping.

3:24:09 It's a nice helper.

3:24:10 It's Clippy on steroids.

3:24:12 What else?

3:24:13 It may be a good education tool and all that kind of stuff,

3:24:19 but computer use turns out extremely difficult to solve.

3:24:24 So I'm trying to frame the cynical case in all

3:24:28 these domains where there's not a really huge economic impact,

3:24:32 but realize how costly it is to train these systems at every level,

3:24:37 both the pre-training and the inference,

3:24:39 how costly the inference is, the reasoning, all of that.

3:24:43 Like, is that possible?

3:24:44 And how likely is that, do you think?

3:24:47 When you look at the models,

3:24:49 there are so many obvious things to improve and it takes

3:24:52 a long time to train these models and to do this art,

3:24:55 and it'll take us with the ideas that we have multiple years

3:24:59 to actually saturate in terms of whatever

3:25:02 benchmark or performance we are searching for.

3:25:05 It might serve very narrow niches;

3:25:07 like the average ChatGPT 800 million user might not get a lot of benefit out

3:25:12 of this, but it is going to serve

3:25:14 different populations by getting better at different things.

3:25:18 But I think what everybody's chasing now

3:25:21 is a general system that's useful to everybody.

3:25:24 So, okay, so if that's not...

3:25:26 That can plateau, right?

3:25:28 I think that dream is actually kind of dying.

3:25:30 As you talked about with the specialized models where it's like...

3:25:34 And multimodal is often...

3:25:36 Video generation is a totally different thing.

3:25:38 Thing.

3:25:39 "That dream is kind of dying" is a big statement,

3:25:42 because I don't know if it's dying.

3:25:44 If you ask the actual frontier lab people, they...

3:25:46 I mean, they're still chasing it, right?

3:25:48 I do think they are still rushing to get the next model out,

3:25:52 which will be much better than the...

3:25:54 "Much" is a relative term, but it will be better than the previous one.

3:25:58 And I can't see them slowing down.

3:26:00 I just think the gains will be made or felt

3:26:03 more through not only scaling the model, but now...

3:26:08 I feel like there's a lot of tech debt.

3:26:10 It's like, "Well, let's just put the better

3:26:12 model in there." Better model, better model.

3:26:15 And now people are like, "Okay,

3:26:17 let's also at the same time improve everything around it

3:26:19 too." Like the engineering of the context and inference scaling.

3:26:23 The big labs will still keep doing that.

3:26:26 And now also the smaller labs will catch up, because now they are hiring more.

3:26:31 There will be more people and LLMs.

3:26:33 It's kind of like a circle.

3:26:35 They also make them more productive and it's just...

3:26:38 It's like amplification.

3:26:39 I think what we can expect is amplification, but not like a change of any...

3:26:43 not like a paradigm change.

3:26:45 I don't think that is true, but everything will be just amplified and amplified,

3:26:48 and I can see that continuing for a long time, you know?

3:26:53 Yeah.

3:26:53 I guess my statement that the dream is dying

3:26:55 depends on exactly what you think it's gonna be doing.

3:26:58 Like, Claude Code is a general model that can do a lot of things,

3:27:02 but it's not necessarily...

3:27:05 It depends a lot on integrations.

3:27:06 I bet Claude Code could do a fairly good job of doing your email,

3:27:10 and the hardest part is figuring out how to give information to it

3:27:13 and how to get it to be able to send your emails.

3:27:17 But that's just kind of like...

3:27:18 I think it goes back to what is the "one model to rule everything" ethos,

3:27:23 which is just like a thing in the cloud that handles

3:27:26 your entire digital life and is way smarter than everybody.

3:27:29 It's like it's operating in a...

3:27:34 So it's an interesting leap of faith to go

3:27:37 from "Claude Code becomes that," which in some ways is...

3:27:41 There are some avenues for that, but I do think

3:27:45 that the rhetoric of the industry is a little bit different.

3:27:49 I think the immediate thing we will feel next as a normal

3:27:52 person using LLMs will probably be related to something trivial,

3:27:57 like making figures.

3:27:58 Right now, LLMs are terrible at making figures.

3:28:01 Is it because we are getting served the cheap

3:28:04 models with much less inference compute than behind the scenes?

3:28:08 Maybe some.

3:28:09 Like, there are some ways to get better figures, but if you ask today,

3:28:13 ..."Draw a flowchart of X, Y, Z," it's most of the time terrible.

3:28:18 And it is a very simple task for a human.

3:28:20 I think it's almost easier sometimes to draw something than to write something.

3:28:25 Yeah, the multimodal understanding does feel like something that is odd...

3:28:28 ...that it's not better solved.

3:28:31 I think we're not saying one obvious thing that we're not realizing,

3:28:35 that's a gigantic thing that's hard to measure,

3:28:37 which is making all of human knowledge accessible— —to the entire world.

3:28:46 One thing that is hard to articulate is

3:28:49 the huge difference between Google Search and an LLM.

3:28:52 I feel like I can basically ask an LLM anything and get an answer,

3:28:59 and it's doing less and less hallucination.

3:29:04 And that means understanding my own life, figuring out a career trajectory,

3:29:09 solving the problems all around me,

3:29:11 learning about anything through human history.

3:29:16 I feel like nobody's really talking about that, because they

3:29:22 just immediately take it for granted that this is awesome.

3:29:25 That's why everybody's using it: because you get answers for stuff.

3:29:29 Think about the impact across time.

3:29:33 This is not just in the United States; it's all across the world.

3:29:37 Kids throughout the world being able to learn

3:29:40 these ideas— the impact that has across time is probably...

3:29:45 ...That's the real impact.

3:29:48 Talk about GDP; it won't be like a leap.

3:29:51 It'll be...

3:29:52 ...that's how we get to Mars, that's how we build these things,

3:29:56 that's how we have a million new OpenAIs and all the innovation from there.

3:30:00 It's this quiet force that permeates everything: human knowledge.

3:30:06 I agree with you.

3:30:08 In a sense, it makes knowledge more accessible,

3:30:10 but it also depends on what the topic is.

3:30:13 For something like math, you can ask it questions and it answers,

3:30:21 but if you want to learn a topic from scratch,

3:30:26 the sweet spot is still elsewhere.

3:30:28 There are really good math textbooks laid out linearly,

3:30:32 and that is a proven strategy to learn a topic.

3:30:36 It makes sense, if you start from zero,

3:30:39 to use information-dense text to soak it up,

3:30:43 but then you use the LLM to make infinite exercises.

3:30:47 Like, you have problems in a certain area

3:30:49 or have questions that something's- uncertain about certain things,

3:30:53 you ask it to generate example problems, you solve them,

3:30:59 and you need more background knowledge, you ask it to generate that.

3:31:03 But then...

3:31:04 it won't give you anything, let's say, that is not in the textbook.

3:31:10 It's just packaging it differently, if that makes sense.

3:31:13 But then there are things I feel like where

3:31:15 it also adds value in a more timely sense,

3:31:18 where there is no good alternative besides a human doing it on the fly.

3:31:24 For example, if you're planning to go to Disneyland and you

3:31:28 try to figure out which tickets to buy for which park when,

3:31:32 well, there is no textbook on that.

3:31:34 There is no information-dense resource.

3:31:36 There's only the sparse internet, and then there is a lot of value in the LLM.

3:31:40 You just ask it.

3:31:42 You have constraints on traveling these days.

3:31:44 I want to go there and there.

3:31:46 Please figure out what I need, when and from where, what it costs and stuff like

3:31:50 that, and it is a very customized, on-the-fly package.

3:31:56 And this is like one of a thousand examples

3:31:58 of personalized- Personalization is essentially

3:32:01 pulling information from the sparse internet,

3:32:04 the non-information-dense thing where there's no better version that exists.

3:32:09 It just doesn't exist.

3:32:10 You make it almost from scratch.

3:32:12 And if it does exist, it's full of- speaking of Disney World,

3:32:15 full of- what would you call it?

3:32:18 Ad slop.

3:32:19 It's impossible.

3:32:20 Take any city in the world, what are the top 10 things to do?

3:32:27 An LLM is just way better to ask than anything on the internet.

3:32:30 Well, for now, that's because they're subsidized

3:32:32 and they're gonna be paid for by ads.

3:32:36 Oh my goodness.

3:32:37 It's coming.

3:32:38 No.

3:32:39 No.

3:32:39 I mean, I'm hoping there's a very clear indication what's

3:32:43 an ad and what's not an ad in that context.

3:32:46 That's something I mentioned a few years ago.

3:32:49 If, I don't know, if you are looking for a new running shoe,

3:32:52 well, is it a coincidence that Nike maybe comes up first?

3:32:56 Maybe, maybe not.

3:32:58 But I think there are clear laws.

3:33:00 You have to be clear about that.

3:33:02 I think that's what everyone fears.

3:33:04 It's the subtle message in there, but that also brings us to the topic of ads,

3:33:11 where I think this was a thing.

3:33:13 Hopefully, I think for- in 2025, just because I think it's they're still

3:33:19 not making money in other ways right now.

3:33:22 Having ad spots in there...

3:33:24 but the thing is, they couldn't,

3:33:26 because there are alternatives without ads and people

3:33:30 would just flock- to the other products.

3:33:33 It's also just crazy how- yeah, how they're one-upping each other,

3:33:38 spending so much money to just get the users.

3:33:41 I think so.

3:33:42 Like, some Instagram ads— I don't use Instagram,

3:33:44 but I understand the appeal of paying a platform

3:33:48 to find users who will genuinely like your product,

3:33:52 and that is the best case of things like Instagram ads.

3:33:56 But there are also plenty of cases

3:33:58 where advertising is very awful for incentives,

3:34:00 and I think that a world where the power

3:34:04 of AI can integrate with that positive view

3:34:06 of, "I am a person and I have a small business and I want to make the best,

3:34:11 I don't know, damn steak knives in the world,

3:34:13 and I want to sell them to somebody who needs them."

3:34:16 And if AI can make that sort of advertising thing work even better,

3:34:20 that's very good for the world, especially with digital infrastructure,

3:34:24 because that's how the modern web has been built.

3:34:27 But that's not to say that addicting feeds so

3:34:31 that you can show people more content is a good thing.

3:34:35 So, I think that's even what OpenAI would say,

3:34:37 is they want to find a way that can make

3:34:40 the monetization upside of ads while still giving their users agency.

3:34:45 And I personally would think that Google is probably going

3:34:47 to be better at figuring out how to do this, because

3:34:50 they already have ad supply and if they figure out how

3:34:54 to turn this demand in their Gemini app into useful ads,

3:34:57 then they can turn it on.

3:34:59 And somebody will figure it out—I don't know if it's this year,

3:35:03 but there will be experiments with it.

3:35:06 I do think what holds companies back right now

3:35:08 is really just that the competition is not doing it.

3:35:11 It's more like a reputation thing.

3:35:13 It's just, I think people are just afraid

3:35:16 right now of ruining or losing their reputation,

3:35:19 losing users, because it would make headlines if someone launched these ads.

3:35:22 But—- Unless they were great, but the first ads won't be great because it's

3:35:26 a hard problem that we don't know how to solve.

3:35:28 Yeah, I think also the first version of that will likely be something like on X,

3:35:32 like the timeline where you have a promoted post sometimes in between.

3:35:35 It'll be something where it will say "promoted" or something small,

3:35:38 and then there will be an image.

3:35:39 I think right now the problem is: who makes the first move?

3:35:43 If we go 10 years out,

3:35:44 the proposition for ads is that you will make so much money on ads by having

3:35:49 so many users that you can use this to fund better R&D and make better models,

3:35:53 which is why YouTube is dominating the market

3:35:57 for any— Netflix is scared of YouTube.

3:36:01 They have the ads, they make—I pay $28 a month for Premium.

3:36:04 They make at least $28 a month off of me and many other people.

3:36:09 And they're just creating such a dominant position in video.

3:36:12 So I think that's the proposition:

3:36:14 that ads can make you have a sustained advantage.

3:36:17 in what you're spending per user.

3:36:20 But there's so much money in it right now that somebody

3:36:24 starting that flywheel is scary because it's a long-term bet.

3:36:29 Do you think there'll be some crazy big moves this year business-wise?

3:36:33 Like Google or Apple acquiring Anthropic or something like this?

3:36:40 Dario will never sell, but we are starting to see some types of consolidation

3:36:44 with Groq for $20 billion and Scale AI for almost $30

3:36:49 billion and countless other deals like this that are structured

3:36:52 in a way that is detrimental to the Silicon Valley ecosystem,

3:36:57 which is this licensing deal where not everybody gets brought along,

3:37:02 rather than a full acquisition that benefits

3:37:05 the rank-and-file employee by getting their stock vested.

3:37:07 That's a big issue for culture to address

3:37:10 because the startup ecosystem is the lifeblood where, if you join a startup,

3:37:16 even if it's not successful, it might get acquired on a cheap premium

3:37:21 and you'll get paid out for this equity.

3:37:24 These licensing deals are taking the top talent a lot of the time.

3:37:27 The deal for Groq to NVIDIA is rumored to be better to the employees,

3:37:32 but it is still this antitrust-avoiding thing.

3:37:35 But I think that this trend of consolidation will continue.

3:37:39 Me and many smart people I respect

3:37:41 have been expecting consolidation to have happened sooner,

3:37:44 but it seems like some of these things are starting to turn,

3:37:49 but at the same time,

3:37:51 companies are raising ridiculous amounts of money for reasons where I'm like,

3:37:56 "I don't know why you're taking that money." So it's mixed this year,

3:38:01 but some consolidation pressure is starting.

3:38:05 What kind of surprising consolidation will we see?

3:38:07 You say Anthropic is a "never." I mean, Groq is a big one.

3:38:10 Groq with a Q, by the way.

3:38:12 Yeah.

3:38:12 There's just a lot of startups and a very high premium on AI startups.

3:38:16 So there could be a lot of- that kind of stuff, yeah.

3:38:19 $10 billion range acquisitions,

3:38:21 which is really big for a startup that was maybe founded a year ago.

3:38:25 I think Manus.ai...

3:38:27 this company based in Singapore that Meta-founded was founded

3:38:30 eight months ago and then had a $2 billion exit.

3:38:33 I think there will be some

3:38:35 other multi-billion dollar acquisitions, like Perplexity.

3:38:39 Like Perplexity, right?

3:38:40 Yeah, people rumor them to Apple.

3:38:41 I think there's a lot of of pressure and liquidity in AI.

3:38:46 There's pressure on big companies to have outcomes and- I would guess that a big

3:38:52 acquisition gives people leeway to then tell the next chapter of that story.

3:38:56 I guess Cursor—we've been talking about code—somebody acquires Cursor.

3:39:00 if somebody acquires Cursor...

3:39:02 They're in such a good position by having so much user data.

3:39:05 And we talked about continual learning.

3:39:07 They had one of the most interesting sentences in a blog post,

3:39:10 which is that they had their new Composer model,

3:39:13 which was a fine-tune of one of these large Mixture of Expert models from China.

3:39:17 You can know that by asking it or because the model

3:39:21 sometimes responds in Chinese— ...which none of the American models do.

3:39:24 And they had a blog post where they're like,

3:39:26 "We're updating the model weights every 90

3:39:28 minutes based on real-world feedback from people

3:39:30 using it." Which is like the closest

3:39:32 thing to real-world RL happening on a model,

3:39:34 and it's just mentioned in one of their blog posts—- That's incredible.

3:39:37 which is super cool.

3:39:38 And by the way, I should say I use Composer a lot

3:39:40 because one of the benefits it has is that it's fast.

3:39:43 I need to try it 'cause everybody says this.

3:39:45 And there'll be some IPOs potentially.

3:39:48 You think Anthropic, OpenAI, xAI.

3:39:51 They can all raise so much money so easily that they don't feel a need to.

3:39:55 So long as fundraising is easy,

3:39:56 they're not going to IPO because public markets apply pressure.

3:40:00 I think we're seeing in China that the ecosystem's a little

3:40:02 different with both MiniMax and Z.ai applying for, filing IPO paperwork,

3:40:08 which will be interesting to see how the Chinese market reacts.

3:40:11 I actually would guess that it's going to be similarly hypey to the US,

3:40:16 so long as all this is going and not based

3:40:18 on the reality that they're both losing a ton of money.

3:40:21 I wish more of the gigantic American AI startups were public because it

3:40:25 would be very interesting to see how

3:40:26 they're spending money and have more insight.

3:40:28 And also just to give people access to investing in these, because I

3:40:33 think they're some of the most

3:40:35 formidable companies—they're the companies of the era.

3:40:38 And the tradition is now for so many

3:40:40 of the big startups in the US to not go public.

3:40:43 It's like we're still waiting for Stripe and the IPO,

3:40:46 but Databricks definitely didn't.

3:40:47 They raised like a Series G or something.

3:40:50 And I just feel like it's kind

3:40:52 of a weird equilibrium for the market where it's like,

3:40:56 I would like to see these companies go public

3:40:58 and evolve in that way that a company can.

3:41:01 Do you think 10 years from now some

3:41:03 of the frontier model companies are still around?

3:41:06 Anthropic, OpenAI?

3:41:08 I definitely don't see it as a winner-takes-all unless there truly is

3:41:12 some algorithmic secret that one of them finds that lets this flywheel.

3:41:16 Because the development path is so similar for all of them.

3:41:19 Google and OpenAI have all the same products, and then Anthropic's more focused,

3:41:24 but when you talk to people it sounds

3:41:26 like they're solving a lot of the same problems.

3:41:28 So I think...

3:41:29 and there's offerings that'll spread out.

3:41:30 There's a lot of...

3:41:31 it's a very big cake being made that people are going to take money out of.

3:41:37 I don't want to trivialize it,

3:41:39 but OpenAI and Anthropic are primarily LLM service providers.

3:41:44 And some of the other companies like Google and xAI,

3:41:48 linked to X, do other stuff too.

3:41:51 And so it's very possible, if AI becomes more commodified,

3:41:56 that the companies just providing LLMs will die.

3:42:00 I think the advantage they have is a lot of users,

3:42:03 and I think they will just pivot.

3:42:05 Like Anthropic, I think, pivoted.

3:42:08 I don't think they originally planned to work on code,

3:42:13 but they found, "Okay, this is a nice niche,

3:42:16 and now we are comfortable and we push

3:42:18 on this niche." I can see the same thing...

3:42:20 Let's say hypothetically, I'm not sure if it will be true,

3:42:24 but let's say Google takes all the market share of the general chatbot.

3:42:27 Maybe OpenAI will then focus on some other sub-topic.

3:42:31 They have too many users to go away in the foreseeable future.

3:42:37 I think Google is always ready to say, "Hold my beer," with AI models.

3:42:41 I think the question is if the companies can support the valuations.

3:42:44 I see the AI companies being looked at in some ways like AWS, Azure,

3:42:50 and GCP are, all competing in the same space and all very successful businesses.

3:42:54 There's a chance that the API market is so unprofitable

3:42:58 that they go up and down the stack to products and hardware.

3:43:01 They have so much cash that they can build power plants and data centers,

3:43:04 which is a durable advantage now.

3:43:06 But there's also a reasonable outcome that these APIs are so

3:43:10 valuable and so flexible for developers that they become something like AWS.

3:43:15 But AWS and Azure are also going to have these APIs,

3:43:20 so having five or six people competing in the API market is hard.

3:43:24 So maybe that's why they get squeezed out.

3:43:27 You mentioned "RIP Llama." Is there a path to winning for Meta?

3:43:32 I think nobody knows.

3:43:34 They're moving a lot, so they're signing licensing deals with Black Forest Labs,

3:43:40 which is an image generation company, or Midjourney.

3:43:43 So I think in some ways on the product and consumer-facing AI front,

3:43:49 it's too early to tell.

3:43:51 I think they have some people who are

3:43:53 excellent and very motivated being close to Zuckerberg.

3:43:56 So I think there's still a story to unfold there.

3:44:00 Llama is a bit different,

3:44:02 where Llama was the most focused expression of the organization.

3:44:06 And I don't see Llama being supported to that extent.

3:44:09 I think it was a very successful brand for them.

3:44:13 So they still might participate in the open ecosystem

3:44:16 or continue the Llama brand into a different service,

3:44:19 because people know what Llama is.

3:44:21 You think there's a Llama 5?

3:44:24 Not an open-weight one.

3:44:27 It's interesting.

3:44:28 I think Llama was the pioneering open-weight model.

3:44:32 With Llama 1, 2, and 3, there was a lot of love.

3:44:37 But I think then, hypothesizing or speculating,

3:44:40 I think the leaders at Meta, like the upper executives, they...

3:44:44 I think they got very excited about Llama because

3:44:47 they saw how popular it was in the community.

3:44:49 And then I think the problem was trying to, let's say,

3:44:53 monetize the open—or not monetize the open source,

3:44:55 but use it to make a bigger splash.

3:44:58 It felt almost forced,

3:45:01 like developing these very big Llama 4 models to be on top of the benchmarks.

3:45:08 But I don't think the goal of Llama models

3:45:10 is to be on top of the benchmarks beating, let's say, ChatGPT or other models.

3:45:14 I think the goal was to have a model that people can use,

3:45:18 trust, modify, and understand.

3:45:20 So that includes having smaller models.

3:45:21 They don't have to be the best models.

3:45:23 And what happened was, these models were, of course...

3:45:27 the benchmarks suggested that they were better than they were because they

3:45:31 had specific models trained on preferences

3:45:33 so that they performed well on benchmarks.

3:45:35 That's kind of, like, this overfitting thing to force it to be the best.

3:45:38 But then at the same time,

3:45:40 they didn't do the small models that people could use.

3:45:42 And I think that no one could run these big models then.

3:45:45 And then there was kind of a weird thing.

3:45:47 I think it's just because people got

3:45:49 too excited about headlines pushing the frontier.

3:45:52 I think that's it.

3:45:54 And too much on the benchmarking side.

3:45:56 It's too much work.

3:45:57 I think it imploded under internal political fighting and misaligned incentives.

3:46:03 The researchers want to build the best models,

3:46:06 but there's a layer of organization— ...and management

3:46:08 that is trying to demonstrate that they do these things.

3:46:11 And then there are rumors about how,

3:46:14 for example, some horrible technical decision was made.

3:46:19 It just seems like it got so bad that it all just crashed out.

3:46:25 Yeah, but we should also give huge props to Mark Zuckerberg.

3:46:28 I think it comes from Mark, actually, from Mark Zuckerberg,

3:46:32 from the top of the leadership, saying open source is important.

3:46:35 The fact that that leadership exists means there could be a Llama 5,

3:46:41 where they learn the lessons from benchmarking and say,

3:46:44 "We're going to be GPT-OSS—" "...and provide a really awesome library of open

3:46:51 source."- What people say is that there's

3:46:53 a debate between Mark and Alexandr Wang,

3:46:56 who is very bright, but much more against open source.

3:46:59 And to the extent that he has a lot of influence over the AI org,

3:47:02 it seems much less likely, because it seems like Mark brought him

3:47:05 in for a fresh leadership eye in directing AI.

3:47:10 And if being open or closed is no longer the defining nature of the model,

3:47:14 I don't expect that to be a defining argument between Mark and Alex.

3:47:19 They're both very bright, but I just have a hard time understanding all

3:47:23 of it because Mark wrote this piece in July of 2024,

3:47:28 which was probably the best blog post at the time,

3:47:33 saying "The Case for Open Source AI." And then July 2025 came around and it was,

3:47:38 "We're reevaluating our relationship with open source." So it's just kind of...

3:47:43 But I think also the problem...

3:47:44 Not the problem, but I think, well,

3:47:46 we may have been a bit too harsh, and that caused some of that.

3:47:50 Because I mean, we as open source developers or the community...

3:47:54 Even though the model was maybe not what

3:47:57 everyone hoped for, it got a lot of backlash.

3:48:00 And I think that was unfortunate because I can see that as a company,

3:48:04 they were hoping for positive headlines.

3:48:06 And instead of just getting no headlines or positive headlines,

3:48:11 in turn they got negative headlines.

3:48:13 And then it kind of reflected bad on the company.

3:48:17 I think that is also something where

3:48:19 it's maybe a spite reaction, almost like, "Okay, we tried to do something nice,

3:48:24 we tried to give you something cool, like an open source model,

3:48:27 and now you are kind of being negative about us,

3:48:31 even for the company." So in that sense,

3:48:34 it looks like, "Well, maybe then we'll change our mind." I guess.

3:48:37 I don't know.

3:48:39 Yeah, that's where the dynamics of discourse on X can lead us,

3:48:46 as a community, astray.

3:48:48 Because sometimes it feels random.

3:48:49 People pick the thing they like and don't like.

3:48:52 I mean, you can see the same thing with Grok 4.1 and Grok Code Fast 1.0.

3:48:59 I don't think, vibe-wise, people love it publicly.

3:49:04 But a lot of people use it.

3:49:09 So if you look to Reddit and X, they don't really give it praise

3:49:13 from the programming community, but they use it.

3:49:17 And the same thing with probably Llama.

3:49:19 I don't understand the dynamics of either positive hype or negative hype.

3:49:23 I don't understand it.

3:49:25 I mean, one of the stories of 2025 is the US filling the gap of Llama,

3:49:29 which is the rise of these Chinese open-weight models,

3:49:33 models- to the point where that was the single

3:49:35 issue I've spent a lot of energy on lately,

3:49:37 trying to do policy work to get the US to invest in this.

3:49:42 So just tell me the story of ADAM.

3:49:43 The ADAM Project started as me calling it the American DeepSeek Project,

3:49:47 which doesn't really work for DC audiences,

3:49:49 but it's the story of the most impactful thing I can do with my career,

3:49:54 which is that these Chinese open-weight models are cultivating a lot of power,

3:49:58 and there is a lot of demand for building on these open models,

3:50:02 especially in enterprises in the US that are very cagey about Chinese models.

3:50:06 The ADAM Project, American Truly Open Models,

3:50:10 is a US-based initiative to build and host high-quality,

3:50:14 genuinely open-weight AI models and supporting

3:50:16 infrastructure explicitly aimed at competing

3:50:19 with and catching up to China's rapidly advancing open-source AI ecosystem.

3:50:25 I think the one-sentence summary would be that...

3:50:28 or two sentences.

3:50:29 One is a proposition that open models are going to be

3:50:32 an engine for AI research because that is what people start with; therefore,

3:50:36 it's important to own them.

3:50:37 And the second one is, therefore, the US should be building the best models

3:50:42 so that the best research happens in the US,

3:50:45 and those US companies take the value from being

3:50:48 the home of where AI research is happening.

3:50:51 And without more investment in open models—we have

3:50:54 plots on the website where it's like, "Qwen, Qwen,

3:50:57 Qwen, Qwen"—it's all these models that are excellent

3:51:01 from these Chinese companies that are cultivating influence internationally.

3:51:05 I think the US is spending way more on AI,

3:51:10 and the ability to create open models that are a generation

3:51:13 beyond what the cutting edge of closed labs costs roughly $100 million,

3:51:18 which is a lot of money, but not a lot of money to these companies.

3:51:22 Therefore, we need a centralizing force of people who want to do this.

3:51:26 And I think we got engagement from people pretty much across the full stack,

3:51:32 whether it's policy.

3:51:34 So there has been support from the administration?

3:51:37 I don't think anyone technically in government has signed it publicly,

3:51:41 but I know people that have worked in AI policy,

3:51:45 in both the Biden and Trump administrations,

3:51:47 are very supportive of promoting open-source models in the US.

3:51:50 I think, for example,

3:51:52 AI2 got a grant from the NSF for $100 million over four years,

3:51:56 which is the biggest CS grant the NSF has ever awarded,

3:52:01 and it's for AI2 to attempt this.

3:52:04 It's a starting point.

3:52:05 But the best thing happens when

3:52:07 there are multiple organizations building models,

3:52:09 because they can cross-pollinate ideas and build this ecosystem.

3:52:13 I don't think it works if it's just Llama releasing models,

3:52:17 because Llama could go away.

3:52:19 The same thing applies for AI2; I can't be the only one building models.

3:52:25 It becomes a lot of time spent on talking to people, whether in policy...

3:52:31 I know NVIDIA is very excited about this.

3:52:34 I think Jensen Huang has been talking about the urgency

3:52:37 for this, and they've done a lot more in 2025,

3:52:40 where the Nemotron models are more of a focus.

3:52:43 They've started releasing some data along with NVIDIA's open models,

3:52:47 and very few companies do this, especially of NVIDIA's size,

3:52:51 so there are signs of progress.

3:52:54 We hear about Reflection AI, where they say their two billion dollar

3:52:58 fundraise is dedicated to building US open models,

3:53:00 and I feel and their announcement tweet reads like a blog post, right?

3:53:06 I think that cultural tide is starting to turn.

3:53:10 In July, four or five DeepSeek-caliber Chinese

3:53:14 open-weight models and and zero from the US.

3:53:18 That's the moment where I realized, like, "Oh,

3:53:20 I guess I have to spend energy on this because nobody else

3:53:23 is gonna do it." So it takes a lot of people contributing together,

3:53:26 and I don't say that, the Adam Project

3:53:28 isn't the thing that's helping to move the ecosystem,

3:53:31 but it's people like me doing this sort of thing to get the word out.

3:53:36 Do you like the 2025 America's AI Action Plan?

3:53:39 That includes open source stuff.

3:53:40 The White House AI Action Plan includes

3:53:43 a dedicated section titled "Encourage Open-Source and Open-Weight

3:53:47 AI," defining such models and arguing they

3:53:49 have unique value for innovation and startups.

3:53:52 Yeah.

3:53:53 I mean, the AI Action Plan is a plan, but largely,

3:53:56 I think it's maybe the most coherent policy

3:54:00 document that has come out of the administration,

3:54:02 and I hope that it largely succeeds.

3:54:05 I know people that have worked on the AI Action

3:54:07 Plan and the challenges of taking policy and making it real.

3:54:10 I have no idea how to do this as an AI researcher,

3:54:13 but largely a lot of things in that were very real,

3:54:16 and there's a huge build-out of AI in the country.

3:54:19 There are a lot of issues that people are hearing about,

3:54:22 from water use to whatever,

3:54:23 and we should be able to build things in this country,

3:54:26 but also, we need to not ruin places

3:54:29 in our country in the process of building it,

3:54:32 and it's worthwhile to spend energy on.

3:54:34 I think that's a role the federal government plays.

3:54:37 They set the agenda.

3:54:38 And with AI, setting the agenda

3:54:40 that open-weight should be a first consideration is

3:54:44 a large part of what they can do and then people think about it.

3:54:49 Also, for education and talent for these companies,

3:54:52 it's very important because otherwise, if there are only closed models,

3:54:56 how do you get the next generation of people contributing at some point?

3:55:01 Because otherwise, you will point only be

3:55:04 able to learn after you joined a company.

3:55:07 But at that point, how do you hire talented people?

3:55:11 How do you identify talented people?

3:55:13 I think open source is essential for a lot of things,

3:55:16 but also even just for educating the population

3:55:19 and training the next generation of researchers.

3:55:21 It's the way, or the only way.

3:55:24 The way that I could've gotten this to go more viral was

3:55:27 to tell a story of Chinese AI integrating with an authoritarian state,

3:55:31 being ASI and taking over the world,

3:55:33 and therefore we need our own American models.

3:55:35 But it's very intentional why I talk about innovation and science

3:55:38 in the US because I think it's both more realistic as an outcome,

3:55:42 but also it's a world that I would like to manifest.

3:55:48 I would say, though, also even any open-weight model,

3:55:52 I do think, is a valuable model.

3:55:55 Yeah.

3:55:55 And my argument is that we should be in a leading position.

3:55:58 But I think it's worth saying it simply because there are still voices in the AI

3:56:04 ecosystem that say we should consider banning

3:56:06 the release of open models due to safety risks.

3:56:09 And I think it's worth adding that, effectively,

3:56:12 that's impossible without making the US have its own great firewall,

3:56:16 which is also known to not work

3:56:19 that well because the cost for training these models,

3:56:22 whether it's one to a hundred million dollars,

3:56:24 is attainable to a huge amount of people

3:56:28 in the world that want to have influence,

3:56:30 so these models will be trained all over the world.

3:56:33 And we want the models, especially when,

3:56:36 like, I mean, there are safety concerns,

3:56:39 but we want this information and tools to flow freely across the world

3:56:43 and into the US so that people can use them and learn from them.

3:56:46 Stopping that would be such a restructuring

3:56:49 of our internet that it seems impossible.

3:56:51 Do you think maybe in that case the big open-weight

3:56:54 models from China are actually a good thing in a sense,

3:56:57 like, for the US companies?

3:56:58 Because maybe the US companies you

3:57:00 mentioned earlier are usually one generation behind

3:57:03 in terms of what they release open source versus what they are using?

3:57:06 For example, gpt-oss might not be the cutting-edge model.

3:57:09 Gemini 3 might not be,

3:57:11 but they do that because they know this is safe to release.

3:57:13 But then when they see, these companies see,

3:57:16 for example, there is DeepSeek-V3.2, which is really awesome,

3:57:20 and it gets used and there is no backlash, there is no security risk,

3:57:24 that could then, again, encourage them to release better models.

3:57:27 Maybe that, in a sense, is a very positive thing.

3:57:30 A hundred percent.

3:57:31 These Chinese companies have set things into motion that I think

3:57:33 would potentially not have happened if they were not all releasing models.

3:57:38 So I think it was like I'm almost

3:57:41 sure that those discussions have been had by leadership.

3:57:45 Is there a possible future where the dominant

3:57:48 AI models in the world are all open source?

3:57:51 Depends on the trajectory of progress that you predict.

3:57:53 If you think saturation in progress is coming within a few years,

3:57:57 so essentially, within the time where financial support is still very good,

3:58:01 then open models will be so optimized and so

3:58:04 much cheaper to run that they'll win out.

3:58:06 This goes back to open source ideas where so many more people will be putting

3:58:10 money into optimizing the serving of these open-weight

3:58:14 common architectures that they will become standards,

3:58:17 and then you could have chips dedicated to them,

3:58:19 and it'll be way cheaper than the offerings

3:58:22 from these closed companies that are custom.

3:58:25 We should say that the AI27 report kinda predicts one of the things it

3:58:30 does from a narrative perspective is that there will be a lot of centralization.

3:58:32 As the AI systems get smarter and smarter,

3:58:36 national security concerns will arise, and you'll centralize the labs,

3:58:40 and they'll become super secretive,

3:58:42 and there'll be this whole race- ...from a military perspective of how do you...

3:58:47 between China and the US.

3:58:48 And so all of these fun conversations we're having about LLMs...

3:58:53 the generals and the soldiers will come into the room and be like, "All right.

3:58:58 We're now in the Manhattan Project stage of this whole thing."- I think in 2025,

3:59:04 '26, '27, I don't think something like that is even remotely possible.

3:59:08 You can make the same argument for computers, right?

3:59:11 You can say, "Computers are capable and we don't

3:59:14 want the general public to get them." Or chips,

3:59:17 even AI chips, but you see how Huawei makes chips now.

3:59:22 It took a few years, but...

3:59:24 and I don't think there is a way you can contain knowledge like that.

3:59:29 I think in this day and age, it is impossible, like the internet.

3:59:34 I don't think this is a possibility.

3:59:38 On the Manhattan Project thing, I think that a Manhattan Project-like thing

3:59:42 for open models would be pretty reasonable, because it wouldn't cost that much.

3:59:46 But I think that will come.

3:59:48 It seems like culturally, the companies are changing.

3:59:51 But I agree with Sebastian on all of that.

3:59:54 I don't see it happening nor being helpful.

3:59:59 Yeah.

3:59:59 The motivating force behind the Manhattan Project was civilizational risk.

4:00:03 It's harder to motivate that for open-source models.

4:00:08 There's no civilizational risk.

4:00:10 On the hardware side, we mentioned NVIDIA a bunch of times.

4:00:15 Do you think Jensen and NVIDIA will keep winning?

4:00:19 I think they have to iterate and manufacture a lot.

4:00:22 And I think they probably...

4:00:25 what they're doing, they do innovate, but I think there's always the chance

4:00:31 that someone does something fundamentally different,

4:00:34 gets very lucky, and then does something.

4:00:37 But the problem is adoption.

4:00:39 The moat of NVIDIA is probably not just the GPU.

4:00:43 It's more like the CUDA ecosystem, and that has evolved over two decades.

4:00:47 Even back when I was a grad student,

4:00:50 I was in a lab doing biophysical simulations,

4:00:53 molecular dynamics, and we had a Tesla GPU back then just for the computations.

4:00:57 It was about 15 years ago now.

4:01:00 And they built this up for a long time, and that's the moat, I think.

4:01:05 It's not the chip itself,

4:01:07 although they have the money to iterate, build, and scale.

4:01:11 But then it's really about compatibility.

4:01:14 If you're at that scale, why would you go with something risky where there

4:01:19 are only a few chips they can make per year?

4:01:21 You go with the big one.

4:01:23 But then I do think with LLMs now,

4:01:25 it will be easier to design something like CUDA.

4:01:30 It took 15 years because it was hard,

4:01:32 but now that we have LLMs, we can maybe replicate CUDA.

4:01:36 And I wonder if there will be a separation

4:01:38 of training and inference compute as we stabilize,

4:01:42 and more compute is needed for inference.

4:01:47 That's supposed to be the point of the Groq acquisition.

4:01:50 And that's why part of what Vera Rubin is-

4:01:52 where they have a new chip with no high-bandwidth memory,

4:01:54 which is one of the- or very little, which is one of the most expensive pieces.

4:01:59 It's designed for pre-fill, which is the part of inference where

4:02:03 you essentially do a lot of matrix multiplications.

4:02:05 And then you only need the memory

4:02:07 when you're doing this autoregressive generation,

4:02:09 and you have the KV cache swaps.

4:02:11 So they have this new GPU that's designed for that specific use case,

4:02:15 and then the cost of ownership per FLOP or whatever is actually way lower.

4:02:20 But I think that NVIDIA's fate lies in the diffusion of AI still.

4:02:25 Their biggest clients are still these hyperscale companies.

4:02:29 Like, Google obviously can make TPUs.

4:02:32 Amazon is making Trainium.

4:02:34 Microsoft will try to do its own things.

4:02:37 And so long as the pace of AI progress is high,

4:02:40 NVIDIA's platform is the most flexible and people will want that.

4:02:43 But if there's stagnation, then creating bespoke chips,

4:02:47 there's more time to do it.

4:02:50 It's interesting that NVIDIA is quite active

4:02:53 in trying to develop all kinds of different products.

4:02:56 They try to create areas of commercial value that will use a lot of GPUs.

4:03:01 Mm-hmm.

4:03:02 But they keep innovating and they're doing a lot of incredible research, so...

4:03:07 Everyone says the company's super oriented around

4:03:09 Jensen and how operationally plugged in he is.

4:03:12 And it sounds so unlike many other big companies that I've heard about.

4:03:16 And so long as that's the culture,

4:03:18 I think that we can expect that to keep progress happening.

4:03:21 And it's like he's still in the Steve Jobs era of Apple.

4:03:24 So long as that is how it operates,

4:03:27 I'm pretty optimistic for their situation because it's like,

4:03:32 it is their top-order problem,

4:03:33 and I don't know if making these chips for the whole

4:03:37 ecosystem is the top goal of all these other companies.

4:03:39 They'll do a good job, but it might not be as good of a job.

4:03:43 Since you mentioned Jensen,

4:03:45 I've been reading a lot about history and about singular figures in history.

4:03:49 What do you guys think about the single man/woman view of history?

4:03:53 How important are individuals for steering

4:03:55 the direction of history in the tech sector?

4:03:58 So, you know, what's NVIDIA without Jensen?

4:04:01 You mentioned Steve Jobs.

4:04:03 What's Apple without Steve Jobs?

4:04:05 What's xAI without Elon or DeepMind without Demis?

4:04:12 People make things earlier and faster, whereas scientifically,

4:04:17 many great scientists credit being in the right place

4:04:19 at the right time and still making the innovation,

4:04:22 where eventually someone else will still have the idea.

4:04:25 So I think that in that way,

4:04:29 Jensen is helping manifest this GPU revolution much faster and much

4:04:34 more focused than it would happen without having a person there.

4:04:38 And this is making the whole AI build-out faster.

4:04:40 But I do still think that eventually, something like ChatGPT would have happened

4:04:45 and a build-out like this would have happened,

4:04:47 but it probably would not have been as fast.

4:04:50 I think that's the sort of flavor that is applied.

4:04:55 These individual people, there are people who are placing bets on something.

4:04:58 Some get lucky, some don't.

4:04:59 But if you don't have these people at the helm, it would be more diffused.

4:05:02 It's almost like investing in an ETF versus individual stocks.

4:05:06 Individual stocks might go up or down more heavily than an ETF,

4:05:11 which is more balanced.

4:05:12 It will eventually go up over time.

4:05:13 We'll get there.

4:05:14 But it's just like, you know, the focus I think is the thing.

4:05:18 Passion and focus.

4:05:20 Isn't there a real case to be made that without Jensen,

4:05:22 there's not a reinvigoration of the deep learning revolution?

4:05:27 It could've been 20 years later, is what I would say.

4:05:30 Or like another AI winter could have come if GPUs weren't around.

4:05:35 That could change history completely because you could think

4:05:37 of all the other technologies that could've come in the meantime,

4:05:42 and the focus of human civilization would get...

4:05:44 Silicon Valley would be captured by different hype.

4:05:48 But I do think there's certainly an aspect

4:05:50 where it was all planned, the GPU trajectory.

4:05:53 But on the other hand, it's also a lot of lucky coincidences or good intuition.

4:05:58 Like the investment into, let's say, biophysical simulations.

4:06:01 I mean, I think it started with video games and then it just happened

4:06:05 to be good at linear algebra because

4:06:07 video games require a lot of linear algebra.

4:06:09 And then you have the biophysical simulations.

4:06:11 But still, I don't think the master plan was AI.

4:06:16 I think it happened to be Alex Krizhevsky.

4:06:19 So someone took these GPUs and said, "Hey,

4:06:22 let's try to train a neural network on that." It happened to work really well,

4:06:26 and I think it only happened because you could purchase those GPUs.

4:06:30 Gaming would've created a demand for faster processors if

4:06:33 NVIDIA had gone out of business in the early days.

4:06:37 That's what I would think.

4:06:37 I think that the GPUs would've been different,

4:06:42 but I think GPUs would still exist at the time

4:06:46 of AlexNet and at the time of the Transformer.

4:06:48 It was just hard to know if it would be

4:06:51 one company as successful or multiple smaller companies with worse chips.

4:06:55 But I don't think that's a 100-year delay.

4:06:59 It might be a decade delay.

4:07:01 Well, it could be one, two, three, four, five-decade delay.

4:07:04 I just can't see Intel or AMD doing what NVIDIA did.

4:07:08 I don't think it would be a company that exists.

4:07:11 I think it would be a different company that would rise.

4:07:13 Like Silicon Graphics or something.

4:07:15 So yeah, some company that has died would have done it.

4:07:19 But just looking at it, it seems like these singular figures,

4:07:23 these leaders, have a huge impact on the trajectory of the world.

4:07:28 Obviously, there are incredible teams behind them.

4:07:31 But, you know, having that kind of very singular,

4:07:36 almost dogmatic focus- -is necessary to make progress.

4:07:41 Yeah, I mean, even with GPT, it wouldn't exist if there wasn't a person,

4:07:44 Ilya, who pushed for this scaling, right?

4:07:47 Yeah, Dario Amodei was also deeply involved in that.

4:07:50 If you read some of the histories from OpenAI,

4:07:52 it seems wild thinking about how early these people were like,

4:07:55 "We need to hook up 10,000 GPUs and take all of OpenAI's compute and train

4:07:58 one model." There were a lot of people who didn't want to do that.

4:08:02 Which is an insane thing to believe.

4:08:05 To believe in scaling before scaling has

4:08:07 any indication that it's going to materialize.

4:08:10 Again, singular figures.

4:08:12 Speaking of which, 100 years from now,

4:08:16 this is presumably post-singularity, whatever singularity is.

4:08:21 When historians look back at our time now,

4:08:24 what technological breakthroughs would they really emphasize

4:08:28 as the breakthroughs that led to the singularity?

4:08:32 So far we have Turing to today, 80 years.

4:08:37 I think it would still be computing,

4:08:39 like the umbrella term "computing." I don't necessarily think

4:08:42 that in 100 or 200 years it would be AI.

4:08:46 It could still very well be computers.

4:08:48 We are now taking better advantage of them, but the fact of computing remains.

4:08:54 It's basically a Moore's Law discussion.

4:08:56 Even the details of CUDA and GPUs won't even be remembered,

4:09:00 nor will all this software turmoil.

4:09:04 It'll just be, obviously, compute.

4:09:07 I generally agree, but is the connectivity

4:09:10 of the internet and compute able to be merged?

4:09:14 Or is it both of them?

4:09:18 I think the internet will probably be related to communication.

4:09:21 It could be a phone, the internet, or satellites.

4:09:25 Compute is more like the scaling aspect of it.

4:09:29 It's possible that the internet is completely forgotten-

4:09:32 -that the internet is wrapped into phone networks, like communication networks.

4:09:38 This is just another manifestation

4:09:40 of that, and the real breakthrough comes from increased compute,

4:09:44 or Moore's Law, broadly defined.

4:09:46 Well, I think the connection of people is very fundamental to it.

4:09:50 it's like, you can talk to anyone.

4:09:52 You want to find the best person in the world for something,

4:09:56 they are somewhere in the world.

4:09:57 And being able to have that flow of information—the AIs will also rely on this.

4:10:02 I've been fixating on when I said

4:10:04 the dream was dead about the one central model.

4:10:07 The thing that is evolving is people having many agents for different tasks.

4:10:11 People already started doing this with different clouds.

4:10:14 It's described as many AGIs in the data center

4:10:18 where each one manages and they talk to each other.

4:10:21 And that is reliant on networking and the free flow of information.

4:10:26 on top of compute.

4:10:27 But networking, especially with GPUs, is such a part of scaling of compute.

4:10:33 The GPUs and the data centers need to talk to each other.

4:10:36 Anything about neural networks will be remembered?

4:10:39 Like, do you think there's something very specific and singular

4:10:42 to the fact that it's neural networks that's seen as a breakthrough,

4:10:46 like a genius, that you're basically replicating,

4:10:48 in a very crude way, the human mind?

4:10:51 The structure of the human brain, the human mind?

4:10:54 I think without the human mind, we probably wouldn't have neural networks,

4:10:58 because it just was an inspiration for that.

4:11:01 But on the other end, I think it's just so, so different.

4:11:04 I mean, it's digital versus biological,

4:11:06 that I do think it will probably be more grouped as an algorithm.

4:11:12 That's massively parallelizable...

4:11:13 ...On this particular kind of compute?

4:11:15 It could have been like genetic computing; genetic algorithms just parallelized.

4:11:19 It just happens that this is more efficient and works better.

4:11:23 And it very well could be that the LLM, the neural networks,

4:11:26 the way we architect them now is just

4:11:29 a small component of the system that leads to singularity.

4:11:34 If you think of it in 100 years, I think society can be changed more

4:11:38 with more compute and intelligence because of autonomy.

4:11:41 But looking at this, what are the things

4:11:45 from the Industrial Revolution that we remember?

4:11:47 We remember the engine,

4:11:48 which is probably the equivalent of the computer in this.

4:11:51 But there's a lot of other physical transformations that people

4:11:55 are aware of, like the cotton gin and all these things,

4:11:59 these machines that are still known: air conditioning, refrigerators.

4:12:04 Some of these things from AI will still be known.

4:12:08 The word "transformer" could still be known.

4:12:11 I would guess that deep learning is definitely still known,

4:12:14 but the transformer might be evolved away

4:12:16 from in 100 years with AGI researchers everywhere.

4:12:21 But I think deep learning is likely to be a term that is remembered.

4:12:28 And I wonder what the air conditioning

4:12:30 and refrigeration of the future is that AI brings.

4:12:32 If we travel forward 100 years from now,

4:12:34 we transport there right now, what do you think is different?

4:12:37 How do you think the world looks different?

4:12:40 First of all, do you think there are humans?

4:12:42 Do you think there are robots everywhere walking around?

4:12:46 I do think specialized robots, for sure, for certain tasks.

4:12:49 Humanoid form?

4:12:51 Maybe half-humanoid.

4:12:52 We'll see.

4:12:54 I think for certain things, yes,

4:12:55 there will be humanoid robots because it's just amenable for the environment.

4:13:00 But for certain tasks, it might make sense.

4:13:03 What's harder to imagine is how we interact

4:13:05 with the devices and what humans do with devices.

4:13:08 Well, I mean, I'm pretty sure it will

4:13:11 probably not be the cellphone or the laptop.

4:13:13 Will it be implants?

4:13:16 I mean, it has to be brain-computer interfaces, right?

4:13:18 I mean, 100 years from now,

4:13:20 given the progress we're seeing now— there has to be...

4:13:25 unless there's legitimately a complete alteration

4:13:29 of how we interact with reality.

4:13:33 On the other hand, cars are older than 100 years, right?

4:13:36 And it's still the same interface.

4:13:38 We haven't replaced cars with something else.

4:13:41 We just made them better,

4:13:42 but it's still a steering wheel, still wheels, you know?

4:13:45 I think we'll still carry around a physical brick

4:13:47 of compute because people want some ability to have a private...

4:13:51 Like, you might not engage with it as much as a phone,

4:13:55 but having private information that is yours

4:13:56 as an interface between the rest of the internet, I think that will still exist.

4:14:01 It might not look like an iPhone and it might be used a lot less,

4:14:05 but I still expect people to carry things around.

4:14:08 Why do you think the smartphone is the embodiment of private?

4:14:11 There's a camera on it.

4:14:14 There's—- Private for you, like encrypted messages, encrypted photos...

4:14:19 know what your life is.

4:14:22 I guess it's a question of how optimistic on brain-machine interfaces you are.

4:14:26 Is all that just going to be stored in the cloud?

4:14:29 Your whole calendar?

4:14:30 It's hard to think about processing all the information that we can process

4:14:37 visually through brain-machine interfaces presenting something

4:14:40 like a calendar or something to you.

4:14:44 It's hard to just think about knowing, without looking, your email inbox.

4:14:49 Like you signal to a computer and then you just know your email inbox.

4:14:53 Is that something that the human brain

4:14:55 can handle being piped into it non-visually?

4:14:58 I don't know exactly how those transformations happen.

4:15:03 Humans aren't changing in 100 years.

4:15:05 I think agency and community are things that people actually want.

4:15:09 A local community, yeah.

4:15:10 People you are close to, being able to do things with them and being

4:15:15 able to ascribe meaning to your life and being able to do things.

4:15:22 In 100 years, I don't think that human biology is changing

4:15:26 away from those on a time scale that we can discuss.

4:15:30 And I think that UBI does not solve agency.

4:15:34 I do expect mass wealth, and I hope that it has spread so

4:15:38 that the average life looks very different in 100 years.

4:15:42 But that's still a lot to happen.

4:15:44 If you think about countries that are early

4:15:47 in their development process to getting access to computing and internet,

4:15:52 to build all the infrastructure and have policy

4:15:57 that shares one nation's wealth with another is...

4:16:01 I think it's an optimistic view to see all that happening in 100 years- ...while

4:16:06 they are still independent entities and not just

4:16:10 like absorbed into some international order by force.

4:16:13 But there could be just better, more elaborate, more effective...

4:16:18 social support systems that help alleviate some

4:16:22 levels of basic suffering from the world.

4:16:24 You know, the transformation of society where a lot

4:16:26 of jobs are lost in the short term, I think we have to really remember that each

4:16:31 individual job that's lost is a human being who's suffering.

4:16:35 That's like a...

4:16:38 When jobs are lost, the scale is a real tragedy.

4:16:41 You can make all kinds of arguments about

4:16:43 economics or how it's all going to be okay.

4:16:46 It's good for the GDP, there's going to be new jobs created.

4:16:50 Fundamentally at the individual level for that human being,

4:16:54 that's real suffering.

4:16:55 That's a real personal sort of tragedy.

4:16:58 And we have to not forget that as the technologies are being developed.

4:17:03 And also my hope for all the AI slop we're seeing is that there will be

4:17:10 a greater and greater premium for the fundamental

4:17:14 aspects of the human experience that are in-person.

4:17:17 The things that we all...

4:17:19 Like seeing each other, talking together in-person.

4:17:23 The next few years are definitely going to be an increased

4:17:26 value on physical goods and events— ...and even more pressure on slop.

4:17:32 So it'll be...

4:17:33 the slop is only starting.

4:17:35 The next few years will be more and more diverse ...versions of slop.

4:17:38 They would be drowning in slop.

4:17:39 Is that what—- So I'm hoping that society drowns

4:17:42 in slop enough to snap out of it and be like, "We can't deal with it.

4:17:47 It just doesn't matter." And then, the physical has such a higher premium on it.

4:17:54 Even like classic examples, I honestly think this is true,

4:17:57 and I think we will get tired of it.

4:17:59 We are already kind of tired of it.

4:18:01 I mean, even art.

4:18:02 I don't think art will go away.

4:18:04 You have paintings, physical paintings.

4:18:06 There's more value, not just monetary value,

4:18:10 but just more value appreciation for the actual

4:18:13 painting than a photocopy of that painting.

4:18:14 It could be a perfect digital reprint,

4:18:16 but there is something when you go to a museum and you look

4:18:19 at that art and you see the real thing and you just think, "Okay.

4:18:22 A human." It's like a craft.

4:18:24 You have like an appreciation for that.

4:18:26 And I think the same is true for writing,

4:18:27 for talking, for any type of experience...

4:18:32 I do unfortunately think it will be like a dichotomy,

4:18:36 like a fork where some things will be automated.

4:18:40 Like, you know, there are not as many paintings as there used to be,

4:18:42 you know, 200 years ago.

4:18:43 There are more photographs, more photocopies.

4:18:46 But at the same time, it won't go away.

4:18:49 There will be value in that.

4:18:51 I think the difference will just be, you know, what's the proportion of that.

4:18:56 But personally, I have a hard time reading

4:18:59 things where I obviously see it's obviously AI generated.

4:19:02 I'm sorry.

4:19:03 It might—it might be really good information there, but I'm just like, "Nah,

4:19:08 not for me."- I think eventually they'll fool you,

4:19:10 and it'll be on platforms that give ways of verifying or building trust.

4:19:15 So you will trust that Lex is not AI generated, having been here.

4:19:19 So then you have trust in this channel.

4:19:22 But it's harder for new people who don't have that trust.

4:19:25 Well, that will get interesting because I think fundamentally it's a solvable

4:19:30 problem by having trust in certain outlets that they won't do it,

4:19:35 but it's all going to be trust-based.

4:19:37 There will be systems to authorize, "Okay, this is real.

4:19:39 This is not real." There will be some telltale signs where

4:19:42 you can obviously tell this is AI generated and this is not.

4:19:45 But some will be so good that it's hard to tell, and then you have to trust.

4:19:50 And well, that will get interesting and a bit problematic.

4:19:54 The extreme case of this is to watermark all human content.

4:19:57 So all photos that we take on our own have

4:20:00 some watermark until they are edited or something like this.

4:20:03 And software can manage communications with the device

4:20:07 manufacturer- device manufacturer to maintain human

4:20:10 editing— which is the opposite of the discussion to try to watermark AI images.

4:20:15 And then you can make a Google image that has

4:20:18 a watermark and use a different Google tool to remove it.

4:20:21 Yep.

4:20:21 It's going to be an arms race, basically.

4:20:23 And we've been mostly focusing on the positive aspects of AI.

4:20:27 All the capabilities that we've been talking about can be used

4:20:32 to destabilize human civilization with even

4:20:34 just relatively dumb AI applied at scale,

4:20:39 and then further, superintelligent AI systems.

4:20:42 Of course, there's the sort of doomer take

4:20:45 that's important to consider as we develop these technologies.

4:20:50 What gives you hope about the future of human civilization,

4:20:53 given everything we've been talking about?

4:20:56 Are we going to be okay?

4:20:59 I think we will.

4:21:00 I'm definitely a worrier, both about AI and non-AI things.

4:21:04 But humans do tend to find a way.

4:21:08 I think that's what humans are built for: to have

4:21:11 community and find a way to figure out problems.

4:21:14 That's what has gotten us to this point.

4:21:16 And to think that the AI opportunity and related technologies is really big.

4:21:23 And I think that there's big social

4:21:25 and political problems to help everybody understand that.

4:21:29 And I think that's what we're staring at a lot of right now,

4:21:33 is like the world is a scary place, and AI is a very uncertain thing.

4:21:36 And it takes a lot of work that is not necessarily building things.

4:21:41 It's like telling people and understanding people,

4:21:44 that the people building AI are historically not motivated or wanting to do.

4:21:50 But it is something that is probably doable.

4:21:52 It just will take longer than people want.

4:21:55 And we have to go through that long period of like hard,

4:22:00 distraught AI discussions if we want to have the lasting benefits.

4:22:05 Yeah.

4:22:05 Through that process,

4:22:06 I'm especially excited that we get a chance to better understand ourselves,

4:22:12 us at the individual level as humans and at the civilization level,

4:22:16 and answer some of the big mysteries,

4:22:18 like what is this whole consciousness thing going on here?

4:22:23 It seems to be truly special.

4:22:24 Like, there's a real miracle in our mind.

4:22:27 And AI puts a mirror to ourselves and we

4:22:29 get to answer some of the big questions about like,

4:22:33 what is this whole thing going on here?

4:22:35 Well, one thing about that is also what I do think makes us

4:22:39 very different from AI and why I don't worry about AI taking over is,

4:22:44 like you said, consciousness.

4:22:45 We humans, we decide what we want to do.

4:22:47 AI in its current implementation, I can't see it changing.

4:22:51 You have to tell it what to do.

4:22:54 And so you have still the agency.

4:22:56 It doesn't take the agency from you because it becomes a tool.

4:23:00 You can think of it as a tool.

4:23:01 You tell it what to do.

4:23:03 It will be more automatic than other previous tools.

4:23:06 It's certainly more powerful than a hammer,

4:23:08 it can figure things out, but it's still you in charge, right?

4:23:12 So the AI is not in charge, you're in charge.

4:23:15 You tell the AI what to do and it's doing it for you.

4:23:18 So in the post-singularity, post-apocalyptic war between humans and machines,

4:23:22 you're saying humans are worth fighting for?

4:23:27 100%.

4:23:27 I mean, this is...

4:23:29 The movie Terminator, they made in the '80s, essentially, and I do think,

4:23:33 well, the only thing I can see going wrong is,

4:23:38 of course, if things are explicitly programmed

4:23:40 to do the thing that is harmful, basically.

4:23:43 I think actually in that, in a Terminator type of setup, I think humans win.

4:23:49 I think we're too clever.

4:23:51 It's hard to explain how we figure it out, but we do.

4:23:56 And we'll probably be using local LLMs,

Study with Looplines Download Captions Watch on YouTube