State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490
Lex Fridman
0:00 The following is a conversation all
0:02 about the state-of-the-art in artificial intelligence,
0:04 including some of the exciting technical breakthroughs and developments
0:08 in AI that happened over the past year,
0:11 and some of the interesting things we think might happen this upcoming year.
0:16 At times, it does get super technical,
0:19 but we do try to make sure that it remains
0:22 accessible to folks outside the field without ever dumbing it down.
0:26 It is a great honor and pleasure to be able to do
0:30 this kind of episode with two of my favorite people in the AI community,
0:35 Sebastian Raschka and Nathan Lambert.
0:38 They are both widely respected machine learning researchers
0:42 and engineers who also happen to be great communicators,
0:46 educators, writers, and X posters.
0:49 Sebastian is the author of two books
0:53 I highly recommend for beginners and experts alike.
0:56 First is Build a Large Language Model
0:59 from Scratch and Build a Reasoning Model from Scratch.
1:04 I truly believe in the machine learning world,
1:08 the best way to learn and understand
1:11 something is to build it yourself from scratch.
1:15 Nathan is the post-training lead at the Allen Institute for AI,
1:21 author of the definitive book on Reinforcement Learning from Human Feedback.
1:26 Both of them have great X accounts, great Substacks.
1:30 Sebastian has courses on YouTube, Nathan has a podcast.
1:34 And everyone should absolutely follow all of those.
1:37 those.
1:38 This is the Lex Fridman podcast.
1:40 To support it, please check out our sponsors in the description,
1:43 where you can also find links to contact me,
1:47 ask questions, get feedback, and so on.
1:50 And now, dear friends, here's Sebastian Raschka and Nathan Lambert.
1:57 So I think one useful lens to look
1:59 at all this through is the so-called DeepSeek moment.
2:03 This happened about a year ago in January 2025,
2:07 when the open-weight Chinese company DeepSeek released DeepSeek R1,
2:11 that I think it's fair
2:14 to say surprised everyone with near-state-of-the-art performance,
2:17 with allegedly much less compute for much cheaper.
2:22 And from then to today, the AI competition has gotten insane,
2:28 both on the research and product level.
2:31 It's just been accelerating.
2:32 discuss all of this today,
2:34 and maybe let's start with some spicy questions if we can.
2:38 Who's winning at the international level?
2:40 Would you say it's the set of companies in China
2:43 or the set of companies in the United States?
2:46 And Sebastian, Nathan, it's good to see you guys.
2:50 guys.
2:50 So Sebastian, who do you think is winning?
2:53 Winning is a very broad term.
2:57 I would say you mentioned the DeepSeek moment,
2:59 and I think DeepSeek is winning the hearts of the people
3:02 who work on open-weight models because they share these as open models.
3:06 Winning, I think, has multiple timescales to it.
3:09 We have today, we have next year, we have in 10 years.
3:13 One thing I know for sure is that I don't think nowadays, in 2026,
3:18 that there will be any company that has access
3:22 to technology that no other company has access to.
3:26 That is mainly because researchers are frequently changing jobs and labs.
3:32 They rotate.
3:32 I don't think there will be a clear winner in terms of technology access.
3:36 However, I do think there will be,
3:39 The differentiating factor will be budget and hardware constraints.
3:43 I don't think the ideas will be proprietary,
3:46 but rather the resources needed to implement them.
3:52 I don't see currently a winner-take-all scenario.
3:55 I can't see that.
3:57 At the moment.
3:59 Nathan, what do you think?
4:00 You see the labs put different energy into what they're trying to do,
4:04 and I think to demarcate the point in time when we're recording this, the hype
4:08 over Anthropic's Claude Opus 4.5 model
4:11 has been absolutely insane, which is just...
4:13 I mean, I've used it and built stuff in the last few weeks, and it's...
4:17 it's almost gotten to the point where it feels like
4:19 a bit of a meme in terms of the hype.
4:21 And it's kind of funny because this is very organic,
4:24 and then if we go back a few months ago,
4:26 we can see the release date and the notes,
4:29 as Gemini 3 from Google got released, and it seemed like the marketing and just,
4:34 like, wow factor of that release was super high.
4:37 But then at the end of November,
4:39 Claude Opus 4.5 was released and the hype has been growing,
4:42 but Gemini 3 was before this.
4:43 And it kind of feels like people don't really talk about it as much,
4:46 even though when it came out, everybody was like,
4:48 this is Gemini's moment to retake Google's structural advantages in AI.
4:53 And Gemini 3 is a fantastic model, and I still use it.
4:56 It's just kind of differentiation is lower.
4:59 And I agree with Sebastian;
5:01 what you're saying with all these, the idea space is very fluid,
5:05 but culturally Anthropic is known for betting very hard on code,
5:09 which is the Claude Code thing, is working out for them right now.
5:12 So I think that even if the ideas flow pretty freely,
5:15 so much of this is bottlenecked
5:16 by human effort and the culture of organizations,
5:19 where Anthropic seems to at least be presenting as the least chaotic.
5:23 It's a bit of an advantage, if they can keep doing that for a while.
5:27 But on the other side of things, there's a lot of ominous technology from China
5:31 where there's way more labs than DeepSeek.
5:34 So DeepSeek kicked off a movement within China,
5:37 I say kind of similar to how ChatGPT kicked off
5:40 a movement in the US where everything had a chatbot.
5:43 There's now tons of tech companies in China
5:46 that are releasing very strong frontier open-weight models,
5:48 to the point where I would say that DeepSeek is kind
5:51 of losing its crown as the preeminent open model maker in China,
5:54 and the likes of Z.ai with their GLM models, Minimax's models,
6:00 Kimi Moonshot, especially in the last few months, has shown more brightly.
6:04 The new DeepSeek models are still very strong, but that's kind of a...
6:08 it could look back as a big narrative point
6:10 where in 2025 DeepSeek came and it provided this platform
6:14 for way more Chinese companies that are releasing these fantastic
6:17 models to kind of have this new type of operation.
6:20 So these models from these Chinese companies are open-weights,
6:23 and depending on this trajectory of business
6:25 models that these American companies are doing, they could be at risk.
6:29 But currently, a lot of people are paying for AI software in the US,
6:33 and historically in China and other parts of the world,
6:36 people don't pay a lot for software.
6:38 So some of these models like DeepSeek have
6:40 the love of the people because they are open-weight.
6:42 How long do you think the Chinese companies keep releasing open-weight models?
6:47 I would say for a few years.
6:49 I think that, like in the US, there's not a clear business model for it.
6:53 I have been writing about open models for a while,
6:55 and these Chinese companies have realized it.
6:57 So I get inbound from some of them.
6:59 And they're smart and realize the same constraints:
7:01 a lot of top US tech companies and other IT companies
7:04 won't pay for an API subscription to Chinese companies for security concerns.
7:08 This has been a long-standing habit in tech,
7:11 and the people at these companies then see open weight models as an ability
7:16 to influence and take part of a huge growing AI expenditure market in the US.
7:20 And they're very realistic about this, and it's working for them.
7:24 I think that the government will see that that is building
7:27 a lot of influence internationally in terms of uptake of the technology,
7:31 so there's going to be a lot of incentives to keep it going.
7:34 But building these models and doing the research is very expensive,
7:37 so at some point, I expect consolidation.
7:39 But I don't expect that to be a story of 2026,
7:42 where there will be more open model
7:45 builders throughout 2026 than there were in 2025.
7:47 And a lot of the notable ones will be in China.
7:50 You were going to say something?
7:51 Yes.
7:52 You mentioned DeepSeek losing its crown.
7:54 I do think to some extent, yes, but we also have to consider though,
8:00 they are still, I would say, slightly ahead.
8:02 And the other ones—it's not that DeepSeek got worse,
8:04 it's just that the other ones are using the ideas from DeepSeek.
8:08 For example, you mentioned Kimi—same architecture, they're training it.
8:11 And then again, we have this leapfrogging where they might be at some
8:14 point in time a bit better because they have the more recent model.
8:17 And I think this comes back to the fact that there won't be a clear winner.
8:22 It will just be like that: one person releases something,
8:25 the other one comes in, and the most
8:27 recent model is probably always the best model.
8:30 Yeah.
8:30 We'll also see the Chinese companies have different incentives.
8:33 Like, DeepSeek is very secretive,
8:35 whereas some of these startups are like the MiniMaxs and Z.ais of the world.
8:40 Those two literally have filed IPO paperwork,
8:42 and they're trying to get Western mindshare and do a lot of outreach there.
8:47 So I don't know if these incentives will change the model development,
8:50 because DeepSeek famously is built by a hedge fund,
8:53 Highflyer Capital, and we don't know exactly what they
8:56 use the models for or if they care about this.
8:59 They're secretive in terms of communication;
9:00 they're not secretive in terms of the technical
9:02 reports that describe how their models work.
9:04 They're still open on that front.
9:05 And we should also say, on the Claude Opus 4.5 hype,
9:10 there's the layer of something being the darling of the X echo chamber,
9:17 on the Twitter echo chamber,
9:18 and the actual amount of people that are using the model.
9:22 I think it's probably fair to say that ChatGPT and Gemini are focused
9:26 on the broad user base that just want to solve problems in their daily lives,
9:32 and that user base is gigantic.
9:34 So the hype about the coding may not be representative of the actual use.
9:39 I would say also a lot of the usage patterns are,
9:43 like you said, name recognition,
9:44 brand and stuff, but also muscle memory almost, where,
9:48 you know, ChatGPT has been around for a long time.
9:51 People just got used to using it, and it's almost like a flywheel:
9:54 they recommend it to other users and that stuff.
9:57 One interesting point is also the customization of LLMs.
10:00 For example, ChatGPT has a memory feature, right?
10:03 And so you may have a subscription and you use it for personal stuff,
10:07 but I don't know if you want to use that same thing at work.
10:10 Because it's a boundary between private and work.
10:12 If you're working at a company,
10:13 they might not allow that or you may not want that.
10:16 And I think that's also an interesting
10:18 point where you might have multiple subscriptions.
10:20 One is just clean code.
10:22 It has nothing of your personal images or hobby projects in there.
10:26 It's just like the work thing.
10:28 And then the other one is your personal thing.
10:30 So I think that's also something where there are two different use cases,
10:32 and it doesn't mean you only have to have one.
10:36 I think the future is also multiple ones.
10:39 What model do you think won 2025,
10:40 and what model do you think is going to win '26?
10:43 I think in the context of consumer chatbots,
10:45 it's a question of: are you willing to bet on Gemini over ChatGPT?
10:50 Which I would say, in my gut,
10:52 feels like a bit of a risky bet because OpenAI has been the incumbent,
10:55 and there are so many benefits to that in tech.
10:58 I think the momentum, if you look at 2025,
11:02 was on Gemini's side, but they were starting from such a low point.
11:07 And RIP Bard and these earlier attempts at getting started.
11:13 Huge credit to them for powering through
11:15 the organizational chaos to make that happen.
11:17 But also it's hard to bet against OpenAI
11:19 because they always come off as so chaotic,
11:23 but they're very good at landing things.
11:25 And I think, personally, I have very mixed reviews of GPT-5,
11:29 but it must have saved them so much money with the high-line feature being
11:33 a router where most users are no longer charging their GPU costs as much.
11:38 So I think it's very hard to dissociate the things that I like out
11:43 of models versus the things that are
11:45 going to actually be a general public differentiator.
11:50 What do you think about 2026?
11:51 Who's going to win?
11:53 I'll say something, even though it's risky.
11:54 I think Gemini will continue to make progress on ChatGPT.
11:56 I think Google's scale,
11:58 when both of these are operating at such extreme scales—and Google
12:02 has the ability to separate research and product a bit better,
12:06 whereas you hear so much about OpenAI
12:08 being chaotic operationally and chasing the high-impact thing,
12:11 which is a very startup culture.
12:13 And then on the software and enterprise side,
12:15 I think Anthropic will have continued success,
12:16 as they've again and again been set up for that.
12:19 And obviously Google Cloud has a lot of offerings,
12:23 but I think this kind of Gemini name brand is important for them to build.
12:27 Google Cloud will continue to do well,
12:30 but that's a more complex thing to explain in the ecosystem,
12:35 because that's competing with the likes of Azure
12:37 and AWS rather than on the model provider side.
12:41 So in infrastructure, you think TPU is giving an advantage?
12:46 Largely because the margin on NVIDIA chips is insane,
12:49 and Google can develop everything from top to bottom
12:51 to fit their stack and not have to pay this margin.
12:54 And they've had a head start in building data centers.
12:57 So all of these things that have both high
12:59 lead times and very hard margins on high costs,
13:02 Google has a just kind of historical advantage there.
13:05 And if there's going to be a new paradigm, it's most likely to come from OpenAI
13:09 where their research division again and again
13:12 has shown this ability to land a new research idea or a product.
13:17 Like Deep Research, Sora,
13:18 o1 thinking models—all these definitional things have come from OpenAI,
13:23 and that's got to be one of their top traits as an organization.
13:27 So it's kind of hard to bet against that, but I think a lot of this year
13:31 will be about scale and optimizing what
13:33 could be described as low-hanging fruit in models.
13:37 And clearly there's a trade-off between intelligence and speed.
13:41 This is what ChatGPT-5 was trying to solve behind the scenes.
13:46 It's like, do people actually want intelligence,
13:49 the broad public, or do they want speed?
13:52 I think it's a nice variety, or the option to have a toggle there.
13:56 I mean, for my personal usage, most of the time when I look something up,
14:00 I use ChatGPT to ask a quick question, get the information I wanted fast.
14:04 For most daily tasks, I use the quick model.
14:07 Nowadays, I think the auto mode is pretty good
14:09 where you don't have to specifically say thinking or non-thinking.
14:12 Then again, I also sometimes want the pro mode.
14:15 Very often what I do is, when I have something written,
14:18 I put it into ChatGPT and say, "Hey, do a very thorough check.
14:23 Are all my references correct?
14:24 Are all my thoughts correct?
14:26 Did I make any formatting mistakes and are
14:28 the figure numbers wrong?" Or something like that.
14:31 And I don't need that right away.
14:33 I finish my stuff, maybe have dinner, let it run, come back and go through this.
14:38 I think this is where it's important to have this option.
14:42 I would go crazy if for each query I
14:43 would have to wait 30 minutes or 10 minutes even.
14:46 That's me.
14:48 I'm sitting over here losing my mind
14:50 that you use the router and the non-thinking model.
14:52 I'm like, "How do you live with that?" That's like my reaction.
14:57 I've been heavily on ChatGPT for a while.
15:01 I never touched ChatGPT-5 non-thinking.
15:03 I find its tone and then its propensity
15:05 for errors—it has a higher likelihood of errors.
15:08 Some of this is from back when OpenAI released o3,
15:11 which was the first model to do this deep
15:14 search and find many sources and integrate them for you.
15:17 I became habituated with that.
15:18 So I will only use GPT-5.2 Thinking or Pro
15:21 when I'm finding any sort of information query for work,
15:25 whether that's a paper or some code reference that I found.
15:28 And I will regularly have like five Pro queries going simultaneously,
15:33 each looking for one specific paper or feedback on an equation or something.
15:38 I have a fun example where I needed the answer as fast
15:41 as possible for this podcast before I was going on the trip.
15:46 like a local GPU running at home and I wanted to run a long RL experiment.
15:50 And usually I also unplug things because you never know if you're not at home,
15:54 you don't want things plugged in.
15:56 And I accidentally unplugged the GPU.
15:57 My wife was already in the car and it's like,
16:00 "Oh dang." Then basically I wanted as fast as possible
16:04 a Bash script that runs my different experiments and the evaluation.
16:09 And it's something I know,
16:10 I learned how to use the Bash interface or Bash terminal,
16:14 but in that moment I just needed like 10 seconds, give me the command.
16:18 This is a hilarious situation but yeah, so what did you use?
16:21 So I did the non-thinking fastest model.
16:23 It gave me the Bash command to chain
16:26 different scripts to each other and then the thing
16:29 is like you have the tee thing where you want to route this to a log file.
16:34 Top of my head I was just like in a hurry, I could have thought about it myself.
16:37 By the way I don't know if there's a representative case,
16:39 wife waiting in the car-...
16:40 you have to run, you know, unplug the GPU.
16:42 You have to generate a Bash script.
16:43 This sounds like a movie, like- Mission Impossible.
16:46 I use Gemini for that.
16:47 So I use thinking for all the information stuff and then
16:50 Gemini for fast things or stuff that I could sometimes Google,
16:52 which is like it's good at explaining things and I trust
16:55 that it has this kind of background of knowledge and it's simple.
16:59 And the Gemini app has gotten a lot
17:00 better and- It's good for those sorts of things.
17:02 And then for code and any sort of philosophical discussion,
17:05 I use Claude Opus 4.5.
17:07 Also always with extended thinking.
17:09 Extended thinking and inference time scaling is just
17:11 a way to make the models marginally smarter.
17:14 And I will always err on that side when the progress is
17:18 very high because you don't know when that'll unlock a new use case.
17:21 And then sometimes use Grok for real-time
17:24 information or finding something on AI Twitter
17:26 that I knew I saw and I need to dig up and I just fixated on.
17:31 Although when Grok 4 came out, the Grok 4 SuperGrok Heavy,
17:35 which was like their pro variant was actually
17:37 very good and I was pretty impressed with it,
17:38 and then it just kind of like muscle memory
17:41 lost track of it with having the ChatGPT app open.
17:44 So I use many different things.
17:46 Yeah.
17:46 I actually do use Grok 4 Heavy for debugging.
17:50 For like hardcore debugging that the other ones can't solve,
17:54 I find that it's the best at.
17:56 And...
17:56 it's interesting 'cause you say ChatGPT is the best interface.
18:00 For me, for that same reason,
18:02 but this could be just momentum- Gemini is the better interface for me.
18:07 I think because I fell in love with their best needle in the haystack.
18:11 If I ever put something that has a lot of context but I'm looking
18:15 for very specific kinds of information to make sure it tracks all of it,
18:19 I find at least that Gemini for me has been the best.
18:24 So it's funny with some of these models,
18:26 if they win your heart over- for one particular feature on one particular day,
18:31 for that particular query, that prompt, you're like,
18:35 "This model's better." And so you'll just stick with it
18:38 for a bit until it does something really dumb.
18:41 There's like a threshold effect.
18:43 Some smart thing and then you fall in love with it
18:45 and then it does some dumb thing and you're like, "You know what?
18:47 I'm gonna switch and try Claude or ChatGPT." And all that kind of stuff.
18:51 This is exactly it: you use it until it breaks,
18:53 until you have a problem, and then you change the LLM.
18:57 And I think it's the same as how we use anything,
19:01 like our favorite text editor, operating systems, or the browser.
19:04 I mean, there are many options: Safari, Firefox, Chrome.
19:07 They're relatively similar, but then there are edge cases,
19:11 extensions you want, and then you switch.
19:14 But I don't think anyone types the same
19:18 thing into different browsers and compares them.
19:21 You only do that when something breaks.
19:23 So that's a good point.
19:25 You use it until it breaks, then you explore other options.
19:28 On the long context thing, I was also a Gemini user,
19:31 but the GPT-5.2 release blog had crazy long context scores.
19:34 People were like, "Did they just figure out some algorithmic change?"
19:38 It went from 30% to 70% in this minor model update.
19:42 It's very hard to keep track of all of these things,
19:46 but now I look more favorably at GPT-5.2's long context.
19:49 So it's just like, "How do I actually
19:53 get to testing this?" It's a never-ending battle.
19:57 Well, it's interesting that none of us talked
19:59 about the Chinese models from a usage perspective.
20:02 What does that say?
20:04 Does it mean the Chinese models are not as good,
20:08 or are we just very biased and US-focused?
20:11 I think currently there's a discrepancy between the model and the platform.
20:15 The open models are more known for the open weights, not the platform yet.
20:19 known for the open weights, not their platform yet.
20:21 Many companies will sell you open-model inference at a very low cost.
20:25 With OpenRouter, it's easy to look at multi-model things.
20:29 You can run DeepSeek on Perplexity.
20:31 Sitting here, we're like, "We use OpenAI GPT-5 Pro consistently." We're all
20:36 willing to pay for the marginal intelligence gain.
20:39 These models from the US are better in terms of the outputs.
20:45 I think the question is,
20:47 will they stay better for this year and for years to come?
20:51 As long as they're better, I'm gonna pay for them.
20:55 There's also analysis showing that the way
20:59 the Chinese models are served—you could argue
21:01 this is due to export controls— is that they use fewer GPUs per replica,
21:05 which makes them slower and have different errors.
21:07 If speed and intelligence are in your favor as a user,
21:10 in the US, a lot of users will go for this.
21:13 And I think that will spur these Chinese
21:15 companies to want to compete in other ways,
21:18 whether it's free or substantially lower costs,
21:21 or it'll breed creativity in terms of offerings,
21:24 which is good for the ecosystem.
21:26 But the simple thing is: the US models are currently better, and we use them.
21:30 I tried these other open models, and I'm like, "Fun,
21:32 but I don't go back." models, and I'm like, "Fun, but not gonna...
21:36 I don't go back to it."- We didn't really mention programming.
21:40 That's another use case that a lot of people deeply care about.
21:45 I use basically half-and-half Cursor and Claude Code, because they're...
21:49 I fundamentally different experiences and both are useful.
21:54 What do you guys...
21:55 You program quite a bit, so what do you use?
21:57 What's the current vibe?
21:59 So, I use the Codeium plugin for VS Code.
22:02 You know, it's very convenient.
22:03 It's just like a plugin,
22:04 and then it's a chat interface that has access to your repository.
22:06 I know that Claude Code is, I think, a bit different.
22:10 It is a bit more agentic.
22:11 It touches more things.
22:12 It does the whole project for you.
22:13 I'm not quite there yet where I'm comfortable
22:16 with that because maybe I'm a control freak,
22:18 but I still would like to see a bit what's going on.
22:21 And Codeium is kind of, right now, for me,
22:24 the sweet spot where it is helping me, but it is not taking completely over.
22:29 I should mention, one of the reasons I do use
22:31 Claude Code is to build the skill of programming with English.
22:34 I mean, the experience is fundamentally different.
22:37 You're...
22:38 As opposed to micromanaging the details
22:40 of the process of the generation of the code, and looking at the diff,
22:45 which you can in Cursor if that's the IDE you use, and in changing, altering.
22:52 Looking and reading the code and understanding the code deeply as you progress,
22:56 versus just thinking in this design space
23:00 and just guiding it at this macro level,
23:05 which I think is another way of thinking about the programming process.
23:10 Also, we should say that Claude Code just seems
23:14 to be somehow a better utilization of Claude Opus 4.5.
23:19 It's a good side-by-side for people to do.
23:20 You can have Claude Code open,
23:21 you can have Cursor open, you can have VS Code open,
23:24 and you can select the same models on all of them— ...and ask questions,
23:27 and it's very interesting.
23:28 Claude Code is way better in that domain.
23:32 It's remarkable.
23:33 All right, we should say that both of you are legit on multiple fronts:
23:36 researchers, programmers, educators, Tweeters.
23:43 And on the book front, too.
23:45 So Nathan, at some point soon, hopefully has an RLHF book coming out.
23:50 It's available for preorder, and there's a full digital preprint.
23:54 I'm just making it pretty and better organized for the physical thing,
23:56 which is a lot of why I do it,
23:58 because it's fun to create things that you think are excellent
24:01 in the physical form when so much of our life is digital.
24:05 I should say, going to Perplexity here,
24:07 Sebastian Raschka is a machine learning researcher
24:09 and author known for several influential books.
24:11 A couple of them that I wanted to mention—which is
24:14 a book I highly recommend—Build a Large Language Model from Scratch,
24:18 and the new one, Build a Reasoning Model from Scratch.
24:21 So, I'm really excited about that.
24:24 Building stuff from scratch is one of the most powerful ways of learning.
24:28 Honestly, building an LLM from scratch is a lot of fun.
24:30 It's also a lot to learn.
24:31 And like you said, it's probably the best
24:33 way to learn how something really works,
24:35 'cause you can look at figures, but figures can have mistakes.
24:38 You can look at concepts and explanations, but you might misunderstand them.
24:43 But if there is code, and the code works, you know it's correct.
24:48 I mean, there's no misunderstanding.
24:50 It's precise.
24:50 Otherwise, it wouldn't work.
24:52 And I think that's the beauty behind coding.
24:54 It doesn't lie.
24:56 It's math, basically.
24:57 So, even though with math,
24:59 I think you can have mistakes in a book you would never notice.
25:02 Because you are not running the math when you are reading the book,
25:05 you can't verify this.
25:06 And with code, what's nice is you can verify it.
25:09 Yeah, I agree with you about the Build an LLM from Scratch book.
25:12 It's nice to tune out everything else,
25:14 the internet and so on, and just focus on the book.
25:16 But, you know, I read several history books.
25:21 It's just less lonely somehow.
25:24 It's really more fun.
25:25 Like for example, on the programming front,
25:28 I think it's genuinely more fun to program with an LLM.
25:31 And I think it's genuinely more fun to read with an LLM.
25:36 But you're right.
25:37 That distraction should be minimized.
25:40 So you use the LLM to basically enrich the experience, maybe add more context.
25:48 I just find the rate of aha moments for me is really high with LLMs.
25:55 100%.
25:55 I also want to correct myself: I'm not suggesting not to use LLMs.
25:58 I suggest doing it in multiple passes.
26:01 Like, one pass just offline, focus mode, and then after that...
26:05 I mean, I also take notes, but I,
26:07 I try to resist the urge to immediately look things up.
26:11 I do a second pass.
26:13 It's just more structured this way.
26:15 Sometimes things are answered in the chapter,
26:18 but sometimes also it just helps to let it sink in and think about it.
26:22 Other people have different preferences.
26:24 I highly recommend using LLMs when reading books.
26:26 For me, it's not the first thing to do; it's the second pass.
26:30 My recommendation is the opposite.
26:31 I like to use the LLM at the beginning to lay out
26:36 the full context of what is this world that I'm now stepping into?
26:40 But I try to avoid clicking out of the LLM into the world of Twitter and blogs,
26:47 because then you're down this rabbit hole.
26:50 You're reading somebody's opinion.
26:51 There's a flame war about a particular topic and all of a sudden
26:55 you're in the realm of the internet and Reddit and so on.
27:00 But if you're purely letting the LLM give you the context of why this matters,
27:05 what are the big picture ideas...
27:07 sometimes books are good at doing that, but not always.
27:12 This is why I like the ChatGPT app,
27:14 because it gives the AI a home on your computer where you can focus on it,
27:17 rather than just being another tab in my mess of internet options.
27:21 And I think Claude Code does a good job of making that a joy,
27:26 where it seems very engaging as a product design to be
27:30 an interface that your AI will then go out into the world.
27:34 It's something that is intangible between it and Codex;
27:37 it just feels warm and engaging, where Codex can often be as good from OpenAI,
27:42 but it just, feel a little bit rough around the edges.
27:46 Whereas Claude Code makes it fun to build things from scratch,
27:50 where you just trust that it'll make something.
27:53 Obviously this is good for websites and kind of refreshing tooling
27:57 and stuff like this, which I use it for, or data analysis.
28:01 For my On my blog, we scrape Hugging Face
28:04 so we keep download numbers for every dataset and model.
28:06 over time, so we have them.
28:07 And Claude was just like, "Yeah,
28:09 I've made use of that data, no problem." And I was like,
28:12 "That would've taken me days." And then
28:14 I have enough situational awareness to be like,
28:16 "Okay, these trends obviously make sense." You can check things.
28:18 But that's just a wonderful interface where you
28:20 can have an intermediary and not have to do
28:23 the kind of awful low-level work that you
28:26 would have to do to maintain different web projects.
28:29 All right.
28:30 So we just talked about a bunch of the closed-weight models.
28:33 Let's talk about the open ones.
28:36 Tell me about the landscape of open LLM models.
28:39 Which are interesting?
28:40 Which stand out to you and why?
28:42 We already mentioned DeepSeek R1.
28:45 Do you wanna see how many we can name off the top of our head?
28:47 Yeah, without looking at notes.
28:49 DeepSeek, Kimi, MiniMax, Z.ai, Moonshot.
28:53 We're just going Chinese.
28:57 Let's throw in Mistral AI, Gemma...
29:01 ...gpt-oss, the open weight model by OpenAI.
29:04 Actually, NVIDIA had a really cool one, Nemotron 3.
29:09 There, there's a lot of stuff especially at the end of the year.
29:11 Qwen might be the one—- Oh, yeah.
29:13 Qwen was the obvious name I was gonna say.
29:15 You can get at least 10 Chinese and at least 10 Western.
29:18 I think that OpenAI released their first open model— ...since GPT-2.
29:23 When I was writing about OpenAI's open model release, they were like,
29:27 "Don't forget about GPT-2," which I thought was
29:29 really funny 'cause it's just such a different time.
29:32 But gpt-oss-120b is actually a very strong model and does
29:35 some things that other models don't do very well.
29:39 Selfishly, I'll promote a bunch of Western companies
29:43 in the US and Europe that have these fully open models.
29:46 I work at the Allen Institute for AI,
29:48 where we've been building OLMo, which releases data and code.
29:51 And now we have actual competition for people that are
29:55 trying to release everything so that others can train these models.
29:58 There's the Institute for Foundation Models/LM360,
30:00 which has had their K2 models of various types.
30:04 Apertus is a Swiss research consortium.
30:07 Hugging Face has SmolLM, which is very popular.
30:12 And NVIDIA's Nemotron 3 has started releasing data as well.
30:15 And then Stanford's Martini Community Project,
30:17 which is kind of making it so there's a pipeline for people to open a GitHub
30:21 issue and implement a new idea and then
30:23 have it run in a stable language modeling stack.
30:26 This space, that list was way smaller in 2024— ...so I think it was just AI2.
30:32 So it's a great thing for more people
30:34 to get involved and to understand language models,
30:36 which doesn't really have a Chinese analog.
30:39 While I'm talking, I'll say that the Chinese
30:44 open language models tend to be much bigger,
30:47 and that gives them higher peak performance as MoEs,
30:49 where a lot of these things that we like a lot,
30:52 whether it was Gemma and Nemotron, have tended to be smaller models from the US,
30:57 which is starting to change from the US and Europe.
30:59 Mistral Large 3 came out, which was a giant MoE model,
31:02 very similar to DeepSeek architecture in December.
31:05 And then a startup, RCAI,
31:08 and both Nemotron and NVIDIA have teased MoE models way bigger than 100
31:15 billion parameters- like this 400 billion parameter
31:17 range coming in this Q1 2026 timeline.
31:20 So I think this kind of balance is set
31:23 to change this year in terms of what people
31:25 are using the Chinese versus US open models for, which
31:28 I'm personally going to be very excited to watch.
31:32 First of all, huge props for being able to name so many of these.
31:36 Did you actually name LLaMA?
31:39 No.
31:39 I feel like...
31:41 RIP.
31:41 This was not on purpose.
31:43 RIP LLaMA.
31:45 All right.
31:45 Can you mention some interesting models that stand out?
31:48 You mentioned Qwen 3 is obviously a standout.
31:51 So I would say the year's almost bookended by both DeepSeek V3 and R1.
31:56 And then on the other hand, in December, DeepSeek-V3.2.
31:59 Because what I like about those is they always
32:01 have an interesting architecture tweak that others don't have.
32:05 But otherwise, if you want to go with the familiar but really good performance,
32:09 Qwen 3 and, like Nathan said, also gpt-oss-120b.
32:13 And I think what's interesting about it is it's kind of like the first
32:18 public or open weight model that was really trained with tool use in mind,
32:22 which I do think is kind of a paradigm
32:25 shift where the ecosystem was not quite ready for it.
32:27 By tool use, I mean that the LLM is able
32:30 to do a web search or to call a Python interpreter.
32:33 And I do think it's a standout because it's a huge unlock.
32:37 Because one of the most common complaints about LLMs are,
32:41 for example, hallucinations, right?
32:43 And so, in my opinion, one of the best ways to solve hallucinations is
32:46 to not try to always remember information or make things up.
32:51 For math, why not use a calculator app or Python?
32:54 If I ask the LLM, "Who won the soccer
32:57 World Cup in 1998?" instead of just trying to memorize, it could go do a search.
33:03 I think mostly it's still a Google search.
33:06 So ChatGPT and gpt-oss-120b, they would do a tool call to Google,
33:09 maybe find the FIFA website.
33:11 Find, okay, it was France.
33:13 It would get you that information reliably
33:15 instead of just trying to memorize it.
33:17 So I think it's a huge unlock which right
33:20 now is not fully utilized yet by the open-source, open-weight ecosystem.
33:24 A lot of people don't use tool call modes because I think,
33:28 first, it's a trust thing.
33:29 You don't want to run this on your computer where it has access to tools,
33:32 could wipe your hard drive or whatever.
33:34 So you want to maybe containerize that.
33:36 But I do think that is like a really
33:40 important step for the upcoming years to have this ability.
33:44 So a few quick things.
33:45 First of all, thank you for defining what you mean by tool use.
33:49 I think that's a great thing to do
33:50 in general for the concepts we're talking about.
33:53 Even things as sort of well-established as MoEs.
33:57 You have to say that means mixture of experts,
34:00 and you kind of have to build up an intuition for people what that means,
34:04 how it's actually utilized, what are the different flavors.
34:06 So what does it mean that there's just such an explosion of open models?
34:11 What's your intuition?
34:13 If you're releasing an open model,
34:14 you want people to use it, is the first and foremost thing.
34:17 And then after that comes things like transparency and trust.
34:20 I think when you look at China, the biggest reason is that they want
34:24 people around the world to use these models,
34:26 and I think a lot of people will not.
34:28 If you look outside of the US, a lot of people will not pay for software,
34:31 but they might have computing resources where you
34:32 can put a model on it and run it.
34:34 I think there can also be data that you don't want to send to the cloud.
34:37 So the number one thing is getting people to use models, use AI,
34:41 or use your AI that might not be able
34:43 to do it without having access to the model.
34:46 I guess we should state explicitly,
34:47 so we've been talking about these Chinese models and open weight models.
34:51 Oftentimes, the way they're run is locally.
34:54 So it's not like you're sending your data
34:58 to China or to whoever developed Silicon Valley, or whoever developed the model.
35:04 A lot of American startups make money
35:06 by hosting- ...these models from China and selling them.
35:09 It's called selling tokens,
35:11 which means somebody will call the model to do some piece of work.
35:15 I think the other reason is for US companies like OpenAI.
35:18 They are so GPU deprived.
35:20 They're at the limits of the GPUs.
35:22 Whenever they make a release, they're always talking about like,
35:25 "Our GPUs are hurting." And I think
35:27 during one of these gpt-oss-120b release sessions,
35:30 Sam Altman said, "Oh, we're releasing this because we can use your GPUs.
35:33 We don't have to use our GPUs, and OpenAI can still get distribution out
35:38 of this," which is another very real thing,
35:41 because it doesn't cost them anything.
35:44 And for the user, I think also,
35:45 there are users who just use the model locally how they would use ChatGPT.
35:49 But also for companies I think it's a huge
35:51 unlock to have these models because you can customize them,
35:53 you can train them, you can add post-training, add more data.
35:58 Like, specialize them into, let's say, law, medical models, whatever you have.
36:02 And the appeal, you mentioned Llama,
36:04 the appeal of the open-weight models from China
36:07 is that the open-weight models' licenses are even friendlier.
36:11 I think they are just unrestricted open source licenses
36:13 where if we use something like Llama or Gemma, there are some strings attached.
36:17 I think it's like an upper limit in terms of how many users you have.
36:20 And then if you exceed, I don't know, so and so many million users,
36:23 you have to report your financial situation to, let's say,
36:26 Meta or something like that.
36:28 And I think while it is a free model, there are strings attached,
36:33 and people do like things where strings are not attached.
36:36 So I think that's also one of the reasons, besides performance,
36:39 why the open-weight models from China are so popular,
36:42 because you can just use them.
36:43 There's no catch in that sense.
36:46 The ecosystem has gotten better on that front,
36:48 but mostly downstream of these new providers providing such open licenses.
36:51 That was funny when you pulled up Perplexity and said,
36:53 "Kimi K2 Thinking hosted in the US." Which is just like an exact...
36:56 I've never seen this, but it's an exact example
36:58 of what we're talking about where people are sensitive to this.
37:01 But Kimi K2 Thinking and Kimi K2 is a model that is very popular.
37:05 People say that has very good creative
37:07 writing and also in doing some software things.
37:09 So it's just these little quirks that people
37:11 pick up on with different models that they like.
37:14 What are some interesting ideas that some of these models have
37:18 explored that you can speak to, that are particularly interesting to you?
37:22 Maybe we can go chronologically.
37:23 I mean, there was, of course, DeepSeek.
37:25 DeepSeek R1 that came out in January of 2025, if we just focus on 2025.
37:29 However, this was based on DeepSeek-V3,
37:31 which came out the year before in December 2024.
37:34 There are multiple things on the architecture side.
37:37 What is fascinating is...
37:38 I mean, that's what I do with my from-scratch coding projects.
37:41 You can still start with GPT-2,
37:43 and you can add things to that model to make it into this other model.
37:47 So it's all still kind of like the same lineage.
37:50 It is a very close relationship between those.
37:53 But top of my head, DeepSeek—what was unique there is the Mixture of Experts.
37:57 Not that they were inventing Mixture of Experts—we
38:00 can maybe talk a bit more about what Mixture of Experts means—but just to list
38:04 these things first before we dive into detail.
38:07 Mixture of Experts, but then they also had Multi-head Latent Attention,
38:11 which is a tweak to the attention mechanism, where this was, I would say,
38:17 the main distinguishing factor between these open-weight models.
38:22 Different tweaks to make inference or KV cache size...
38:25 We can also define KV cache in a few moments,
38:29 but to kind of make it more economical to have long context,
38:32 to shrink the KV cache size.
38:34 So what are tweaks that we can do?
38:36 And most of them focused on the attention mechanism.
38:38 There is Multi-head Latent Attention in DeepSeek.
38:41 There is Group Query Attention, which is still very popular.
38:44 It's not invented by any of those models.
38:46 It goes back a few years.
38:47 But that would be the other option.
38:50 Sliding window attention—I think OLMo 3 uses it, if I remember correctly.
38:54 So there are these different tweaks that make the models different.
38:57 Otherwise, I put them all together
39:00 in an article once where I just compared them.
39:03 They are very, surprisingly similar.
39:05 It's just different numbers in terms of how many
39:08 repetitions of the transformer block you have in the center.
39:11 And, like, just little knobs that people tune.
39:14 But what's so nice about it is it works no matter what.
39:17 You can tweak things.
39:19 You can move the normalization layers around to get some performance gains.
39:22 And OLMo is always very good in ablation studies,
39:26 showing what it actually does to the model if you move something around.
39:30 Ablation studies: does it make it better or worse?
39:32 But there are so many, let's say,
39:33 ways you can implement a transformer and make it still work.
39:36 The big ideas that are still prevalent is Mixture of Experts,
39:40 multi-head latent attention, sliding window attention, group query attention.
39:44 And then at the end of the year, we saw a focus on making the attention
39:49 mechanism scale linearly with inference token prediction.
39:52 So there was Qwen2-VL, for example, which added a gated delta net.
39:57 It's kind of inspired by State space models,
40:00 where you have a fixed state that you keep updating.
40:02 But it makes essentially this attention cheaper,
40:06 or it replaces attention with a cheaper operation.
40:08 And it may be useful to step
40:11 back and talk about transformer architecture in general.
40:14 Yeah, so maybe we should start with the GPT-2 architecture.
40:17 The transformer that was derived from the "Attention Is All You Need" paper.
40:21 The "Attention Is All You Need" paper
40:23 had a transformer architecture that had two parts, an encoder and a decoder.
40:28 And GPT went just focusing in on the decoder part.
40:32 It is essentially still a neural network
40:35 and it has this attention mechanism inside.
40:37 And you predict one token at a time.
40:40 You pass it through an embedding layer.
40:43 There's the transformer block.
40:44 The transformer block has attention modules and a fully connected layer.
40:47 And there are some normalization layers in between.
40:50 But it's essentially neural network layers with this attention mechanism.
40:53 So coming from GPT-2 when we move on to gpt-oss-120b,
40:57 there is, for example, the Mixture of Experts layer.
41:00 It's not invented by gpt-oss-120b.
41:02 It's a few years old.
41:04 But it is essentially a tweak to make the model
41:09 larger without consuming more compute in each forward pass.
41:13 So there is this fully connected layer,
41:15 and if listeners are familiar with multi-layer perceptrons,
41:19 you can think of a mini multi-layer perceptron,
41:22 a fully connected neural network layer inside the transformer.
41:25 And it's very expensive, because it's fully connected.
41:27 If you have a thousand inputs and a thousand outputs,
41:30 that's like one million connections.
41:31 And it's a very expensive part in this transformer.
41:34 And the idea is to kind of expand that into multiple feedforward networks.
41:39 So instead of having one, let's say you have 256,
41:43 but it would make it way more expensive,
41:45 because now you have 256, but you don't use all of them at the same time.
41:49 So you now have a router that says, "Okay, based on this input token,
41:52 it would be useful to use this fully connected network." And in that context,
41:57 it's called an expert.
41:58 So a Mixture of Experts means you have multiple experts.
42:01 And depending on what your input is, let's say it's more math-heavy,
42:05 it would use different experts, compared to, let's say,
42:09 translating input text from English to Spanish.
42:11 It would maybe consult different experts.
42:13 It's not quite clear, I mean, not as clear-cut to say, "Okay,
42:16 this is only an expert for math and for Spanish." It's a bit more fuzzy.
42:20 But the idea is essentially that you pack more knowledge into the network,
42:25 but not all the knowledge is used all the time.
42:27 That would be very wasteful.
42:29 So, during the token generation, you are more selective.
42:32 There's a router that selects which tokens should go to which expert.
42:36 It adds more complexity.
42:38 It's harder to train.
42:39 There's a lot that can go wrong, like collapse and everything.
42:42 So I think that's why OLMo 3 still uses dense...
42:45 I mean, you have OLMo models with Mixture of Experts,
42:48 but dense models, where dense means...
42:50 So also, it's jargon.
42:52 There's a distinction between dense and sparse.
42:55 So Mixture of Experts is considered sparse, because we have a lot of experts,
42:59 but only a few of them are active.
43:01 So that's called sparse.
43:01 And then dense would be the opposite,
43:03 where you only have one fully connected module, and it's always utilized.
43:08 So maybe this is a good place to also talk about KV cache.
43:11 But actually, before that, even zooming out, like fundamentally,
43:14 how many new ideas have been implemented from GPT-2 to today?
43:22 Like, how different really are these architectures?
43:25 Take the Mixture of Experts.
43:27 The attention mechanism in gpt-oss-120b,
43:29 that would be the Group Query Attention mechanism.
43:31 So it's a slight tweak from Multi-Head Attention to Group Query Attention.
43:35 So that we have too...
43:36 I think they replaced LayerNorm by RMSNorm,
43:39 but it's just like a different normalization there and not a big change.
43:43 It's just like a tweak.
43:45 The nonlinear activation function— people familiar with deep neural networks,
43:49 I mean, it's the same as changing sigmoid with ReLU.
43:52 It's not changing the network fundamentally.
43:55 It's just a little tweak.
43:56 And that's about it, I would say.
43:59 It's not really fundamentally that different.
44:01 It's still the same architecture.
44:03 So you can go from one into the other by just adding these changes basically.
44:10 It fundamentally is still the same architecture.
44:12 Yep.
44:12 For example, you mentioned my book earlier.
44:14 That's a GPT-2 model in the book because it's simple and it's very small,
44:18 so 124 million parameters approximately.
44:20 But in the bonus materials, I do have OLMo from scratch,
44:25 Gemini 3 from scratch, and other types of from-scratch models.
44:28 And I always start it with my GPT-2 model and just tweak the—well,
44:31 add different components and you get from one to the other.
44:34 It's kind of like a lineage in a sense.
44:38 Can you build up an intuition for people?
44:40 Because when you zoom out, you look at it,
44:43 there's so much rapid advancement in the AI world.
44:46 And at the same time, fundamentally the architectures have not changed.
44:51 So where is all the turbulence, the turmoil of the advancement happening?
44:58 Where are the gains to be had?
45:01 So there are different stages where you
45:03 develop the network or train the network.
45:05 You have the pre-training.
45:06 Now back in the day, it was just pre-training with GPT-2.
45:09 Now you have pre-training, mid-training, and post-training.
45:12 So I think right now we are in the post-training focus stage.
45:17 Pre-training still gives you advantages if you scale it up with better,
45:23 higher quality data.
45:24 But then we have capability unlocks that were not there with GPT-2,
45:28 for For example, ChatGPT is basically a GPT-3 model.
45:32 And GPT-3 is the same as GPT-2 in terms of architecture.
45:36 What was new was adding supervised
45:39 fine-tuning and reinforcement learning with human feedback.
45:41 So it's more on the algorithmic side than the architecture.
45:45 I would say that the systems also change a lot.
45:47 If you listen to NVIDIA's announcements,
45:48 they talk about things like, "You now do FP8,
45:51 you can now do FP4." What's happening is these labs are figuring
45:55 out how to utilize more compute to put it into one model,
45:58 which lets them train faster and put more data in.
46:01 And then you can find better configurations faster by doing this.
46:05 So you can look at, essentially, tokens per second per GPU as a metric
46:09 that you look at when you're doing large-scale training.
46:12 You can go from 10k to 13k by turning on FP8 training,
46:16 which means you're using less memory per parameter in the model.
46:20 By saving less information, you do less communication and train faster.
46:24 So all of these system things underpin
46:27 way faster experimentation on data and algorithms.
46:35 It's a loop that keeps going where it's hard to describe
46:38 when you look at architectures and they're exactly the same,
46:40 but the code base used to train
46:42 these models is vastly different- -and you could probably...
46:45 the GPUs are different but you probably train gpt-oss-20b way
46:49 faster in wall-clock time than GPT-2 was trained at the time.
46:54 Yeah.
46:54 Like you said, they had, for example,
46:56 in Mixture of Experts this FP4 optimization where you get more throughput.
47:00 But I do think, for speed this is true,
47:04 but it doesn't give the model new capabilities.
47:07 It's just: how much can we make the computation
47:11 coarser without suffering in terms of model performance degradation?
47:15 But I do think- I mean, there are alternatives popping up to the transformer.
47:20 Text diffusion models, a completely different paradigm.
47:23 And there is also...
47:24 I mean, although text diffusion models might use transformer architectures,
47:27 it's not an autoregressive transformer.
47:30 And also Mamba models.
47:32 It's a state space model.
47:34 But they do have trade-offs, and nothing has yet replaced
47:39 the autoregressive transformer as the state-of-the-art model.
47:42 For state-of-the-art, you would still go with that, but there are now
47:46 alternatives for the cheaper end—alternatives
47:49 that are kind of making compromises.
47:52 It's not just one architecture anymore.
47:54 There are little ones coming up.
47:57 But if we talk about the state-of-the-art,
47:59 it's pretty much still the transformer architecture,
48:02 autoregressive, derived from GPT-2 essentially.
48:06 I guess the big question here is,
48:07 we talked quite a bit about the architecture behind the pre-training.
48:11 Are the scaling laws holding strong across pre-training,
48:16 post-training, inference, context size, data, and synthetic data?
48:21 I'd like to start with the technical definition
48:22 of a scaling law- -which informs all of this.
48:24 The scaling law is the power law relationship between...
48:27 You can think of the x-axis,
48:28 so kind of what you are scaling as a combination of compute and data,
48:32 which are kind of similar,
48:34 and then the y-axis is like the held-out prediction accuracy over next tokens.
48:38 We talked about models being autoregressive.
48:39 It's like if you keep a set of text that the model has not seen,
48:45 how accurate will it get when you train?
48:47 And the idea of scaling laws came when people
48:50 figured out that that was a very predictable relationship.
48:53 And I think that that technical term is continuing,
48:57 and then the question is, what do users get out of it?
49:01 Then there are more types of scaling where,
49:03 OpenAI's o1 was famous for introducing inference time scaling.
49:06 And I think less famously for also
49:08 showing that you can scale reinforcement learning training
49:11 and get kind of this log x-axis
49:14 and then a linear increase in performance on y-axis.
49:16 So there's kind of these three axes now where
49:19 the traditional scaling laws are talked about for pre-training,
49:21 which is how big your model is and how big your dataset is,
49:25 and then scaling reinforcement learning, which is like how long can you do
49:28 this trial and error learning that we'll talk about.
49:30 We'll define more of this, and then this inference time compute,
49:33 which is just letting the model generate more tokens on a specific problem.
49:36 So I'm kind of bullish, but they're all really still working,
49:40 but the low-hanging fruit has mostly been taken,
49:43 especially in the last year on reinforcement learning with verifiable rewards,
49:46 which is this RLVR, and then inference time scaling,
49:50 which is just why these models feel so different to use,
49:53 where previously you would get that first token immediately.
49:55 And now they'll go off for seconds, minutes, or even hours,
49:59 generating these hidden thoughts before giving
50:01 you the first word of your answer.
50:03 And that's all about this inference time scaling,
50:05 which is such a wonderful kind of step
50:08 function in terms of how the models change abilities.
50:11 They kind of enabled this tool use stuff and enabled
50:13 this much better software engineering that we were talking about.
50:17 And this, when we say enabled, is almost entirely downstream of the fact
50:21 that this reinforcement learning with verifiable
50:23 rewards training just kind of let the models pick up these skills very easily.
50:27 So let the models learn, so if you look at the reasoning process
50:32 when the models are generating a lot of tokens,
50:34 what it'll often be doing is: it tries a tool, it looks at what it gets back.
50:37 It tries another API, it sees what it gets back and if it solves the problem.
50:41 So the models, when you're training them, very quickly learn to do this.
50:45 And then at the end of the day,
50:47 that gives this kind of general foundation where the model
50:49 can use CLI commands very nicely in your repo
50:52 and handle Git for you and move things around
50:55 and organize things or search to find more information,
50:57 which if we were sitting in these chairs a year ago
51:00 is something that we didn't really think of the models doing.
51:03 So this is just kind of something that has
51:05 happened this year and has totally transformed how
51:07 has totally transformed how we think of using
51:10 AI which evolution and just unlocks so much value.
51:13 But it's like, just so- pr- unlocks so much value.
51:18 But it's- it's like, it's not clear what the next avenue will
51:21 be in terms of unlocking stuff like this.
51:23 I think there's...
51:24 we'll get to continual learning later,
51:25 but there's a lot of buzz around certain areas of AI,
51:28 but no one knows when the next step function will really come.
51:32 So you've actually said quite a lot of things there,
51:35 and said profound things quickly.
51:37 It would be nice to unpack them a little bit.
51:40 You say you're bullish basically on every version of scaling.
51:43 So can we just even start at the beginning?
51:47 Pre-training, are we kind of implying that the low-
51:51 hanging fruit on pre-training scaling has been picked?
51:55 Has pre-training hit a plateau,
51:58 or is even pre-training still something you're bullish on?
52:01 Pre-training has gotten extremely expensive.
52:03 I think to scale up pre-training,
52:05 it's also implying that you're gonna serve a very large model to the users.
52:10 So I think that it's been loosely established the likes of GPT-4
52:14 and similar models were around one trillion parameters at the biggest size.
52:18 There's a lot of rumors that they've actually
52:20 gotten smaller as training has gotten more efficient.
52:23 You want to make the model smaller because
52:25 then your costs of serving go down proportionately.
52:28 These models, the cost of training them is really low relative
52:31 to the cost of serving them to hundreds of millions of users.
52:34 I think DeepSeek had this famous number of about
52:36 five million dollars for pre-training at cloud market rates.
52:40 In OLMo 3, section 2.4 in the paper, we just detailed how long we had the GPU
52:46 clusters sitting around for training which includes engineering issues,
52:50 multiple seeds, and it was like about two million dollars to rent
52:53 the cluster to deal with all the headaches of training a model.
52:56 So these models are pretty— like,
52:59 a lot of people could get one to 10 million dollars to train a model,
53:02 but the recurring costs of serving millions
53:05 of users is really billions of dollars of compute.
53:08 I think that you can look at a thousand
53:11 GPU rental you can pay 100 grand a day for.
53:14 And these companies could have millions of GPUs.
53:16 Like you can look at how much these things cost to sit around.
53:19 So that's kind of a big thing, and then it's like,
53:23 if scaling is actually giving you a better model,
53:25 is it gonna be financially worth it?
53:27 And I think we'll slowly push it out as AI solves more compelling tasks,
53:31 so like the likes of Claude Opus 4.5, making Claude Code just work for things.
53:36 I— I launched this project called the ATOM project,
53:39 which is American Truly Open Models in July,
53:42 and that was like a true vibe coded website,
53:45 and like, I have a job to make plots and stuff.
53:49 And then I came back to refresh it in the last few weeks
53:51 and it's like Claude Opus 4.5 versus whatever model at the time was like,
53:55 just crushed all the issues that it had from building in June and July and like,
53:59 it might be a bigger model.
54:01 There's a lot of things that go into this, but there's still progress coming.
54:04 So what you're speaking to is the nuance of the y-axis
54:07 of the scaling laws—the way it's experienced versus on a benchmark,
54:11 the actual intelligence might be different.
54:13 But still, your intuition about pre-training,
54:16 if you scale the size of compute, will the models get better?
54:21 Not whether it's financially viable but just from the law aspect of it,
54:26 do you think the models will get smarter?
54:28 Yeah.
54:29 And I think that there's...
54:30 And this sometimes comes off as almost like disillusionment from people,
54:34 leadership at AI companies saying this, but they're like,
54:37 "It's held for 13 orders of magnitude of compute,
54:40 why would it ever end?" So I think fundamentally it is pretty unlikely to stop,
54:44 it's just eventually we're not even gonna be able to test
54:47 the bigger scales because of all the problems that come with more compute.
54:50 I think that there's a lot of talk on how
54:54 2026 is a year when very large Blackwell compute clusters,
54:58 like gigawatt-scale facilities at hyperscalers, are coming online.
55:03 These were all contracts for power and data centers
55:06 that were signed and sought out in 2022 and 2023.
55:11 So before or right after ChatGPT.
55:13 It took this two-to-three-year lead time to build
55:16 these bigger clusters to train the models.
55:18 While there's obviously immense interest in building
55:20 even more data centers than that.
55:21 So that is the crux that people are saying: these new clusters are coming.
55:25 The labs are gonna have more compute for training.
55:28 They're going to utilize this, but it's not a given.
55:31 I've seen so much progress that I expect it,
55:34 and I expect a little bit bigger models, and I expect...
55:39 I would say it's more like we'll see a $2,000 subscription this year.
55:42 We've seen $200 subscriptions.
55:43 That could 10X again, and these are the kind of things that could come,
55:47 and they're all downstream of this bigger model
55:50 that offers just a little bit more cutting edge.
55:53 So, you know, it's reported that xAI
55:55 is gonna hit that one-gigawatt scale early '26,
55:59 and a full two gigawatts by year end.
56:03 How do you think they'll utilize that in the context of scaling laws?
56:09 Is a lot of that inference?
56:10 Is a lot of that training?
56:13 It ends up being all of the above.
56:15 So I think that all of your decisions
56:17 when you're training a model come back to pre-training.
56:20 So if you're going to scale RL on a model,
56:22 you still need to decide on your architecture that enables this.
56:25 We were talking about other architectures
56:27 and using different types of attention, or a mixture of experts models.
56:31 The sparse nature of MoE models makes it much more efficient to do generation,
56:37 which becomes a big part of post-training,
56:40 and you need to have your architecture ready
56:42 so that you can actually scale up this compute.
56:45 I still think most of the compute is going in at pre-training.
56:48 Because you can still make a model better,
56:51 you still want to go and revisit this.
56:53 You still want the best base model you can.
56:55 And in a few years that'll saturate and the RL compute will just go longer.
57:00 Are there people who disagree with you and say pre-training is dead?
57:06 It's all about scaling inference, scaling post-training,
57:09 scaling context, continual learning, scaling data, synthetic data?
57:15 People vibe that way and describe it in that way,
57:17 but I think it's not the practice that is happening.
57:19 It's just the general vibe of people saying
57:21 this thing is dead-- The excitement is elsewhere.
57:23 So the low-hanging fruit- ...in RL is elsewhere.
57:26 For example, we released our model in November...
57:28 Every company has deadlines.
57:30 Our deadline was November 20th, and for that, our run was five days,
57:34 which compared to 2024 is a very long time to just
57:37 be doing post- training at a model of 30 billion parameters.
57:40 It's not a big model.
57:41 And then in December, we had another release,
57:43 where we let the RL run for another three and a half weeks,
57:47 and the model got notably better, so we released it.
57:50 And that's a to just allocate to something
57:53 that is going to be your peak- ...for the year.
57:56 So it's like-- The reasoning is-- There's these types
57:59 of decisions when training a model where they just...
58:01 They can't leave it forever.
58:03 You have to keep pulling in the improvements from researchers.
58:07 So you redo pre-training, you'll do this post-training for a month,
58:11 but then you need to give it to your users.
58:14 You need to do safety testing.
58:15 So it's just...
58:16 I think there's a lot in place
58:18 that reinforces this cycle of updating the models.
58:21 Things improve.
58:22 You get a new compute cluster that lets you do something more stably or faster.
58:27 It's like you hear a lot about Blackwell having rollout issues, where at AI2,
58:32 most of the models we're pre-training are on 1,000 to 2,000 GPUs.
58:35 But when pre-training on 10,000 or 100,000 GPUs,
58:38 you hit very different failures.
58:40 GPUs break in weird ways, and on a 100,000 GPU run,
58:44 you're pretty much guaranteed to have one GPU that is down.
58:48 Your training code must handle that redundancy,
58:50 which is a very different problem.
58:51 Whereas what we're doing, like playing with post-training on a cluster,
58:55 or for people learning ML, what they're battling to train these biggest
59:00 models is just- ...mass distributed scale, and it's very different.
59:05 But that's somewhat different than- That's a systems
59:10 problem- ...in order to enable scaling laws, especially at pre-training.
59:15 You need all these GPUs at once.
59:17 When we shift to RL, it actually lends itself to heterogeneous compute
59:21 because you have many copies of the model.
59:24 To do a primer for language model reinforcement learning,
59:28 what you're doing is having two sets of GPUs.
59:31 One you can call the actor, and one you call the learner.
59:34 The learner is where your actual reinforcement learning updates happen.
59:38 These are traditionally policy gradient algorithms.
59:42 Proximal Policy Optimization, PPO, and Group Relative Policy Optimization,
59:46 GRPO, are the two popular classes.
59:50 And on the other side you have actors which are generating completions,
59:54 and these completions are what you're going to grade.
59:57 Reinforcement learning is all about optimizing reward.
1:00:00 In practice, you can have a lot of different actors
1:00:03 in different parts of the world doing different types of problems,
1:00:06 and then you send it back to this highly networked compute
1:00:10 cluster to do this actual learning where you take the gradients.
1:00:14 You need to have a tightly meshed network to do different
1:00:18 types of parallelism and spread out your model for efficient training.
1:00:22 Every different type of training and serving has these considerations to scale.
1:00:28 We talked about pre-training and RL,
1:00:30 and then inference time scaling- how do you serve
1:00:33 a model that's thinking for an hour to 100 million users?
1:00:35 I don't know about that, but I know that's a hard problem.
1:00:39 In order to give people this intelligence, there's all these systems problems,
1:00:42 and we need more compute and you need more stable compute to do it."-
1:00:46 But you're bullish on all of these kinds of scaling is what I'm hearing.
1:00:49 On the inference, on the reasoning, even on the pre-training?
1:00:54 Yeah, so that's a big can of worms here, but there are basically two...
1:00:58 The knobs are the training and the inference scaling where you can get gains.
1:01:02 In a world where we had, let's say,
1:01:05 infinite compute resources, you want to do all of them.
1:01:08 So you have training, you have inference scaling,
1:01:10 and training is like a hierarchy:
1:01:12 it's pre-training, mid-training, post-training.
1:01:13 Changing the model size, more training data,
1:01:16 training a bigger model gives you more knowledge in the model.
1:01:20 Then the model, let's say, has a better base model.
1:01:23 Back in the day, or still, we call it a foundation model, and it unlocks...
1:01:28 But you don't, let's say, have the model be able to solve
1:01:32 your most complex tasks during pre-training or after pre-training.
1:01:36 You still have these other unlock phases
1:01:38 where you have mid-training or, for example,
1:01:41 post-training with RL that unlocks capabilities that the model
1:01:43 has in terms of knowledge in the pre-training.
1:01:45 And I think, sure, if you do more pre-training,
1:01:50 you get a better base model that you can unlock later.
1:01:53 But like Nathan said, it just becomes too expensive.
1:01:55 We don't have infinite compute, so you have to decide,
1:01:58 do I want to spend that compute more on making the model larger?
1:02:01 It's like a trade-off.
1:02:02 In an ideal world, you want to do all of them.
1:02:05 And I think in that sense, scaling is still pretty much alive.
1:02:08 You would still get a better model,
1:02:09 but like we saw with Claude Opus 4.5, it's just not worth it.
1:02:12 Because you can unlock more performance
1:02:16 with other techniques at that current moment,
1:02:19 especially if you look at inference scaling.
1:02:21 That's one of the biggest gains this year with o1,
1:02:24 where it took a smaller model further than
1:02:29 pre-training a larger model like Claude Opus 4.5.
1:02:31 So I wouldn't say pre-training scaling is dead,
1:02:33 it's just that there are other more attractive ways to scale right now.
1:02:37 But at some point, you will still
1:02:39 want to make some progress on the pre-training.
1:02:42 The thing also to consider is where you want to spend your money.
1:02:46 If you spend it more on the pre-training, it's like a fixed cost.
1:02:50 You train the model, and then it has this capability forever.
1:02:53 You can always use it.
1:02:56 With inference scaling, you don't spend money during training,
1:02:58 you spend money later per query, and then it's also like math.
1:03:02 How long is my model gonna be on the market if I replace it in half a year?
1:03:06 Maybe it's not worth spending $5 million,
1:03:08 $10 million, $100 million on training it longer.
1:03:12 Maybe I will just do more inference scaling and get performance there.
1:03:17 It maybe costs me $2 million in terms of user queries.
1:03:19 It becomes a question of how many users you have and doing the math,
1:03:23 and I think that's also where it's interesting where ChatGPT is in a position.
1:03:26 I think they have a lot of users where they need to go a bit cheaper,
1:03:29 where they have that GPT-5 model that is a bit smaller.
1:03:32 Other companies that have...
1:03:34 Let's say, if your customers have other trade-offs.
1:03:38 For example, there was also the Math Olympiad or some
1:03:41 of these math problems where ChatGPT or they had a proprietary model,
1:03:46 and I'm pretty sure it's just like a model
1:03:49 that has been fine-tuned a little bit more, but most of it was during inference
1:03:53 scaling to achieve peak performance in certain tasks.
1:03:56 need that all the time.
1:03:57 But yeah, long story short, I do think all of these pre-training,
1:04:02 mid-training, post-training, inference scaling,
1:04:04 they are all still things you want to do.
1:04:06 It's just finding—at the moment, in this year,
1:04:08 it's finding the right ratio that gives
1:04:10 you the best bang for the buck, basically.
1:04:13 I think this might be a good place to define pre-training,
1:04:16 mid-training, and post-training.
1:04:18 So, pre-training is the classic training one next token prediction at a time.
1:04:21 You have a big corpus of data.
1:04:23 And Nathan probably also has very interesting insights there because of OLMo 3.
1:04:27 A big portion of the paper focuses on the right data mix.
1:04:30 So, pre-training is essentially just, you know, training cross entropy loss,
1:04:34 training on next token prediction on a vast corpus of internet data,
1:04:39 books, papers and so forth.
1:04:41 It has changed a little bit over the years
1:04:43 in the sense people used to throw in everything they can.
1:04:46 Now, it's not just raw data.
1:04:48 It's also synthetic data where people, let's say, rephrase certain things.
1:04:54 So synthetic data doesn't necessarily mean purely AI-made data.
1:04:58 It's also taking something from an article, a Wikipedia article,
1:05:02 and then rephrasing it as a Q&A question or summarizing it,
1:05:07 rewording it, and making better data that way.
1:05:12 Because I think of it also like with humans.
1:05:15 If someone, let's say, reads a book compared to a messy—no offense,
1:05:19 but like—Reddit post or something like that, I do think
1:05:24 you learn—- There's going to be a post about this, Sebastian.
1:05:28 Some Reddit data is very coveted and excellent for training.
1:05:31 You just have to filter it.
1:05:33 And I think that's the idea.
1:05:35 I think it's like if someone took that and rephrased it in a, let's say,
1:05:40 more concise and structured way,
1:05:42 I think it's higher quality data that gets the LLM there faster.
1:05:46 You get the same LLM out of it at the end,
1:05:49 but it trains faster because if the grammar and the punctuation are correct,
1:05:54 it already learns the correct way,
1:05:57 versus getting information from a messy source
1:05:59 and then learning later how to correct that.
1:06:02 So, I think that is how pre-training evolved and why scaling still works.
1:06:09 It's not just about the amount of data,
1:06:13 it's also the tricks to make that data better for you, in a sense.
1:06:17 And then mid-training is...
1:06:18 I mean, it used to be called pre-training.
1:06:21 I think it's called mid-training because it was awkward
1:06:23 to have pre-training and post-training but nothing in the middle, right?
1:06:26 It sounds a bit weird.
1:06:27 You have pre-training and post-training, but what's the actual training?
1:06:29 So, the mid-training is usually similar to pre-training,
1:06:33 but it's a bit more specialized.
1:06:35 It's the same algorithm, but what you do is you focus,
1:06:39 for example, on long-context documents.
1:06:42 The reason you don't do that during pre-training is
1:06:46 because you don't have that many long context documents.
1:06:49 We have a specific phase.
1:06:50 And one problem of LLMs is still that it's a neural network.
1:06:54 It has the problem of catastrophic forgetting.
1:06:56 So, you teach it something, it forgets other things.
1:06:58 And you wanna...
1:06:59 I mean, it's not 100% forgetting,
1:07:01 but it's like "no free lunch." It's the same with humans.
1:07:04 If you ask me some math I learned 10 years ago,
1:07:07 I would have to look at it again.
1:07:09 Nathan was actually saying that he's consuming so
1:07:11 much content that there's a catastrophic forgetting issue.
1:07:14 Yeah, I'm trying to learn so much about AI,
1:07:16 and it's like I was learning about pre-training parallelism.
1:07:18 I'm like, "I lost something and I don't know
1:07:21 what it was."- I don't want to anthropomorphize LLMs,
1:07:23 but it's the same kind of thing in how humans learn.
1:07:27 I mean, quantity is not always better because you have to be selective.
1:07:32 And mid-training is being selective in terms of quality content at the end.
1:07:36 So the last thing the LLM has seen is the quality stuff.
1:07:39 And then post-training is all the fine-tuning, supervised fine-tuning, DPO,
1:07:45 Reinforcement Learning with Verifiable Rewards (RLVR),
1:07:49 with human feedback, and so forth.
1:07:51 So the refinement stages.
1:07:52 And it's also interesting, it's a cost thing.
1:07:54 You spend a lot of money on pre-training right now.
1:07:57 RL a bit less.
1:07:59 With RL, you don't really teach it knowledge.
1:08:02 It's more like unlocking the knowledge; it's more like a skill learning,
1:08:05 like how to solve problems with the knowledge that it has from pre-training.
1:08:09 There are actually three papers this year,
1:08:11 or last year, 2025, on RL for pre-training.
1:08:14 But I don't think anyone does that in production.
1:08:17 Toy examples for now.
1:08:18 Toy examples, right?
1:08:19 But to generalize, RL post-training is more like the skill unlock,
1:08:23 where pre-training is like soaking up the knowledge.
1:08:27 A few things that could be helpful.
1:08:29 A lot of people think of synthetic data as being bad for training the models.
1:08:34 You mentioned how DeepSeek got almost...
1:08:37 OCR, which is Optical Character Recognition.
1:08:39 A lot of labs did it.
1:08:41 Ai2 had one, Meta had multiple.
1:08:44 And the reason each of these labs has
1:08:47 these is because there are vast amounts of PDFs
1:08:49 and other digital documents on the web that aren't
1:08:52 in formats that are encoded with text easily.
1:08:54 So you use these Almost-OCR, DeepSeek OCR, or what we called our Almost-OCR,
1:08:59 to extract trillions of tokens of candidate data for pre-training.
1:09:04 Pre-training dataset size is measured in trillions of tokens.
1:09:08 Smaller models from researchers can be something like five to 10 trillion.
1:09:11 researchers can be something like five to 10 trillion.
1:09:15 Um, Qwen is documented going up to 50 trillion,
1:09:17 and there are rumors that these closed labs can go to 100 trillion tokens.
1:09:21 Getting this potential data to put in—they have a very big funnel,
1:09:24 and the data you actually train on is a small percentage of this.
1:09:29 This character recognition data would be described
1:09:32 as synthetic data for pre-training in a lab.
1:09:34 And then there's also the fact that ChatGPT now gives wonderful answers,
1:09:38 and you can train on those best answers, and that's synthetic data.
1:09:41 It's very different than early ChatGPT with lots of hallucination data.
1:09:46 when people became grounded in synthetic data.
1:09:49 One interesting question is, if I recall correctly,
1:09:51 OLMo 3 was trained with less data than
1:09:53 specifically some other open-weight models, maybe even OLMo 2.
1:09:56 But you still got better performance,
1:09:58 and that might be one example of how the data helped.
1:10:01 It's mostly down to data quality.
1:10:02 I think if we had more compute, we would train for longer.
1:10:05 I think we'd ultimately see that as something we would want to do.
1:10:09 And especially with big models, you need more compute,
1:10:11 because we talked about having more parameters and we talked about knowledge.
1:10:15 Essentially, there's a ratio where big models can absorb more from data,
1:10:19 and then you get more benefit out of this.
1:10:22 Any logarithmic graph in your mind is like a small
1:10:25 model will level off sooner if you're measuring tons of tokens,
1:10:28 and bigger models need more.
1:10:30 But mostly, we aren't training that big of models right now at AI2,
1:10:34 and getting the highest quality data we can is the natural starting point.
1:10:38 Is there something to be said about the topic of data quality?
1:10:41 Is there some low-hanging fruit there still where the quality could be improved?
1:10:46 It's like turning the crank.
1:10:47 Historically, in the open,
1:10:49 there's been a canonical best pre-training dataset that has moved around
1:10:54 between who has the most recent one or the best recent effort.
1:10:56 Like AI2's Dolma was very early with the first OLMo,
1:10:59 and Hugging Face had FineWeb.
1:11:01 And there's a DCLM project,
1:11:02 which has been kind of like a, which stands for Data Comp Language Model.
1:11:07 There's been Data Comp for other machine learning projects,
1:11:10 and they had a very strong dataset.
1:11:12 And a lot of it is the internet is becoming fairly closed off,
1:11:17 so we have Common Crawl,
1:11:18 which is hundreds of trillions of tokens, and you filter it.
1:11:21 It looks like scientific work where you're training
1:11:24 classifiers and making decisions based on how you
1:11:27 prune down this dataset into the highest quality
1:11:30 stuff and the stuff that suits your tasks.
1:11:33 Previously, language models were tested a lot
1:11:35 more on knowledge and conversational things,
1:11:37 but now they're expected to do math and code.
1:11:40 To train a reasoning model, you need to remix your whole dataset.
1:11:43 And there's a lot of wonderful scientific methods here where you can,
1:11:47 you can take your gigantic dataset,
1:11:49 sample really tiny things from different sources,
1:11:52 such as GitHub, Stack Exchange, Reddit, Wikipedia.
1:11:56 You can sample small things from them,
1:11:57 and train small models on each of these mixes
1:12:00 and measure their performance on your evaluations.
1:12:02 You can just do basic linear regression, and it's like,
1:12:04 "Here's your optimal dataset." But if your evaluations change,
1:12:07 your dataset changes a lot.
1:12:08 So a lot of OLMo 3 was new sources for reasoning to be better at math and code,
1:12:13 and then you do this mixing procedure and it gives you the answer.
1:12:17 I think that's happened at labs this year;
1:12:19 there's new hot things, whether it's coding environments or web navigation,
1:12:23 and you need to bring in new data,
1:12:24 change your whole pre-training so that your post-training can work better.
1:12:28 And that's like the constant evolution and the redetermining
1:12:31 of what they care about for their models.
1:12:35 Are there fun anecdotes of what sources of data
1:12:38 are particularly high quality that we wouldn't expect?
1:12:41 You mentioned Reddit sometimes can be a source.
1:12:45 Reddit was very useful.
1:12:47 I think PDFs is definitely one.
1:12:51 Oh, especially arXiv.
1:12:52 Yeah, so AI2 has run Semantic Scholar for a long time,
1:12:56 which is what you can say is a competitor
1:12:59 to Google Scholar with a lot more features.
1:13:01 And to do this, AI2 has found and scraped a lot of PDFs for openly
1:13:06 accessible papers that might not be behind
1:13:09 the closed walled garden of a certain publisher.
1:13:11 So, truly open scientific PDFs.
1:13:13 And if you sit on all of these and you process it, you can get value out of it.
1:13:17 And I think that like,
1:13:19 a lot of that style of work has been done by the frontier labs did much earlier.
1:13:24 You just need to have a pretty
1:13:26 skilled researcher that understands how things change models;
1:13:30 they bring it in, clean it, and it's a lot of labor.
1:13:33 When frontier labs scale researchers, a lot more goes into data.
1:13:38 If you join a frontier lab and you want to have impact,
1:13:41 the best way to do it is just find new data that's better.
1:13:45 And then, the fancy, glamorous algorithmic things like figuring out how to make
1:13:50 o1 is like the sexiest thought of a scientist.
1:13:52 It's like, "Oh, I figured out how to scale RL." There's a group
1:13:55 that did that, but most of the contribution is like—- On the dataset-
1:13:58 ..."I'm gonna make the data better,"
1:14:00 or, "I'm gonna make the infrastructure better
1:14:01 so everyone on my team can run experiments 5% faster."- At the same time,
1:14:05 I think it's also one of the closest guarded secrets,
1:14:07 what your training data is, for legal reasons.
1:14:09 And so there's also, I think,
1:14:10 a lot of work that goes into hiding what your training data was essentially.
1:14:14 Like training the model to not give
1:14:17 away the sources because you have legal reasons.
1:14:19 The other thing, to be complete,
1:14:20 is that some people are trying to train on only licensed data,
1:14:23 whereas Common Crawl is a scrape of the whole internet.
1:14:26 So if I host multiple websites, I'm happy to have them train language models,
1:14:32 but I'm not explicitly licensing what governs it.
1:14:35 And therefore, Common Crawl is largely unlicensed,
1:14:38 which means that your consent really hasn't
1:14:41 been provided for how to use the data.
1:14:43 There's another idea where you can train language
1:14:44 models only on data that has been licensed explicitly,
1:14:47 so that the kind of governing contract is provided,
1:14:50 and I'm not sure if Apertus is the copyright thing or the license thing.
1:14:53 I know that the reason that they did it was for an EU compliance thing,
1:14:56 where they wanted to make sure that their model fit one of those checks.
1:15:05 On that note, there's also the distinction in licensing.
1:15:09 Some people just purchase the license.
1:15:12 Let's say they buy an Amazon Kindle book,
1:15:15 or a Manning book, and then use that in training.
1:15:18 That is a gray zone 'cause you paid
1:15:19 for the content and you might want to train on it.
1:15:22 But then there are also restrictions where even that shouldn't be allowed.
1:15:25 And so that is where it gets a bit fuzzy.
1:15:28 And yeah, I think that is right now still a hot topic.
1:15:33 Big companies like OpenAI approached private companies
1:15:36 for their proprietary data and private companies,
1:15:39 they become more and more, let's say,
1:15:42 protective of their data because they know, "Okay,
1:15:44 this is going to be my moat in a few
1:15:47 years." And I do think that's like the interesting question,
1:15:50 where if LLMs become more commoditized,
1:15:53 and I think a lot of people learn about LLMs,
1:15:56 there will be a lot more people able to train LLMs.
1:15:58 Of course, there are infrastructure challenges.
1:16:00 But if you think of big industries like pharmaceutical industries,
1:16:04 law, finance industries, I do think they, at some point,
1:16:07 will hire people from other frontier labs
1:16:10 to build their in-house models on their proprietary data,
1:16:13 which will be then, again,
1:16:14 another unlock with pre-training that is currently not there.
1:16:17 Because even if you wanted to, you can't get that data.
1:16:21 You can't get access to clinical trials most
1:16:23 of the time and these types of things.
1:16:25 So, I do think scaling, in that sense, might be still pretty much alive
1:16:28 if you also look in domain-specific applications,
1:16:31 because we are still right now, in this year,
1:16:33 just looking at general purpose LLMs on, on ChatGPT, Anthropic and so forth.
1:16:37 They are just general purpose, they're not even, I think,
1:16:40 scratching the surface of what an LLM can do if
1:16:43 it is really specifically trained and designed for a specific task.
1:16:47 I think on the data thing—this is one of the things that happened in 2025,
1:16:50 and we totally forget it—is Anthropic lost
1:16:52 in court and owed $1.5 billion to authors.
1:16:55 Anthropic, I think, bought thousands of books and scanned them
1:16:59 and was cleared legally for that because they bought the books,
1:17:03 and that is kind of going through the system.
1:17:04 And then the other side, they also torrented some books,
1:17:07 and I think this torrenting was the path where the court said
1:17:10 that they were then culpable to pay these billions of dollars to authors,
1:17:13 which is just such a mind-boggling lawsuit that kind of just came and went.
1:17:17 That is so much money-...
1:17:20 from the VC ecosystem.
1:17:22 These are court cases that will define the future of human civilization,
1:17:25 because it's clear that data drives a lot
1:17:27 of this, and there's this very complicated human tension of...
1:17:30 I mean, you can empathize.
1:17:33 You're both authors.
1:17:34 And there's some degree to which, I mean,
1:17:36 you put your heart and soul and your sweat
1:17:39 and tears into the writing that you do.
1:17:42 It feels a little bit like theft for somebody
1:17:46 to train your data without giving you credit.
1:17:49 And there are, like Nathan said, also two layers to it.
1:17:51 Someone might buy the book and then train on it,
1:17:54 which could be argued fair or not fair,
1:17:56 but then there are the straight-up companies who use
1:18:00 pirated books where it's not even compensating the author.
1:18:03 That is, I think, where people got
1:18:04 a bit angry about it specifically, I would say.
1:18:06 Yeah, but there has to be some kind of compensation scheme.
1:18:09 This is like moving towards-...
1:18:11 towards something like Spotify streaming did-...
1:18:13 originally for music.
1:18:14 You know, what does that-...
1:18:15 compensation look like?
1:18:16 You have to define those kinds of models.
1:18:17 You have to think through all of that.
1:18:19 One other thing I think people are generally curious about,
1:18:22 I'd love to get your thoughts, as LLMs are used more and more.
1:18:26 If you look at even arXiv, but GitHub,
1:18:29 more and more of the data is generated by LLMs.
1:18:32 What do you do in that kind of world?
1:18:36 How big of a problem is that?
1:18:39 Largest problem's the infrastructure and systems,
1:18:41 but from an AI point of view, it's kind of inevitable.
1:18:45 So it's basically LLM-generated data
1:18:47 that's curated by humans essentially, right?
1:18:49 Yes, and I think that a lot
1:18:51 of open source contributors are legitimately burning out.
1:18:53 If you have a popular open source repo,
1:18:55 somebody's like, "Oh, I want to do open source AI.
1:18:57 It's good for my career," and they just
1:19:00 vibe- -code something and they throw it in.
1:19:02 You might get more of this-- I have a--- than I do.
1:19:05 Yeah, so I have actually a case study here.
1:19:09 I have a repository called MLxtend that I
1:19:11 developed as a student around 10 years ago,
1:19:14 and it is a reasonably popular library still for certain algorithms,
1:19:18 I think especially like frequent data mining stuff.
1:19:22 And there were recently two or three people who submitted
1:19:25 a lot of PRs in a very short amount of time.
1:19:28 I do think LLMs have been involved in submitting these PRs.
1:19:31 Me, as the maintainer, there are two things.
1:19:33 First, I'm a bit overwhelmed.
1:19:35 I don't have time to read through it because,
1:19:37 especially as an older library, that is not a priority for me.
1:19:40 At the same time, I kind of also appreciate it because
1:19:43 I think something people forget is it's not just using the LLM.
1:19:46 There's still a human layer that verifies something,
1:19:48 and that is in a sense also how data is labeled, right?
1:19:53 One of the most expensive things is getting
1:19:56 labeled data for RL from human feedback phases.
1:19:59 And this is kind of like that, where it goes through phases,
1:20:03 and then you actually get higher quality data out of it.
1:20:06 And so I don't mind it in a sense.
1:20:08 It can feel overwhelming, but I do think there is also value in it.
1:20:12 It feels like there's a fundamental
1:20:14 difference between raw LLM-generated data and LLM-generated
1:20:16 data with a human in the loop that does some kind of verification,
1:20:21 even if that verification is a small percent of the lines of code.
1:20:26 I think this goes with anything where people think also sometimes, "Oh, yeah.
1:20:31 I can just use an LLM to learn about XYZ," which is true.
1:20:34 You can, but there might be a person who is
1:20:37 an expert who might have used an LLM to write specific code.
1:20:41 There is kind of like this human work that went into it to make it nice,
1:20:45 throwing out the not-so-nice parts to kind of pre-digest it for you,
1:20:50 and that saves you time.
1:20:52 I think that's the value-add,
1:20:53 where you have someone filtering things or even using the LLMs correctly.
1:20:59 This is still labor that you get for free.
1:21:02 For example, if you read a Substack article,
1:21:04 I could maybe ask an LLM to give me opinions
1:21:07 on that, but I wouldn't even know what to ask.
1:21:10 And I think there is still value in reading that article
1:21:13 compared to me going to the LLM because you are the expert.
1:21:17 You select what knowledge is actually spot on, should be included,
1:21:20 and you give me this very...
1:21:23 this executive summary.
1:21:24 And this is a huge value-add because now I don't have
1:21:28 to waste three to five hours to go through this myself,
1:21:31 maybe get some incorrect information and so on.
1:21:34 And so I think that's also where the future
1:21:37 still is for writers even though there are LLMs that...
1:21:41 Can kind of save you time.
1:21:44 It's kinda fascinating to watch.
1:21:45 I'm sure you guys do this, but for me,
1:21:48 I look at the difference between a summary and the original content.
1:21:53 Even if it's a page-long summary of a page-long content,
1:21:57 it's interesting to see how the LLM-based summary takes the edge off.
1:22:03 What is the signal it removes from the thing?
1:22:07 The voice is what I talk about a lot.
1:22:09 Voice?
1:22:10 Well, voice...
1:22:10 I would love to hear what you mean by voice,
1:22:13 but sometimes there's literally insights.
1:22:16 By removing an insight, you're changing the meaning of the thing.
1:22:21 So I'm continuously disappointed how bad LLMs
1:22:24 are at really getting to the core insights, which is what a great summary does.
1:22:31 Yet even if I have these extremely elaborate prompts
1:22:36 where I'm really trying to dig for the insights,
1:22:40 it's still not quite there, which...
1:22:42 I mean, that's a whole deep philosophical question about what is human knowledge
1:22:46 and wisdom and what does it mean to be insightful and so on.
1:22:49 But when you talk about the voice, what do you mean?
1:22:52 When I write, I think a lot of what I'm trying to do
1:22:55 is take what you think as a researcher, which is very raw.
1:22:59 A researcher is trying to encapsulate an idea at the frontier
1:23:02 of their understanding and they're trying to put what is a feeling into words.
1:23:07 I try to do this in my writing,
1:23:11 which makes it come across as raw but also high-information
1:23:14 in a way that some people will get it and some won't.
1:23:17 And that's the nature of research.
1:23:19 And language models don't do this well.
1:23:21 They're all trained with Reinforcement Learning from Human Feedback,
1:23:25 which takes feedback from many people
1:23:27 and averages how the model behaves from this.
1:23:30 And I think it's going to be hard for a model
1:23:34 to be very incisive when there's that sort of filter.
1:23:37 This is a wonderful fundamental problem for researchers in RLHF.
1:23:43 This provides so much utility in making the models better,
1:23:47 but also the problem formulation is kind of...
1:23:51 there's this knot in it that you can't get past.
1:23:55 These language models don't have this prior
1:23:57 in their deep expression they're trying to get at.
1:24:00 I don't think it's impossible.
1:24:01 There are stories of models that really shock people.
1:24:04 Like, I think of...
1:24:05 I would love to have tried Bing Sydney.
1:24:08 Did that have more voice?
1:24:10 Because it would so often go off the rails,
1:24:13 which is historically obviously a scary way—like telling a reporter to leave
1:24:17 his wife—is a crazy model to potentially put in general adoption.
1:24:21 But that's kind of like a trade-off;
1:24:23 is this RLHF process in some ways adding limitations?
1:24:28 That's a terrifying place to be as one of these frontier labs and companies,
1:24:33 because millions of people are using them.
1:24:36 There was a lot of backlash last year with GPT-4o getting removed,
1:24:39 and I've personally never used the model,
1:24:41 but I've talked to people at OpenAI where they get emails from users
1:24:47 that might be detecting subtle differences
1:24:50 in the deployments in the middle of the night.
1:24:52 And they email them like, "My friend is different." And they find
1:24:56 these employees' emails and send them things because they
1:24:58 are so attached to this set of model
1:25:02 weights and configuration that is deployed to the users.
1:25:05 We see this with TikTok.
1:25:06 You open it...
1:25:07 I don't use TikTok, but supposedly in five minutes the algorithm gets you.
1:25:11 It's locked in.
1:25:12 And those are language models doing recommendations.
1:25:15 I think there are ways you can do this.
1:25:18 Within five minutes of chatting with it, the model just gets you.
1:25:21 And that is something that people aren't really ready for.
1:25:26 Like, don't give that to kids.
1:25:28 At least until we know what's happening.
1:25:30 But there's also going to be this mechanism...
1:25:32 What's going to happen with these LLMs as they're used more and more...
1:25:36 Unfortunately, the nature of the human
1:25:37 condition is such that people commit suicide.
1:25:39 And so what journalists will do is they
1:25:42 will report extensively on the people who commit suicide.
1:25:45 And they would very likely link it to the LLMs
1:25:48 because they have that data about the conversations.
1:25:50 If you're really struggling in your life, if you're depressed,
1:25:54 if you're thinking about suicide,
1:25:55 you're going to probably talk to LLMs about it.
1:25:58 And so what journalists will do is say, "Well,
1:26:01 the suicide was committed because of the LLM."
1:26:03 And that's going to lead to the companies,
1:26:06 because of legal issues and so on, more and more taking the edge off of the LLM.
1:26:13 So it's going to be as generic as possible.
1:26:15 It's so difficult to operate in this space because you don't
1:26:19 want an LLM to cause harm to humans at that level,
1:26:23 but also this is the nature of the human experience,
1:26:27 is to have a rich conversation,
1:26:29 a fulfilling conversation, one that challenges you from which you grow.
1:26:33 You need that edge.
1:26:35 And that's something extremely difficult for AI researchers
1:26:39 on the RLHF front to actually have to solve,
1:26:44 because you're dealing with the human condition.
1:26:47 A lot of researchers at these companies are so well-motivated,
1:26:50 and definitely Anthropic and OpenAI culturally want to do good for the world.
1:26:56 And it's such a—I'm like, "Ooh,
1:26:58 I don't wanna work on this," because on the one hand,
1:27:01 a lot of people see AI as a health ally,
1:27:04 as somebody they can talk to about their health confidentially,
1:27:08 but then it bleeds all the way into this, like talking about mental health,
1:27:14 where it's heartbreaking that this will be
1:27:17 the thing where somebody goes over the edge, but other people might be saved.
1:27:21 And I'm like, "I don't..." As a researcher, it's like,
1:27:25 I don't want to train image generation
1:27:26 models and release them openly because I don't
1:27:29 want to enable somebody to have a tool
1:27:31 on their laptop that can harm other people.
1:27:34 I don't have the infrastructure in my company to do that safely.
1:27:37 But there's a lot of areas like this where it
1:27:40 just needs people that will approach it with complexity and conviction.
1:27:44 It's just such a hard problem.
1:27:48 But also, we as a society, as users of these technologies,
1:27:50 need to make sure that we're having the complicated conversation about it versus
1:27:54 just fearmongering— that Big Tech is causing
1:27:57 harm to humans or stealing your data.
1:27:59 It's more complicated than that.
1:28:01 And you're right.
1:28:03 There's a very large number of people inside these companies,
1:28:05 many of whom I know, who deeply care about helping people.
1:28:09 They are considering the full human experience of people from across the world,
1:28:13 not just Silicon Valley.
1:28:14 People across the United States and the world, what their needs are.
1:28:18 It's really difficult to design this one system that is able
1:28:23 to help all these different kinds of people across different age groups,
1:28:26 cultures, and mental states.
1:28:31 I wish that the timing of AI was different relative
1:28:34 to the relationship of Big Tech to the average person.
1:28:37 Big Tech's reputation was so low, and with how AI is so expensive,
1:28:40 it's inevitably going to be a Big Tech thing.
1:28:42 It takes so many resources, and people say the US is,
1:28:46 "betting the economy on AI" with this build-out.
1:28:48 To have these be intertwined at the same
1:28:51 time makes for such a hard communication environment.
1:28:54 It would be good for me to go talk to more people
1:28:56 in the world who hate Big Tech and see AI as a continuation of this.
1:29:03 And one of the things you recommend,
1:29:05 one of the antidotes that you talk about, is to find agency in this system.
1:29:11 As opposed to sitting back in a powerless way and consuming
1:29:16 the AI slop as it rapidly takes over the internet.
1:29:20 Find agency by using AI to build stuff, build apps, build...
1:29:25 One, that actually helps you build intuition, but two,
1:29:29 it's empowering because you can understand how it works,
1:29:32 what the weaknesses are.
1:29:34 It gives your voice power to say, "This is a bad use of the technology,
1:29:38 and this is a good use." And you're more plugged into the system then,
1:29:44 so you can understand it better and you can steer it better as a consumer.
1:29:48 I think that's a good point you brought up about agency.
1:29:51 Instead of ignoring it and saying, "Okay,
1:29:52 I'm not going to use it," I think it's probably long-term healthier to say,
1:29:57 "Okay, it's out there.
1:29:58 I can't put it back." when they came out.
1:30:01 How do I make best use of it, and how does it help me to up-level myself?
1:30:06 The one thing I worry about here, though, is,
1:30:08 if you just fully use it for something you love to do,
1:30:10 the thing you love to do is no longer there.
1:30:13 And that could potentially, I feel, lead to burnout.
1:30:16 For example, if I use an LLM to do all my coding for me, now there's no coding.
1:30:21 I'm just managing something that is coding for me.
1:30:23 Two years later, let's say, if I just do that eight hours a day,
1:30:27 having something code for me, do I feel fulfilled still?
1:30:36 Is this hurting me in terms of being excited about my job,
1:30:39 excited about what I'm doing?
1:30:40 Am I still proud to build something?
1:30:43 On that topic of enjoyment, it's quite interesting.
1:30:45 We should just throw this in there,
1:30:47 that there's this recent survey of about 791 professional developers,
1:30:52 meaning 10-plus years of experience.
1:30:56 That's a long time.
1:30:58 As a junior developer?
1:31:01 Yeah, in this day and age.
1:31:03 So, there's also many fronts that are surprising.
1:31:06 They break it down by junior and senior developers.
1:31:10 But, I mean, it just shows that both junior
1:31:14 and senior developers use AI-generated code in code they ship.
1:31:22 So this is not just for fun or intermediate learning things.
1:31:26 This is code they ship.
1:31:28 25%—like, most of them use around 50% or more.
1:31:32 And what's interesting is,
1:31:33 for the category of over 50% of your code that you ship is AI-generated,
1:31:38 senior developers are much more likely to do so.
1:31:42 But you don't want AI to take away the thing you love.
1:31:46 I think this speaks to my experience, these results I'm about to say.
1:31:49 Together, about 80% of people find it either somewhat more enjoyable
1:31:54 or significantly more enjoyable to use AI as part of the work.
1:31:59 I think it depends on the task.
1:32:01 From my personal usage, for example,
1:32:03 I have a website where I sometimes tweak things.
1:32:07 I personally don't enjoy this.
1:32:09 So in that sense, if the AI can help me
1:32:13 to implement something on my website, I'm all for it.
1:32:15 It's great.
1:32:16 But then, at the same time,
1:32:17 when I solve a complex problem— well, if there's a bug,
1:32:21 and I hunt this bug, and I find it, it's the best feeling in the world.
1:32:26 You feel great.
1:32:27 But now, if you don't even think about the bug,
1:32:31 you just go directly to the LLM, well,
1:32:33 you never have this kind of feeling, right?
1:32:35 But then there could be the middle ground where, well,
1:32:39 you try yourself, you can't find it, you use the LLM,
1:32:42 and then you don't get frustrated because it helps
1:32:44 you and you move on to something that you enjoy.
1:32:46 And so, looking at these statistics, I think what is not factored in is
1:32:51 that it's averaging over all the different scenarios.
1:32:54 We don't know if it's for the core task or if
1:32:59 it's for something mundane that people would not have enjoyed otherwise.
1:33:02 So, in a sense, AI is really great
1:33:04 for doing mundane things that take a lot of work.
1:33:06 For example, my wife the other day—she has
1:33:09 a podcast for book discussions, a book club,
1:33:13 and she was transferring the show notes from Spotify to YouTube,
1:33:19 and then the links somehow broke.
1:33:21 And she had in some episodes, because it is so many books, like 100 links,
1:33:25 and it would have been really painful to go in there and fix each link manually.
1:33:29 So I suggested, "Hey,
1:33:30 let's try ChatGPT." We copied the text into ChatGPT, and it fixed them.
1:33:35 Instead of two hours going from link to link fixing
1:33:39 that, it made that type of work much more seamless.
1:33:43 I think everyone has a use case where AI is
1:33:47 useful for something that would be really boring, really mundane.
1:33:51 For me personally, since we're talking about coding,
1:33:56 and you mentioned debugging...
1:33:58 the source of enjoyment for me, more on the Cursor side than Claude Code,
1:34:02 is that I have a friend, I have a pair programmer.
1:34:08 It's less lonely.
1:34:10 You made debugging sound like this great joy.
1:34:15 No, I would say debugging is like a drink
1:34:18 of water after you've been going through a desert for days.
1:34:22 You skip the whole desert part where you're suffering.
1:34:27 Sometimes it's nice to have a friend who can't really find the bug,
1:34:31 but can give you some intuition about the code,
1:34:35 and together you're going through the desert and finding that drink of water.
1:34:40 At least for me, maybe it speaks
1:34:43 to the loneliness of the programming experience.
1:34:45 That is a source of joy.
1:34:48 It's maybe also related to delayed gratification.
1:34:51 I'm a person who, even as a kid,
1:34:54 I liked the idea of Christmas presents—having them,
1:34:57 getting them—better than actually receiving the presents.
1:35:01 I would look forward to the day I get the presents,
1:35:04 but then it's over and I'm disappointed.
1:35:06 And maybe it's the same with food.
1:35:08 I think food tastes better when you're really hungry.
1:35:12 You're right, with debugging, it is not always great.
1:35:17 It's often frustrating, but then if you can solve it, then it's great.
1:35:23 But there's a sweet Goldilocks zone;
1:35:24 if it's too hard, then it's just wasting your time.
1:35:28 But I think another challenge is how will people learn?
1:35:33 We looked at the chart and saw that more senior
1:35:37 developers are shipping more AI-generated code than the junior ones.
1:35:42 It's very interesting,
1:35:42 because intuitively you would think it's the junior developers
1:35:45 because they don't know how to do the thing yet,
1:35:49 and so they use AI to do that thing.
1:35:51 It could mean the AI is not good enough yet to solve that task,
1:35:55 but it could also mean experts are more effective at using it.
1:35:59 They know how to use it better,
1:36:02 review the code, and then they trust the code more.
1:36:06 One issue for society in the future will be:
1:36:08 how do you become an expert if you never try to do the thing yourself?
1:36:14 One way I always learned is by trying things myself.
1:36:18 If you look at math textbooks and the solutions,
1:36:21 you learn something, but you learn actually better if you try first.
1:36:27 Then you appreciate the solution differently because you
1:36:29 know how to put it into your mental framework.
1:36:32 If LLMs are here all the time,
1:36:35 would you actually go to the length of struggling?
1:36:38 Would you be willing to struggle?
1:36:40 Struggle is not nice, right?
1:36:42 But if you use the LLM to do everything,
1:36:44 at some point you will never really take the next step,
1:36:47 and then you will maybe not get that unlock
1:36:50 that you would get as an expert using an LLM.
1:36:52 So, I think there's like a Goldilocks sweet spot where maybe the trick
1:36:58 here is you make dedicated offline time where you study two hours a day,
1:37:01 and the rest of the day use LLMs.
1:37:03 But I think it's important also for people to still invest in themselves,
1:37:07 in my opinion, to not just LLM everything.
1:37:11 Yeah, as a civilization, we each individually have to find that Goldilocks zone.
1:37:16 And in the programming context as developers.
1:37:18 Now, we've had this fascinating conversation
1:37:20 that started with pre-training and mid-training.
1:37:24 Let's get to post-training.
1:37:25 A lot of fun stuff in post-training.
1:37:28 So, what are some of the interesting ideas in post-training?
1:37:32 The biggest one from 2025 is
1:37:34 learning this reinforcement learning with verifiable rewards.
1:37:37 You can scale up the training there,
1:37:39 which means doing a lot of this kind of iterative generate-grade loop,
1:37:43 and that lets the models learn both interesting
1:37:47 behaviors on the tool use and software side.
1:37:49 This could be searching, running commands on their own and seeing the outputs,
1:37:53 and then also that training enables this inference time scaling very nicely.
1:37:57 And it just turned out that this paradigm was very nicely linked,
1:38:02 where this kind of RL training enables inference time scaling.
1:38:05 But inference time scaling could have been found in different ways.
1:38:07 So, it was kind of this perfect storm where the models change a lot,
1:38:10 and the way that they're trained is a major factor in doing so.
1:38:15 And this has changed how people approach post-training dramatically.
1:38:20 Can you describe RLVR, popularized by DeepSeek R1?
1:38:23 Can you describe how it works?
1:38:26 Yeah.
1:38:26 Fun fact: I was on the team that came up with the term RLVR,
1:38:29 which is from our Tulu 3 work before DeepSeek.
1:38:33 We don't take a lot of credit for being the people to popularize the scaling RL,
1:38:37 but as fun as what academics get,
1:38:40 as an aside, is the ability to name and influence the discourse,
1:38:44 because the closed labs can only say so much.
1:38:47 That one of the things you can do as an academic
1:38:49 is you might not have the compute to train the model,
1:38:51 but you can frame things in a way that ends up being described
1:38:55 as a community coming together around this RLVR term, which is very fun.
1:39:00 And then DeepSeek are the people that did the training breakthrough,
1:39:03 which is, they scaled the reinforcement learning.
1:39:06 You have the model generate answers and then
1:39:09 grade the completion if it was right,
1:39:11 and then that accuracy is your reward for reinforcement learning.
1:39:16 So reinforcement learning is classically an agent that acts in an environment,
1:39:20 and the environment gives it a state and a reward back,
1:39:24 and you try to maximize this reward.
1:39:26 In the case of language models,
1:39:28 the reward is normally accuracy on a set of verifiable tasks,
1:39:31 whether it's math problems or coding tasks.
1:39:34 And it starts to get blurry with things like factual domains.
1:39:38 That is also, in some ways, verifiable, or constraints on your instruction,
1:39:44 like respond only with words that start with A." All
1:39:47 of these things are verifiable in some way,
1:39:50 and the core idea of this is you find a lot more of these problems
1:39:55 that are verifiable and you let the model
1:39:56 try it many times while taking these RL steps, these RL gradient updates.
1:40:02 The infrastructure evolved from reinforcement learning from human feedback,
1:40:06 where in that era the score they were trying
1:40:10 to optimize was a learned reward model of human preferences.
1:40:13 So you kind of changed the problem domains
1:40:15 and that let the optimization go on to much bigger scales,
1:40:19 which kind of kickstarted a major change in what
1:40:22 the models can do and how people use them.
1:40:25 What kind of domains is RLVR amenable to?
1:40:28 Math and code are the famous ones,
1:40:30 and then there's a lot of work on what is called the rubrics,
1:40:34 which is related to a word people might have heard: LLM-as-a-judge.
1:40:38 For each problem, I'll have a set of problems in my dataset.
1:40:42 I will then have an LLM and ask it,
1:40:45 "What would a good answer to this problem look like?" And then you could
1:40:49 try the problem over and over again and assign a score based on this rubric.
1:40:54 That's not necessarily verifiable like math and code domains,
1:40:57 but this rubrics idea and other scientific problems that might
1:41:00 be a little bit more vague is where the attention is,
1:41:04 where they're trying to push this set of methods into these kind
1:41:07 of more open-ended domains so the models can learn a lot more.
1:41:11 I think that's called reinforcement learning with AI feedback, right?
1:41:14 That's the older term for it coined in Anthropic's Constitutional AI paper.
1:41:18 It's like a lot of these things come in cycles.
1:41:21 Also, just one step back for RLVR.
1:41:24 I think the interesting thing here is that you
1:41:27 ask the LLM a, let's say, math question, and then you know the correct answer,
1:41:31 and you let the LLM, as you said, figure it out.
1:41:35 How it does it—you don't constrain it much.
1:41:37 There are some constraints like "use the same language,
1:41:40 don't switch between Spanish and English."
1:41:42 But let's say you're pretty much hands-off.
1:41:45 You only give the question and the answer,
1:41:47 and then the LLM has the task to arrive at the right answer,
1:41:51 but the beautiful thing here is what happens in practice:
1:41:55 the LLM will do a step-by-step description,
1:41:57 like as a student or as a mathematician would derive the solution.
1:42:02 It will use those steps, and that helps the model to improve its own accuracy.
1:42:07 And then, like you said, the inference scaling.
1:42:11 Inference scaling loosely means spending more compute during inference,
1:42:16 and here the inference scaling is that the model would use more tokens.
1:42:22 In the DeepSeek R1 paper,
1:42:24 they showed the longer they train the model, the longer the responses are.
1:42:28 They grow over time.
1:42:29 They use more tokens, so it becomes more expensive.
1:42:32 It becomes expensive for simple tasks, but these explanations help accuracy.
1:42:36 There are also papers showing what the model explains does not
1:42:40 necessarily have to be correct or maybe it's unrelated to the answer,
1:42:44 but for some reason, it still helps the model that it is explaining.
1:42:48 And I think it's also—again, I don't want to anthropomorphize these LLMs,
1:42:52 but it's kind of like how we humans operate.
1:42:55 If there's a complex math problem in a math class, class,
1:42:59 you usually have a note paper and you do it step by step.
1:43:02 You cross out things.
1:43:03 And the model also self-corrects, and that was,
1:43:05 I think, the aha moment in the DeepSeek R1 paper.
1:43:08 They called it the aha moment because the model
1:43:10 itself recognized it made a mistake and then said, "Ah, I did something wrong,
1:43:13 let me try again." And I think that's just so cool that this falls out
1:43:18 of just giving it the correct answer and having it figure out how to do it,
1:43:22 that it kind of does in a sense what a human would do.
1:43:26 Although LLMs don't think like humans,
1:43:28 it's kind of like an interesting coincidence and it...
1:43:31 And the other nice side effect is it's
1:43:33 great for us humans often to see these steps.
1:43:36 It builds trust, but also we us humans to see these steps.
1:43:38 It builds trust, but also we learn and can double check things.
1:43:40 There's a lot in here.
1:43:40 I think some of the debate...
1:43:42 There's been a lot of debate this year on if the language models like these...
1:43:45 I think the aha moments are kind of fake
1:43:48 because in pre-training you essentially have seen the whole internet.
1:43:51 so you have definitely seen people explaining their work,
1:43:54 even verbally, like a transcript of a math lecture.
1:43:57 "You try this, oh, I messed this up."
1:43:58 And what RLVR is very good at doing is amplifying
1:44:02 these behaviors because they're very useful in enabling
1:44:04 the model to think longer and to check its work.
1:44:07 And I agree that it is very beautiful that this training kind of...
1:44:11 The model learns to amplify this in a way
1:44:13 that is so useful for the final answers being better.
1:44:17 I can give you also a hands-on example.
1:44:18 I was training the Qwen 3 base model with RLVR on MATH-500.
1:44:22 The base model had an accuracy of about 15%.
1:44:26 Just 50 steps, like in a few minutes with RLVR,
1:44:30 the model went from 15% to 50% accuracy.
1:44:33 And the model...
1:44:34 You can't tell me it's learning anything fundamentally about math
1:44:38 in- The Qwen example is weird because there've been two papers this year,
1:44:41 one of which I was on, about data contamination in Qwen
1:44:44 and specifically that they train on a lot of this special
1:44:47 mid-training phase that we should take a minute on, because it's
1:44:50 weird because they train on problems that are almost identical to MATH.
1:44:53 Exactly.
1:44:54 And so you can see that basically the RL,
1:44:57 it's not teaching the model any new knowledge about math.
1:44:59 You can't do that in 50 steps.
1:45:01 So the knowledge is already there,
1:45:02 in the pre-training, you're just unlocking it.
1:45:04 I still disagree with the premise because there's a lot
1:45:06 of weird complexities that you can't prove because one
1:45:10 of the things that points to weirdness is that if
1:45:12 you take the Qwen 3 so-called base model and you...
1:45:15 You could Google like "math dataset, Hugging Face",
1:45:18 and you could take a problem and what you do if you put it into Qwen 3 base...
1:45:23 All these math problems have words,
1:45:24 so it'd be like "Alice has five apples and takes one...
1:45:27 and gives three to whoever," and there are these word problems.
1:45:30 With these Qwen-based models, why people are suspicious of them is if you
1:45:34 change the numbers but keep the words- Qwen will produce,
1:45:38 without tools, will produce a very
1:45:41 high accuracy decimal representation of the answer, which means there's some...
1:45:45 At some time, it was shown problems that were almost identical to the test set,
1:45:50 and it was using tools to get a very high precision answer,
1:45:53 but a language model without tools will never actually have this.
1:45:57 So it's kind of been this big debate in the research community:
1:46:01 how much of these reinforcement learning papers that are
1:46:04 training on Qwen and measuring specifically on this math benchmark,
1:46:07 where there's been multiple papers talking about contamination,
1:46:10 is like, how much can you believe them?
1:46:12 And I think this is what caused the reputation of RLVR being about formatting,
1:46:15 because you can get these gains so quickly,
1:46:18 therefore it must already be in the model.
1:46:20 But there's a lot of complexity here that we...
1:46:22 It's not really like controlled experimentation, so we don't really know.
1:46:27 But if it weren't true, I would say distillation wouldn't work, right?
1:46:30 I mean, distillation can work to some extent, but the thing is that is, I think,
1:46:35 the biggest problem,
1:46:36 and I research this contamination because we don't know what's in the data.
1:46:38 Unless you have a new dataset, it is really impossible.
1:46:42 And the same, you mentioned the math dataset,
1:46:44 where you have a question and then answer and an explanation is given,
1:46:48 but then also even something simpler like MMLU,
1:46:51 which is a multiple-choice benchmark.
1:46:53 If you just change the format slightly, like, I don't know,
1:46:57 if you use a dot instead of a parenthesis
1:47:00 or something like that, the model accuracy will vastly differ.
1:47:05 I think that that could be like a model issue rather than a general issue.
1:47:09 It's not even malicious by the developers of the LLM, like, "Hey,
1:47:11 we want to cheat at that benchmark." It has seen something at some point.
1:47:14 I think the only fair way to evaluate an LLM is to have
1:47:18 a new benchmark that is after the cutoff date when the LLM was deployed.
1:47:22 Can we lay out what would be the recipe
1:47:25 of all the things that go into post-training?
1:47:28 And you mentioned RLVR was a really exciting, effective thing.
1:47:32 Maybe we should elaborate.
1:47:34 RLHF still has a really important component to play.
1:47:37 What kind of other ideas are there on post-training?
1:47:40 I think you can kind of take this in order.
1:47:41 I think you could view it as what made o1,
1:47:45 which is this first reasoning model, possible, or what will the latest model be?
1:47:49 And they actually...
1:47:51 You're going to have similar interventions
1:47:53 at these, where you start with mid-training,
1:47:56 and the thing that is rumored to enable
1:47:59 o1 and similar models is really careful data curation,
1:48:02 where you're providing a broad set of what is called reasoning traces,
1:48:07 which is just the model generating words
1:48:10 in a forward process that is reflecting,
1:48:13 like breaking down a problem into intermediate steps and trying to solve them.
1:48:16 So at mid-training, you need to have data that is similar
1:48:20 to this to make it so that when you move into post-training,
1:48:23 primarily with these verifiable rewards, it can learn.
1:48:27 And then what is happening today is you're figuring out which problems
1:48:32 to give the model and how out which problems to give the model
1:48:35 and how long you can train it for and how much inference
1:48:37 you can enable the model to use when solving these verifiable problems.
1:48:41 So as models get better,
1:48:43 certain problems models get better, certain problems are no longer...
1:48:47 The model will solve them 100% of the time,
1:48:48 and therefore there's very little signal in this.
1:48:51 If we look at the GRPO equation, this one is famous for this because essentially
1:48:55 the reward given to the agent is based on how
1:48:59 good a given action—an action is a completion—is
1:49:02 relative to the other answers to that same problem.
1:49:05 So if all the problems get the same answer,
1:49:07 there's no signal in these types of algorithms.
1:49:08 So what they're doing is they're finding harder problems,
1:49:11 which is why you hear about things like scientific domains,
1:49:14 where it's so hard to get anything right.
1:49:17 If you have a lab or something,
1:49:19 it just generates so many tokens or much harder software problems.
1:49:22 So the frontier models are all pushing into these harder domains when they
1:49:26 can train on more problems and the model will learn more skills at once.
1:49:30 The RLHF link to this is that RLHF has been
1:49:33 and still is kind of like the finishing touch on the models,
1:49:36 where it makes the models more useful
1:49:38 by improving the organization or style or tone.
1:49:41 There are different things that resonate with different audiences,
1:49:43 like some people like a really quirky model
1:49:45 and RLHF could be good at enabling that personality,
1:49:49 and some people hate this markdown bulleted list thing that the models do,
1:49:54 but it's actually really good for quickly parsing information.
1:49:57 In RLHF, this human feedback stage is really great for putting
1:50:02 this into the model at the end of the day.
1:50:05 It's what made ChatGPT so magical for people.
1:50:07 And that use has actually remained fairly stable.
1:50:10 This formatting can also help the models
1:50:14 get better at math problems, for example.
1:50:16 So it's like the border between style and formatting,
1:50:20 and like the method that you use to answer a problem is
1:50:24 actually all very closely linked in terms of when you're training these models,
1:50:29 which is why RLHF can still make a model better at math,
1:50:32 but these verifiable domains are a much more direct process
1:50:35 to doing this because it makes more sense with the problem formulation,
1:50:39 which is why it ends up all forming together.
1:50:42 But to summarize, it's like mid-training is give
1:50:44 the model the skills it needs to then learn.
1:50:47 RL with verifiable rewards is letting the model try a lot of times,
1:50:52 so put a lot of compute into trial-and-error learning across hard problems.
1:50:55 And then RLHF would be like finishing the model,
1:50:57 making it easy to use and kind of just rounding the model out.
1:51:02 Can you comment on the amount of compute required for RLVR?
1:51:06 It's only gone up and up.
1:51:08 I think Ilya Sutskever was famous for saying they
1:51:10 use a similar amount of compute for pre-training and post-training.
1:51:12 Back to the scaling discussion,
1:51:15 they involve very different hardware for scaling.
1:51:17 Pre-training is very compute-bound, which is like this FLOPs discussion,
1:51:20 which is just how many matrix multiplications can you get through at once.
1:51:24 And because with RL you're generating these answers,
1:51:26 you're trying the model in real-world environments,
1:51:29 it ends up being much more memory-bound because
1:51:31 you're generating long sequences and the attention mechanisms
1:51:35 have this behavior where you get a quadratic
1:51:38 increase in memory as you're getting to longer sequences.
1:51:42 So the compute becomes very different.
1:51:43 In pre-training we would talk about a model—if we
1:51:46 go back to like the Biden administration executive order,
1:51:48 it's like 10 to the 25th FLOPs to train a model.
1:51:51 If you're using FLOPs in post-training,
1:51:53 it's a lot weirder because the reality is just like:
1:51:56 how many hours are you allocating?
1:51:58 How many GPUs for?
1:51:59 And I think in terms of time, the RL compute is getting much closer because
1:52:04 you just can't put it all into one system.
1:52:06 Pre-training is so computationally dense where all the GPUs
1:52:09 are talking to each other and it's extremely efficient,
1:52:11 whereas RL has all these moving parts and can take
1:52:13 a long time to generate a sequence of 100,000 tokens.
1:52:17 If you think about GPT-5.2 Pro taking an hour, it's like,
1:52:20 what if your training run has a sample for an hour
1:52:23 and you have to make sure that's handled efficiently?
1:52:25 So I think in GPU hours or just wall-clock hours,
1:52:29 the RL runs are probably approaching the same number of days as pre-training,
1:52:32 but they probably aren't using as many GPUs at the same time.
1:52:36 There are rules of thumb where in labs you don't want
1:52:40 your pre-training runs to last more
1:52:41 than a month because they fail catastrophically.
1:52:43 And if you are planning a huge cluster to be
1:52:46 held for two months and then it fails on day 50,
1:52:49 the opportunity costs are just so big.
1:52:52 So people don't want to put all their eggs in one basket.
1:52:56 GPT-4 was the ultimate YOLO run, and nobody ever wanted to do it before,
1:53:01 where it took three months to train and everybody was shocked that it worked.
1:53:04 I think people are a little bit more cautious and incremental now.
1:53:07 So RLVR is more, let's say,
1:53:10 unlimited in how much you can train or still get benefit, where RLHF,
1:53:14 because it's a preference tuning,
1:53:15 you reach a certain point where it doesn't really
1:53:17 make sense to spend more RL budget on that.
1:53:20 So just a step back with preference tuning:
1:53:22 there are multiple people that can give multiple, let's say,
1:53:26 explanations for the same thing and they can both be correct,
1:53:29 but at some point you learn a certain style
1:53:31 and it doesn't make sense to iterate on it.
1:53:34 My favorite example is: if relatives ask me what laptop they should buy,
1:53:38 I give them an explanation or ask,
1:53:41 "What is your use case?" They, for example, prioritize battery life and storage.
1:53:46 Other people like us, for example, we would prioritize RAM and compute.
1:53:50 Both answers are correct, but different people require different answers.
1:53:55 With preference tuning, you're trying to average somehow.
1:53:58 You are asking the data labelers to give you, not the right,
1:54:02 but the preferred answer and then you train on that.
1:54:04 But at some point you learn that average preferred answer.
1:54:07 And there's no reason to keep training longer on it because it's just a style,
1:54:13 whereas with RLVR, you let the model
1:54:16 solve more and more complex, difficult problems.
1:54:19 So I think it makes more sense to allocate more budget long-term to RLVR.
1:54:25 Also, right now we are in an RLVR 1.0 blend where
1:54:31 it's still that simple thing where we have a question and answer,
1:54:34 but we don't do anything with the stuff in between.
1:54:38 There were multiple research papers, also by Google for example,
1:54:42 on process reward models that also give
1:54:44 scores for the explanation—how correct is the explanation.
1:54:47 And I think that will be the next thing, let's say RLVR 2.0 for this year,
1:54:53 focusing in between question and answer, like how to leverage that information,
1:54:58 the explanation, to help it get better accuracy.
1:55:01 So that's one angle.
1:55:03 And there was a DeepSeek Math-V2 paper where
1:55:07 they also had interesting inference scaling there where,
1:55:11 first, they had developed models that grade themselves, a separate model.
1:55:16 And I think that will be one aspect.
1:55:18 And the other, like Nathan mentioned,
1:55:20 it will be for RLVR branching into other domains.
1:55:24 The place where people are excited are value functions, which is pretty similar.
1:55:28 So process reward models are kind of like...
1:55:30 Process reward models assign how good something is
1:55:34 at each intermediate step in a reasoning process,
1:55:37 where value functions apply value to every token the language model generates.
1:55:41 Both of these have been largely unproven
1:55:44 in the language modeling and reasoning model era.
1:55:48 People are more optimistic about value functions for whatever reason now.
1:55:52 I think process reward models were tried a lot more in this pre-o1,
1:55:57 pre-reasoning model era, and a lot of people had a lot of headaches with them.
1:56:00 So I think a lot of it is human nature...
1:56:03 Value models have a very deep history in reinforcement learning.
1:56:06 They're one of the first things core to deep reinforcement learning existing,
1:56:10 is training value models.
1:56:12 So right now people are excited about trying value models,
1:56:16 but there's very little proof.
1:56:18 And there are negative examples in trying to scale up process reward models.
1:56:22 These things don't always hold in the future.
1:56:24 We came to this discussion by talking about scaling.
1:56:27 The simple way to summarize what you're saying
1:56:29 is you don't want to do too much RLHF, where the signal doesn't scale.
1:56:33 People have worked on RLHF for language models for years,
1:56:36 especially with intense interest after ChatGPT.
1:56:39 And the first release of a reasoning model trained with RLVR, OpenAI's o1,
1:56:44 had a scaling plot where if you increase training compute logarithmically,
1:56:47 you get a linear increase in evaluations.
1:56:50 This has been reproduced multiple times.
1:56:52 DeepSeek had a plot like this.
1:56:54 But there's no scaling law for RLHF where if you log-increase the compute,
1:56:58 you get performance.
1:56:59 In fact, the seminal scaling paper for RLHF
1:57:02 is scaling laws for reward model over-optimization.
1:57:05 So that's a big line to draw with RLVR and the methods we have now.
1:57:10 In the future, they will follow this scaling paradigm:
1:57:13 where you can let the best runs run for an extra 10x and you get performance,
1:57:18 but you can't do this with RLHF.
1:57:20 And that is just going to be field-defining in how people approach them.
1:57:24 While I'm a shill for people to academically do RLHF,
1:57:28 to do the best RLHF you might not need the extra 10 or 100x of compute,
1:57:34 but to do the best RLVR you do.
1:57:38 I think there's a seminal paper from a Meta internship.
1:57:42 It's called something like "The Art of Scaling Reinforcement Learning
1:57:46 with Language Models." What they describe as a framework is Scale-RL.
1:57:50 Their incremental experiment was like 10,000 V100 hours,
1:57:54 which is like thousands or tens of thousands of dollars per experiment.
1:57:58 They do a lot of them, and This cost is not accessible to the average academic,
1:58:04 which is a hard equilibrium where it's trying
1:58:08 to figure out how to learn from each community.
1:58:11 I was wondering if we could take a bit
1:58:13 of a tangent and talk about education and learning.
1:58:16 If you're someone listening to this who's
1:58:20 a smart person interested in programming and AI,
1:58:23 I presume building something from scratch is a good beginning.
1:58:28 So can you take me through what you would recommend people do?
1:58:32 I would personally start, like you said,
1:58:34 implementing a simple model from scratch that you can run on your computer.
1:58:38 The goal is not, when you build a model from scratch,
1:58:41 to have something for every day use.
1:58:43 It's not going to be your personal
1:58:46 assistant replacing an existing open-weight model or ChatGPT.
1:58:49 It's to see what exactly goes into the LLM, what comes out,
1:58:53 and how the pre-training works on your own computer, preferably.
1:58:59 Then you learn about pre-training,
1:59:01 supervised fine-tuning, and the attention mechanism.
1:59:03 You get a solid understanding of how things work,
1:59:06 but at some point you reach a limit, because small models can only do so much.
1:59:11 The problem with learning about LLMs at scale is
1:59:14 that it's exponentially more complex to make a larger model,
1:59:18 because the model isn't just larger—you have
1:59:21 to shard your parameters across multiple GPUs.
1:59:24 Even for the KV cache, there are multiple ways to implement it.
1:59:27 One is just to understand how it works, just to grow the cache.
1:59:31 You grow it step-by-step by, let's say, concatenating lists,
1:59:35 but then that wouldn't be optimal on GPUs.
1:59:39 You would pre-allocate a tensor and then fill it in.
1:59:42 But that adds another 20 or 30 lines of code.
1:59:45 And for each thing, you add so much code.
1:59:47 The goal with the book is basically to understand how the LLM works.
1:59:51 It's not going to be a production-level LLM,
1:59:53 but once you have that, you can understand the production-level LLM.
1:59:56 So you're trying to always build an LLM that's going to fit on one GPU?
2:00:00 Yes.
2:00:00 Most of them do.
2:00:02 I have some bonus materials on some MoE models.
2:00:04 One or two of them may require multiple GPUs,
2:00:08 but the goal is to have it on one GPU.
2:00:10 And the beautiful thing is, you can self-verify.
2:00:13 It's almost like RLVR.
2:00:14 When you code these from scratch,
2:00:17 you can take an existing model from the Hugging Face Transformers library.
2:00:21 The library is great, but if you want to learn about LLMs,
2:00:26 it's not the best place to start because the code
2:00:28 is so complex to fit so many use cases.
2:00:32 Because people use it in production,
2:00:34 it has to be really sophisticated, really intertwined, and hard to read.
2:00:38 It's not linear.
2:00:39 It started as a fine-tuning library,
2:00:40 and then it grew to be the standard representation of every model architecture.
2:00:45 Hugging Face is the default place to get a model,
2:00:48 and Transformers is the software.
2:00:49 It enables it so people can easily load a model and do something basic with it.
2:00:57 And all frontier labs that have open-weight
2:00:59 models have a Transformers version of it, like from DeepSeek to gpt-oss-120b.
2:01:03 That's the canonical weight format you can load.
2:01:06 But even even Transformers, the library, is not used in production.
2:01:10 People use SGLang or vLLM, and it adds another layer of complexity.
2:01:15 We should say that the Transformers library has like 400 models.
2:01:19 So it's the one library that tries to implement a lot of LLMs,
2:01:22 and so you have a huge codebase, basically.
2:01:25 It's huge.
2:01:26 It's like, I don't know, maybe millions—- That's crazy.
2:01:30 hundreds of thousands of lines of code.
2:01:33 Understanding the part you want to understand
2:01:35 is finding the needle in the haystack.
2:01:36 But what's beautiful is you have a working implementation,
2:01:39 so you can work backwards.
2:01:40 What I would recommend doing, or what I also do,
2:01:44 is if I want to understand, for example, how OLMo is implemented,
2:01:47 I would look at the weights in the model hub, the config file,
2:01:50 and then you can see, "Oh, they used so many layers.
2:01:53 They use, let's say,
2:01:55 Group Query Attention or Multi-Head Attention in that case."
2:01:58 Then you see all the components in a human-readable, 100-line config file.
2:02:01 And then you start, let's say, with your GPT-2 model and add these things.
2:02:05 The cool thing here is you can then load
2:02:08 the pretrained weights and see if they work in your model.
2:02:12 You want to match the same output that you get with a Transformer model,
2:02:15 and then you can use that, basically
2:02:18 as a verifiable reward to make your architecture correct.
2:02:21 Sometimes it takes me a day.
2:02:23 With OLMo 3, the challenge was RoPE for the position embeddings.
2:02:26 They had a YaRN extension and there was some custom scaling there,
2:02:32 and I couldn't quite match these things.
2:02:35 In this struggle, you kind of understand things.
2:02:38 At the end, you know you have it correct because you can unit test it.
2:02:42 You can check against the reference implementation.
2:02:44 I think that's one of the best ways to learn, really.
2:02:48 To basically reverse-engineer something.
2:02:51 I think that is something everyone interested
2:02:53 in getting into AI today should do.
2:02:56 That's why I liked your book.
2:02:58 I came to language models from the RL and robotics field.
2:03:01 I had never taken the time to just learn all the fundamentals.
2:03:06 This transformer architecture is so fundamental,
2:03:09 just as deep learning was in the past, and people need to do this.
2:03:14 I think where a lot of people get overwhelmed is,
2:03:19 "How do I apply this to have impact or find
2:03:21 a career path?" Because language models
2:03:23 make this fundamental stuff so accessible,
2:03:26 and people with motivation will learn it.
2:03:29 Then it's like, "How do I get cycles on goal to contribute to research?"
2:03:34 I'm actually fairly optimistic because the field moves so fast that a lot
2:03:38 of times the best people don't fully solve a problem because there's
2:03:42 a bigger problem to solve that's very low-hanging fruit, so they move on.
2:03:46 I think that a lot of what I was trying to do in this RLHF book is
2:03:51 take post-training techniques and describe how people think about
2:03:54 them influencing the model and what people are doing.
2:03:57 Then it's remarkable how many things I
2:04:01 just think people stop studying or don't pursue.
2:04:05 I think people trying to go narrow after doing the fundamentals is good,
2:04:08 and then reading the relevant papers and being engaged in the ecosystem.
2:04:14 It's like you actually...
2:04:16 actually...
2:04:16 The proximity that random people online have
2:04:19 to the leading researchers—no one knows who all the...
2:04:23 The anonymous accounts on X and ML are very popular,
2:04:26 and no one knows who all these people are.
2:04:28 It could just be random people that study this stuff deeply,
2:04:31 especially with the AI tools.
2:04:32 To just be like, "I don't understand this, keep
2:04:34 digging into it," is a very useful thing.
2:04:36 But there's a lot of research areas that maybe
2:04:39 have three papers that you need to read,
2:04:42 and then one of the authors will probably email you back.
2:04:45 But you have to put in a lot
2:04:47 of effort into these emails to understand the field.
2:04:50 I think it would be for a newcomer easily weeks of work
2:04:53 to feel like they can truly grasp what is a very narrow area,
2:04:57 but I think going narrow after you have the fundamentals will be
2:05:00 very useful to people because I've become very interested in character training,
2:05:05 which is how you make the model funny or sarcastic or serious,
2:05:11 and what do you do to the data to do this?
2:05:14 A student at Oxford reached out to me and was like,
2:05:16 "Hey, I'm interested in this," and I advised him.
2:05:18 And that paper now exists.
2:05:20 There's like two or three people in the world that were very interested in this.
2:05:25 He's a PhD student, which gives him an advantage, but for me,
2:05:28 that was a topic I was waiting for someone to be like,
2:05:31 "Hey, I have time to spend cycles on this." I'm sure
2:05:33 there's a lot more very narrow things where you're just like,
2:05:36 "It doesn't make sense that there was no answer to this." I
2:05:38 think it's just there's so much information coming that people are like,
2:05:42 "I can't grab onto any of these," but if you just stick in an area,
2:05:46 I think there's a lot of interesting things to learn.
2:05:48 Yeah, I think you can't try to do it all
2:05:51 because it would be very overwhelming and you would burn out.
2:05:53 For me, for example, I haven't kept up with computer vision in a long time;
2:05:57 I just focused on LLMs.
2:05:58 But coming back to your book,
2:06:00 I think this is a really great book and a really good
2:06:03 bang for the buck because if you want to learn about RLHF,
2:06:06 I wouldn't go out there and read RLHF papers because
2:06:08 you would be spending two years—- Some of them contradict.
2:06:11 I've just edited the book, and there's no chapter where I had to be like,
2:06:15 "X papers say one thing and Y papers say another,
2:06:18 and we'll see what comes out to be
2:06:21 true."- Just to go through the table of contents,
2:06:23 what are some ideas we might have missed in the bigger picture of post-training?
2:06:26 First of all, you did the problem setup,
2:06:28 training overview, what are preferences, preference data,
2:06:31 and the optimization tools, reward modeling, regularization,
2:06:35 instruction tuning, rejection sampling, and reinforcement learning.
2:06:42 Then, Constitutional AI and AI feedback, reasoning,
2:06:46 and inference-time scaling to use in function calling,
2:06:49 synthetic data and distillation,
2:06:51 evaluation, and then an open questions section, over-optimization,
2:06:54 style and information, and then product UX, character and post-training.
2:06:59 What are some ideas worth mentioning that connect
2:07:03 both the educational and the research components?
2:07:06 You mentioned character training, which is pretty interesting.
2:07:08 Character training is interesting because there's so little on it.
2:07:10 We talked about how people engage with these models.
2:07:13 We feel good using them because they're positive,
2:07:16 but that can go too far; it can be too positive.
2:07:19 And it's like, essentially, it's:
2:07:20 How do you change your data and decision-making
2:07:23 to make it exactly what you want?
2:07:26 And like, OpenAI has this thing called a model spec,
2:07:29 which is essentially their internal guideline
2:07:31 for what they want the model to do, and they publish this to developers.
2:07:35 So, essentially, you can know what is a failure
2:07:38 of OpenAI's training—where they have the intentions and they
2:07:41 haven't met them yet— versus what is something that they
2:07:43 actually wanted to do and that you don't like.
2:07:46 And that transparency is very nice,
2:07:47 but all the methods for curating these documents and how
2:07:50 easy it is to follow them is not very well known.
2:07:53 I think the way the book is designed is that the RL
2:07:56 chapter is obviously what people want
2:07:57 because everybody hears about it with RLVR,
2:07:59 and it's the same algorithms and the same math,
2:08:01 but you can use it in very different documents.
2:08:05 So I think the core of RLHF is like how messy preferences are.
2:08:09 It's essentially a rehash of a paper I wrote years ago,
2:08:13 but this is essentially the chapter that'll tell
2:08:15 you why RLHF is never ever fully solvable because,
2:08:21 the way that even RL is set up, it assumes that preferences can be quantified
2:08:29 and that multiple preferences can be reduced to single values.
2:08:33 And I think it relates in the economics
2:08:35 literature to the Von Neumann-Morgenstern utility theorem,
2:08:38 and that is the chapter where all of that philosophical, economic,
2:08:43 and psychological context tells you what gets compressed into doing RLHF.
2:08:47 So it's like you have all of this and then later in the book it's like:
2:08:50 You use this RL map to make the number go up.
2:08:52 And I think that's why it'll be very
2:08:54 rewarding for people to do research on, because
2:08:57 quantifying preferences is something that humans have designed
2:09:01 a problem in order to make preferences studyable.
2:09:04 But there's kind of fundamental debates, like, an example is in a language model
2:09:09 response you have different things you care about, like accuracy or style.
2:09:12 And when you're collecting the data, they all get compressed into: "I like
2:09:16 this more than another." And that is happening,
2:09:19 and there's a lot of research in other areas
2:09:22 of the world that go into how you should actually do this.
2:09:26 I think social choice theory is the subfield
2:09:30 of economics around how you should aggregate preferences.
2:09:33 And I went to a workshop that published a white paper
2:09:37 on: "How can you think about using social choice theory for RLHF?" So
2:09:41 I mostly would want people that get excited about the math
2:09:44 to come and find things where they could stumble into this broader context.
2:09:48 I think there's a fun thing:
2:09:49 I just keep a list of all the tech reports of reasoning models I like.
2:09:54 So in Chapter 14, where there's a short summary of RLVR,
2:09:57 there's just a gigantic table where I
2:09:59 list every single reasoning model that I like.
2:10:03 I think in education, a lot of it needs to be like, at this point,
2:10:07 what I like, because the language models are so good at the math.
2:10:11 For example, the famous paper, Direct Preference Optimization,
2:10:13 which is a much simpler way of solving the problem than RL.
2:10:17 The derivations in the appendix skip steps of math.
2:10:21 And for this book, I redid the derivations and I'm like,
2:10:24 "What the heck is this log trick that they use
2:10:26 to change the math?" But doing it with language models, they're like,
2:10:29 "This is the log trick." And I'm like,
2:10:31 "I don't know if I like this, that the math is so commoditized." I think
2:10:35 some of the struggle in reading this appendix-
2:10:38 ...and following the math is good for learning.
2:10:44 Yeah, we're returning to this often on the topic of education.
2:10:47 You both have brought up the word "struggle" quite a bit.
2:10:51 So there is value.
2:10:52 If you're not struggling as part of this process,
2:10:55 you're not fully following the proper process for learning.
2:10:59 proper process for learning, I suppose.
2:11:02 Some providers are working on models
2:11:04 for education designed to not give- actually, I haven't used them,
2:11:07 but I'd guess they're designed to not give all the information at once.
2:11:12 And make people work for it.
2:11:13 Training models to do this would be a wonderful contribution.
2:11:16 Where, like all of the stuff in the book, you had to reevaluate every decision.
2:11:19 decision for it- It's a great example.
2:11:21 There's a chance we work on it at Ai2, which I thought would be so fun.
2:11:26 It makes sense.
2:11:27 I did something like that the other day for video games.
2:11:30 Sometimes for pastime I play video games,
2:11:32 like I like- Video games with puzzles, like Zelda and Metroid.
2:11:36 And there's this new game where I really got stuck and was okay with it.
2:11:41 I don't want to struggle for two days, so I used an LLM.
2:11:45 But then you say, "Hey, please don't add spoilers.
2:11:47 Just, you know, I'm here and there.
2:11:49 What do I have to do next?" You can do the same thing for math where you say,
2:11:53 "Okay, I'm stuck at this point.
2:11:55 Don't give me the full solution,
2:11:57 but what is something I could try?" Where you carefully probe it.
2:12:01 But the problem here is I think it requires discipline.
2:12:05 Many people enjoy math,
2:12:06 but there are also a lot of people who need to do it for their homework,
2:12:11 and then it's like a shortcut.
2:12:13 We could develop an educational LLM, but other LLMs are still there,
2:12:17 and there's still a temptation to use the other LLMs.
2:12:20 I think many people in college understand the stuff they're passionate
2:12:23 about- about- ...they're self-aware and they understand it shouldn't be easy.
2:12:27 I think we just have to develop a good taste- ...talk about research taste,
2:12:33 school taste about stuff that you should
2:12:36 be struggling on- ...and stuff you shouldn't be.
2:12:38 It's tricky, because you don't have
2:12:40 good long-term vision sometimes you don't have
2:12:43 good long-term vision about what would be actually useful to you in your career.
2:12:48 But you have to develop that taste, yeah.
2:12:52 I was talking to my fiancee or friends about this, there's this brief
2:12:56 10-year window where all of the homework and all the exams could be digital.
2:13:00 Before that, everybody had to do all the exams
2:13:02 in blue books because there was no other way.
2:13:04 And now after AI, everyone's going to need to be
2:13:06 in blue books and oral exams because everyone could cheat so easily.
2:13:09 It's like this brief generation that had
2:13:11 a different education system where everything could be digital,
2:13:15 but you still couldn't cheat.
2:13:16 And now it's just going back.
2:13:18 It's just very funny.
2:13:21 You mention character training.
2:13:22 Just zooming out on a more general topic,
2:13:24 for that project how much compute was required?
2:13:28 And in general, to contribute as a researcher,
2:13:31 are there places where not too much compute is
2:13:35 required where you can actually contribute as an individual researcher?
2:13:39 For the character training thing, I think this research is built on fine-tuning
2:13:43 about 7 billion parameter models with LoRA,
2:13:46 which is essentially only fine-tuning a small
2:13:48 subset of the weights of the model.
2:13:51 I don't know exactly how many GPU hours that would take.
2:13:55 But it's doable.
2:13:56 Not doable for every academic.
2:13:57 The situation for some academics is so dire that the only
2:14:00 work you can do is doing inference where you have
2:14:02 closed models or open models and you get completions from them
2:14:05 and you can look at them and understand the models.
2:14:07 And that's very well-suited to evaluation,
2:14:09 where you want to be the best at creating representative
2:14:14 problems that the models fail on or show certain abilities,
2:14:17 which I think that you can break through with this.
2:14:21 I think that the top-end goal for a researcher working on evaluation,
2:14:25 if you want to have career momentum,
2:14:27 is that Frontier Labs pick up your evaluation.
2:14:30 You don't need to have every project do this.
2:14:32 But if you go from a small university with no compute and find something
2:14:36 that Claude struggles with, and then the next
2:14:39 Claude model has it in the blog post, there's your career rocket ship.
2:14:42 I think that's hard,
2:14:44 but if you want to scope the maximum possible impact with minimum compute,
2:14:48 it's something like that, which is just get very narrow
2:14:51 and it takes learning of where the models are going.
2:14:54 So you need to build a tool that tests where Claude 4.5 will fail.
2:14:59 If I'm going to start a research project,
2:15:02 I need to think where the models in eight months are going to be struggling.
2:15:06 But what about developing totally novel ideas?
2:15:08 This is a trade-off.
2:15:09 I think that if you're doing a PhD,
2:15:11 you could also be like, "It's too risky to work in language models.
2:15:15 I'm going way longer term," which is like what is— what is
2:15:19 the thing that's going to define language model development in 10 years?
2:15:22 I end up being a person that's pretty practical.
2:15:25 I mean, I went to my PhD where it was like, "I got into Berkeley.
2:15:28 Worst case, I get a master's,
2:15:29 and then I go work in tech." I'm very practical about it,
2:15:32 so I'm like the life afforded to people
2:15:36 to work at these AI companies, the amount of...
2:15:38 OpenAI's average compensation is over a million
2:15:40 dollars in stock a year per employee.
2:15:43 For any normal person in the US,
2:15:46 to get into this AI lab is transformative for your life.
2:15:49 So I'm pretty practical about it.
2:15:50 there's still a lot of upward mobility
2:15:52 working in language models if you're focused.
2:15:54 And look at these jobs.
2:15:55 But from a research perspective,
2:15:57 the transformative impact in these academic awards...
2:16:01 to be the next Yann LeCun is from not
2:16:04 working on— not caring about language model development very much.
2:16:07 It's a big financial sacrifice in that case.
2:16:09 So I work with some awesome students, and they're like,
2:16:12 "Should I go work at an AI lab?" And I'm like,
2:16:14 "You're getting a PhD at a top school.
2:16:16 Are you gonna leave to go to a lab?" I don't know.
2:16:19 If you go work at a top lab, I don't blame you.
2:16:22 Don't go work at some random startup that might go to zero.
2:16:24 But if you're going to OpenAI, I'm like, "It could be worth leaving a PhD
2:16:29 for."- Let's more rigorously think through this.
2:16:32 So where would you give a recommendation
2:16:34 for people to do a research contribution?
2:16:36 So the options are academia: get a PhD.
2:16:40 Spend five years publishing.
2:16:44 Compute resources are constrained.
2:16:46 There's— there's research labs that are more
2:16:50 focused on open-weight models, and working there.
2:16:57 Or closed frontier research labs.
2:17:01 So OpenAI, Anthropic, xAI, and so on.
2:17:04 The two gradients are: the more closed,
2:17:06 the more money you tend to get, but you also get less credit.
2:17:10 In terms of building a portfolio of things that you've done,
2:17:17 it's very clear what you have done as an academic.
2:17:20 Versus if you are going to trade this fairly
2:17:25 reasonable progression for being a cog in the machine,
2:17:28 which could also be very fun.
2:17:30 So I think it's very different career paths.
2:17:33 But the opportunity cost for being a researcher is
2:17:36 very high because PhD students are paid essentially nothing.
2:17:38 So it ends up rewarding people that have a fairly stable safety net,
2:17:42 and they realize that they can operate in the long term,
2:17:45 wanting to do very interesting work and get a very interesting job.
2:17:49 So it is a privileged position to be like,
2:17:53 "I'm gonna see out my PhD and figure it out
2:17:56 after because I want to do this." At the same time,
2:18:00 the academic ecosystem is getting bombarded by funding getting cut and stuff.
2:18:04 So there's just so many different trade-offs where I understand
2:18:06 plenty of people that are like, "I don't enjoy it.
2:18:08 I can't deal with this funding search.
2:18:10 My grant got cut for no reason by the government," or, "I don't know what's
2:18:15 gonna happen." So I think there's a lot
2:18:17 of uncertainty and trade-offs that, in my opinion,
2:18:20 favor just taking the well-paying job with meaningful impact.
2:18:24 It's not like you're getting paid to sit around at OpenAI.
2:18:27 You're building the cutting edge of things
2:18:29 that are— changing millions of people's relationship to tech.
2:18:35 But publication-wise, they're being more secretive, increasingly so.
2:18:38 So you're publishing less and less.
2:18:40 And so you are having a positive impact at scale,
2:18:44 but you're a cog in the machine.
2:18:48 I think it honestly hasn't changed that much.
2:18:51 I have been in academia.
2:18:53 I'm not in academia anymore.
2:18:55 wouldn't want to miss my time in academia.
2:18:57 But what I wanted to say before I get
2:18:59 to that is that I think it hasn't changed that much.
2:19:02 I was working in computational biology,
2:19:04 using AI or machine learning methods with collaborators,
2:19:08 and a lot of people went from academia directly to Google.
2:19:15 And I think it's the same.
2:19:17 Back then, professors were sad that their students went
2:19:21 into industry because they couldn't carry on their legacy.
2:19:25 I think it's the same.
2:19:26 It hasn't changed that much.
2:19:28 The only thing that has changed is the scale.
2:19:32 Cool stuff was always developed in industry that was closed.
2:19:36 You couldn't talk about it.
2:19:38 And I think the difference now is your preference.
2:19:42 Do you like to publish your work, or are you more in a closed lab?
2:19:47 That's one difference.
2:19:48 The compensation, of course, is another, but it's always been like that.
2:19:54 It depends on where you feel comfortable.
2:19:56 And nothing is forever.
2:19:58 Right now, there's a third option, which is launching a startup.
2:20:02 A lot of people are doing that.
2:20:05 It's a very risky move, but it can be a high-risk,
2:20:10 high-reward situation, whereas joining an industry lab is pretty safe.
2:20:14 You also have upward mobility.
2:20:17 I think once you've been at an industry lab, it's easier to find future jobs.
2:20:22 But then again, how much do you enjoy the team
2:20:28 and working on proprietary things versus how much you like publishing work?
2:20:33 I mean, publishing is stressful.
2:20:35 Acceptance rates at conferences can be arbitrary and very frustrating,
2:20:40 but it's high reward if you have a paper published.
2:20:43 You feel good because your name is on there.
2:20:46 It's a high accomplishment.
2:20:48 I feel like my friends who are professors seem happier
2:20:51 than those who work at a frontier lab, to be honest.
2:20:55 There's a grounding there.
2:20:57 The frontier labs definitely do this 9-9-6,
2:21:00 which is shorthand for working all the time.
2:21:03 Can you describe 9-9-6?
2:21:05 It's a culture invented, I believe, in China and adopted in Silicon Valley.
2:21:10 What is 9-9-6?
2:21:11 It's 9:00 AM to 9:00 PM,- Six days a week.
2:21:15 six days a week.
2:21:16 What is that, 72 hours?
2:21:18 Okay.
2:21:18 So, is this basically the standard in AI companies in Silicon Valley?
2:21:24 This kind of grind mindset.
2:21:27 Yeah, I mean, maybe not exactly like
2:21:28 that, but I think there is a trend towards it.
2:21:30 And it's interesting.
2:21:31 I think it almost flipped because when I was in in academia, I felt like that.
2:21:36 As a professor, you write grants, you teach, and you do research.
2:21:39 It's like three jobs in one,
2:21:41 and it's more than a full-time job if you want to be successful.
2:21:45 successful.
2:21:46 And I feel like now, like Nathan just said,
2:21:49 the professors, in comparison to a lab,
2:21:52 I think they have less pressure or workload than
2:21:55 at a frontier lab because—- I think they work a lot.
2:21:58 They're just so fulfilled.
2:21:59 By working with students— and having a constant runway
2:22:02 of mentorship and a mission that is very people-oriented,
2:22:05 I think in a era when things are moving very fast and are very chaotic,
2:22:09 it's very rewarding to people.
2:22:11 Yeah, and I think at a startup, it's this pressure.
2:22:14 It's like you have to make it.
2:22:16 And it's really important that people put in the time,
2:22:19 but it is really hard because you have to deliver constantly,
2:22:22 and I've been at a startup.
2:22:24 I had a good time, but I don't know if I could do it forever.
2:22:28 It's an interesting pace and it's exactly like we talked about in the beginning.
2:22:33 These models are leapfrogging each other,
2:22:35 and they are just constantly trying to take
2:22:38 the next step compared to their competitors.
2:22:40 It's just ruthless right now.
2:22:42 I think this leapfrogging nature and having
2:22:44 multiple players is actually an underrated
2:22:46 driver of language modeling progress where
2:22:48 competition is so deeply ingrained in people,
2:22:52 and these companies have intentionally created very strong cultures.
2:22:56 Like, Anthropic is known to be so culturally,
2:23:00 like, deeply committed and organized.
2:23:02 I mean, we hear so little from them,
2:23:05 and everybody at Anthropic seems very aligned.
2:23:07 And it's like being in a culture that is super tight and having this competitive
2:23:13 dynamic is a thing that's gonna make you
2:23:16 work hard and create things that are better.
2:23:20 But that comes at the cost of human capital,
2:23:22 which is like you can only do this for so long,
2:23:26 and people are definitely burning out.
2:23:28 I wrote a post on burnout as I've tread in and out of this myself,
2:23:33 especially trying to be a manager, full-mode training.
2:23:36 It's a crazy job doing this.
2:23:37 The book Apple in China by Patrick McGee,
2:23:40 he talked about how hard the Apple engineers
2:23:42 worked to set up the supply chains in China,
2:23:44 and he was like, they had "saving marriage" programs, and he told in a podcast,
2:23:49 he was like, "People died from this level of working hard." So I think
2:23:53 it's just like it's a perfect environment
2:23:56 for creating progress based on human expense, and there's gonna be a lot of...
2:24:02 the human expense is the 996 that we started this with, which is like—...
2:24:07 people do really grind.
2:24:08 I also read this book.
2:24:09 I think they had a code word for if someone had to go
2:24:12 home to spend time with their family to save the marriage, and it's crazy.
2:24:15 Then the colleagues said, "Okay, this is like red alert for this situation.
2:24:19 We have to let that person go home this weekend." But at the same time,
2:24:23 I don't think they were forced to work.
2:24:25 They were so passionate about the product,
2:24:27 I guess, that you get into that mindset.
2:24:29 And I had that sometimes as an academic,
2:24:32 but also as an independent person, I have that sometimes.
2:24:35 I overwork, and it's unhealthy.
2:24:37 I had back issues, I had neck issues,
2:24:39 because I did not take the breaks that I maybe should have taken.
2:24:43 But no one forced me to; it's because I wanted to work,
2:24:45 because it's exciting stuff.
2:24:46 That's what OpenAI and Anthropic are like.
2:24:47 They want to do this work.
2:24:49 Yeah, but there's also a feeling of fervor that's building,
2:24:53 especially in Silicon Valley, aligned with the scaling laws idea,
2:24:56 where there's this hype where the world will be transformed in a scale
2:25:00 of weeks and you want to be at the center of it.
2:25:03 And then, you know, I have this great fortune
2:25:07 of having conversations with a wide variety of human beings,
2:25:12 and from there I get to see all
2:25:14 these bubbles and echo chambers across the world.
2:25:17 It's fascinating to see how we humans form them.
2:25:19 And I think it's fair to say that Silicon Valley is a kind of echo chamber,
2:25:25 a kind of silo and bubble.
2:25:27 I think bubbles are actually really useful and effective.
2:25:31 It's not necessarily a negative thing because you could be ultra-productive.
2:25:34 It could be the Steve Jobs reality distortion field,
2:25:39 because you just convince each other that breakthroughs are imminent,
2:25:42 and by convincing each other of that, you make the breakthroughs imminent.
2:25:49 Byrne Hobart wrote a book classifying bubbles.
2:25:51 One of them is financial bubbles, which is like speculation, which is bad,
2:25:54 and the other one is for build-outs,
2:25:56 because it pushes people to build these things.
2:25:58 And I do think AI is in this, but I
2:26:01 worry about it transitioning to a financial bubble, which is- Yeah,
2:26:05 but also in the space of ideas,
2:26:07 that bubble—you are doing a reality distortion field,
2:26:12 and that means you are deviating from reality.
2:26:14 And if you go too far from reality while also working, you know, 996,
2:26:22 you might miss some fundamental aspects of the human experience,
2:26:26 including beyond Silicon Valley.
2:26:27 This is a common problem in Silicon Valley:
2:26:30 it's a very specific geographic area.
2:26:32 You might not understand the Midwest perspective,
2:26:34 the full experience of all the other humans
2:26:38 in the United States and across the world,
2:26:40 and you speak a certain way to each other,
2:26:42 you convince each other of a certain thing,
2:26:44 and that can get you into real trouble.
2:26:47 Whether AI is a big success and becomes a powerful technology or it's not,
2:26:53 in either trajectory you can get yourself into trouble.
2:26:56 So you have to consider all of that.
2:26:58 Here you are, a young person trying to decide
2:27:00 what you want to do with your life.
2:27:02 The thing that is...
2:27:03 I don't even really understand this, but the SF AI memes
2:27:07 have gotten to the point where "permanent underclass" was one of them,
2:27:11 which was the idea that the last six months of 2025 was
2:27:14 the only time to build durable value in an AI startup or model.
2:27:18 Otherwise, all the value will be captured by existing
2:27:21 companies and you will therefore be poor, which...
2:27:24 that's an example of the SF thing that goes so far.
2:27:28 I still think for young people going to be able to tap into it,
2:27:31 if you're really passionate about wanting to have an impact in AI,
2:27:35 being physically in SF is the most likely place where you're going to do this.
2:27:39 But it has has trade-offs.
2:27:42 I think SF is an incredible place, but there is a bit of a bubble.
2:27:46 And if you go into that bubble, which is extremely valuable, just get out also.
2:27:52 Read history books, read literature, visit other places in the world.
2:27:57 Twitter and Substack are not the entire world.
2:28:01 I would say, one of the people I worked with is moving to SF,
2:28:04 and it's like, I need to get him a copy of Season of the Witch,
2:28:07 which is a history of SF from 1960 to 1985,
2:28:10 which goes through the hippie revolution,
2:28:14 like all the gays taking over the city and that culture emerging,
2:28:19 and then the HIV/AIDS crisis and other things.
2:28:22 And it's just like, that is so recent,
2:28:24 and so much turmoil and hurt, but also love in SF.
2:28:28 And it's like, no one knows about this.
2:28:30 It's a great book, Season of the Witch.
2:28:31 I recommend it.
2:28:32 A bunch of my SF friends who get out recommended it to me.
2:28:37 And I think that's just like living there...
2:28:39 I lived there and I didn't appreciate this context, and it's just so recent.
2:28:46 Yeah.
2:28:47 Okay, let's...
2:28:48 We talked a lot about a lot of things.
2:28:52 Certainly about the things that were exciting last year.
2:28:56 But this year, One of the things you guys
2:28:59 mentioned that's exciting is the scaling of text diffusion models,
2:29:02 and just a different exploration of text diffusion.
2:29:04 Can you talk about what that is and what the possibility it holds?
2:29:09 So, different kinds of approaches than the current LMs?
2:29:13 Yeah, so we talked a lot about the transformer
2:29:16 architecture and the autoregressive
2:29:17 transformer architecture specifically, like GPT.
2:29:19 And it doesn't mean no one else is working on anything else.
2:29:23 So, people are always on the, let's say, lookout for the next big thing.
2:29:27 Because I think it would be almost stupid not to.
2:29:30 Because sure, right now, the transformer architecture is the thing,
2:29:33 and it works best, and there's, right now, nothing else out there.
2:29:37 But, you know, it's always a good idea to not put all your eggs into one basket.
2:29:41 So, people are developing other alternatives to the autoregressive transformer.
2:29:45 One of them would be, for example, text diffusion models.
2:29:49 And listeners may know diffusion models from image generation,
2:29:52 like Stable Diffusion popularized it.
2:29:54 There was a paper on generating images.
2:29:57 Back then, people used GANs, Generative Adversarial Networks.
2:30:00 And then there was this diffusion
2:30:02 process where you iteratively denoise an image,
2:30:04 and that resulted in really good quality images over time.
2:30:08 Stable Diffusion was a company.
2:30:09 Other companies build their own diffusion models.
2:30:11 And then people are now like, "Okay,
2:30:13 can we try this also for text?" Doesn't, you know,
2:30:16 make intuitive sense yet, because it feels like, okay,
2:30:18 it's not something continuous like a pixel that we can differentiate.
2:30:21 It's discrete text, so how do we implement that denoising process?
2:30:26 It's kind of similar to the BERT models by Google.
2:30:31 Like, when you go back to the original transformer,
2:30:33 they were the encoder and the decoder.
2:30:35 The decoder is what we are using right now in GPT and so forth.
2:30:39 The encoder is more like a parallel technique where
2:30:43 you have multiple tokens that you fill in in parallel.
2:30:47 GPT models, they do autoregressive generation,
2:30:49 completing the sentence one token at a time.
2:30:52 And in BERT models, you have a sentence that has gaps.
2:30:57 You mask them out, and then one iteration is filling in these gaps.
2:31:02 Text diffusion is kind of like that, where
2:31:04 you are starting with some random text, and then you are filling in the missing
2:31:10 parts or refining them iteratively over multiple iterations.
2:31:12 And the cool thing here is that this can do multiple tokens at the same time.
2:31:18 It's like the promise of having it more efficient.
2:31:21 Now, the trade-off is, of course, how good is the quality?
2:31:25 It might be faster, and now you have this dimension of the denoising process.
2:31:29 The more steps you do, the better the text becomes.
2:31:32 And people...
2:31:34 I mean, you can scale in different ways.
2:31:37 They try to see if that is maybe a valid alternative to the autoregressive
2:31:41 model in terms of giving you the same quality for less compute.
2:31:46 Right now, there are papers that suggest if you want to get the same quality,
2:31:51 you have to crank up the denoising steps, and then you end up spending the same
2:31:56 compute you would spend on an autoregressive model.
2:31:58 The other downside is, while being parallel sounds appealing,
2:32:01 some tasks are not parallel.
2:32:03 Like reasoning tasks or tool use, maybe where you have to ask a code
2:32:08 interpreter to give you an intermediate result.
2:32:10 That is tricky with diffusion models.
2:32:12 So, there are some hybrids, but the main idea is: how can we parallelize it?
2:32:16 It's an interesting avenue.
2:32:18 I think right now, there are mostly research models out there,
2:32:22 like LaMDA and some other ones.
2:32:24 I saw some by startups, some deployed models.
2:32:27 There is no big diffusion model at scale yet,
2:32:30 like on the Gemini or ChatGPT level.
2:32:32 But there was an announcement by Google,
2:32:36 a site where they said they are launching Gemini Diffusion,
2:32:39 and they put it into context of their Gemini Nano 2 model,
2:32:44 and they said basically:
2:32:45 for the same quality on most benchmarks, we can generate things much faster.
2:32:50 You mentioned what's next.
2:32:52 I don't think the text diffusion model is going to replace autoregressive LLMs,
2:32:55 but it will be something maybe for quick, cheap, at-scale tasks.
2:33:00 Maybe the free tier in the future will be something like that.
2:33:08 I think there are examples where it's already being used.
2:33:10 To paint an example of why this is better, for example,
2:33:13 when GPT-5 is taking 30 minutes to respond, it's generating one token at a time.
2:33:17 And this diffusion idea is essentially to generate all
2:33:20 of those tokens and the completion in one batch,
2:33:23 which is why it could be way faster.
2:33:25 And I think it could be suited for...
2:33:27 the startups I'm hearing about are code startups where you have a code base,
2:33:31 and you have somebody that's effectively "vibe coding," and they say,
2:33:34 "Make this change." And a code diff is essentially a huge reply from the model,
2:33:39 but it doesn't have to have that much external context,
2:33:42 and you can get it really fast by using these diffusion models.
2:33:45 One example I've heard is that they
2:33:47 use text diffusion to generate really long diffs,
2:33:50 because doing it with an autoregressive model would take minutes,
2:33:53 and that time for a user-facing product causes a lot of churn.
2:33:57 Every second, you lose a lot of users.
2:33:59 So, I think it's going to be this thing
2:34:01 where it's going to— ...grow and have some applications,
2:34:03 but I actually thought that different types of models were going
2:34:06 to be used for different things much sooner than they have been,
2:34:10 so I kind of trade off.
2:34:11 I think the tool-use point is the one
2:34:13 that's stopping them from being most general purpose because,
2:34:18 for Claude Code and ChatGPT search,
2:34:22 the autoregressive chain is interrupted with some external tool,
2:34:25 and I don't know how to do that with the diffusion setup.
2:34:29 So what's the future of tool use this year and then in the coming years?
2:34:32 Do you think there's going to be a lot of developments there,
2:34:35 and how that's integrated into the entire stack?
2:34:37 I do think right now, it's mostly on the proprietary LLM side,
2:34:41 but I think we will see more of that in the open-source tooling.
2:34:44 And I think it is a huge unlock because then you
2:34:48 can really outsource certain tasks from just memorization to actual— you know,
2:34:54 instead of having the LLM memorize what is 23 plus 5, just use a calculator.
2:34:59 So do you think that can help solve hallucination?
2:35:02 Not solve it, but reduce it.
2:35:03 So the LLM still needs to know when to ask for a tool call.
2:35:09 And the second one is, well, it doesn't mean the internet is always correct.
2:35:13 You can do a web search,
2:35:14 but let's say I asked who won the World Cup in, let's say,
2:35:18 1998; it still needs to find the right website and get the right information.
2:35:21 You can still go to the incorrect website and give me incorrect information.
2:35:25 So I don't think it will fully solve that, but it is improving it in that sense.
2:35:31 And so another cool paper earlier this year—I think it was December 31st,
2:35:36 so it's not technically 2026, but close—the recursive language model.
2:35:43 That's a cool idea to kind of take this even a bit further.
2:35:47 Just to explain, Nathan, you also mentioned earlier,
2:35:51 it's harder to do cool research in academia because of the compute budget.
2:35:54 If I recall correctly, they did everything with GPT-5,
2:35:57 so they didn't even use local models,
2:35:59 but the idea is, let's say you have a long-context task;
2:36:01 instead of having the LLM solve all of it in one shot or even in a chain,
2:36:06 you break it down into sub-tasks.
2:36:08 You have the LLM decide what is a good sub-task,
2:36:12 and then recursively call an LLM to solve that.
2:36:16 And I think something like that, adding tools—you know,
2:36:20 each one maybe you have a huge Q&A task,
2:36:23 so each one goes to the web and gathers information,
2:36:26 and then you pull it together at the end and stitch it back together.
2:36:29 I think there's going to be a lot of unlock using
2:36:33 things like that where you don't necessarily improve the LLM itself;
2:36:37 you improve how the LLM is used and what the LLM can use.
2:36:41 One downside right now with tool use is you
2:36:43 have to give the LLM permission to use tools.
2:36:46 And that will take some trust,
2:36:49 especially if you want to unlock things like having
2:36:51 an LLM answer emails for you—or not even answer,
2:36:54 but just sort them for you or select them for you or something like that.
2:36:57 I don't know if I would today give an LLM access to my emails, right?
2:37:01 I mean, this is a huge risk.
2:37:03 I think there's a cool...
2:37:04 one last point on the tool use thing.
2:37:06 I think that you hinted at this, and we've both come at this in our own ways,
2:37:10 is that the open versus closed models use
2:37:12 tools in very different ways, where open models,
2:37:15 people go to Hugging Face and download the model,
2:37:17 and then the person's going to be like,
2:37:18 "What tool do I want?" I don't know, Exa is my preferred search provider,
2:37:22 but somebody else might care for a different search startup.
2:37:25 Where you release a model,
2:37:26 it needs to be useful for multiple tools, for multiple use cases,
2:37:29 which is really hard because you're making a general reasoning engine model,
2:37:33 which is actually what gpt-oss-120b is good for.
2:37:36 But on the closed models,
2:37:38 you're deeply integrating the specific tool into your experience,
2:37:41 and I think that open models will struggle to replicate some
2:37:45 of the things that I like to do with closed models,
2:37:47 which will be like, you can reference a mix of public and private information.
2:37:51 And something that I keep trying every three to six months,
2:37:55 I try Claude Code on the web, which is just prompting a model to make
2:37:59 an update to some GitHub repository that I have.
2:38:02 And it's just like that set of secure cloud environments is just so nice
2:38:06 for just sending it off to do this thing and then come back to me,
2:38:10 and these will probably help define some of the local open and closed niches.
2:38:18 But I think initially, because there was such a rush to get tool use working,
2:38:22 the open models were on the back foot, which is kind of inevitable.
2:38:25 I think there's so much research, so many resources in these frontier labs,
2:38:28 but it will be fun when the open models solve
2:38:31 this because it's going to necessitate a bit more flexible
2:38:34 and potentially interesting model that might work with this recursive
2:38:37 idea to be an orchestrator and a tool use model,
2:38:41 so hopefully the necessity drives some interesting innovation there.
2:38:45 So, continual learning—this is a longstanding topic, important problem.
2:38:51 I think that increases in importance as the cost of training the models goes up.
2:38:56 So can you explain what continual learning is and how important it
2:38:59 might be this year and in the coming years to make progress?
2:39:03 This relates a lot to this kind of SF zeitgeist of, what is AGI,
2:39:07 which is Artificial General Intelligence,
2:39:08 and what is ASI, Artificial Superintelligence,
2:39:11 and what are the language models that we have today capable of doing?
2:39:15 I think the language models can solve a lot of tasks,
2:39:18 but a key milestone among the AI community
2:39:21 is essentially when AI could replace any remote worker,
2:39:25 taking in information and solving digital tasks and doing them.
2:39:29 And the limitation that's highlighted by people is that a language model
2:39:33 will not learn from feedback the same way that an employee does.
2:39:36 So if you hire an editor, the editor will mess up, but you will tell them.
2:39:41 And if you hired a good editor, they don't do it again.
2:39:43 But language models don't have this ability
2:39:45 to modify themselves and learn very quickly.
2:39:47 So the idea is, if we are going to actually get to something that is a true,
2:39:52 general adaptable intelligence that can go into any remote work scenario,
2:39:55 it needs to be able to learn quickly from feedback and on-the-job learning.
2:40:00 I'm personally more bullish on language models being
2:40:02 able to just provide them with very good context.
2:40:06 You said, maybe offline,
2:40:07 that you can write extensive documents to models where you say,
2:40:11 "I have all this information.
2:40:13 Here are all the blog posts I've ever written.
2:40:15 I like this type of writing.
2:40:17 My voice is based on this." But many people don't provide this to models,
2:40:20 and the models weren't designed to take this amount of context previously.
2:40:24 Agentic models are just starting.
2:40:26 So it's this kind of trade-off: do we need to update the weights of this model
2:40:31 with this continual learning thing to make them learn fast?
2:40:34 Or the counterargument is we just need
2:40:36 to provide them with more context and information,
2:40:38 and they will have the appearance of learning fast
2:40:40 by having a lot of context and being smart.
2:40:43 So we should mention the terminology here.
2:40:45 Continual learning refers to changing the weights continuously so
2:40:49 that the model adapts and adjusts based on the new incoming information,
2:40:57 doing so continually, rapidly, and frequently.
2:41:00 And then the thing you mentioned on the other side
2:41:03 of it generally will be referred to as in-context learning.
2:41:07 As you learn stuff, there's a huge context window.
2:41:11 You can just keep loading it with extra
2:41:13 information every time you prompt the system,
2:41:15 which I think both legitimately can be seen as learning.
2:41:21 It's just a different place where you're doing the learning.
2:41:24 I think, to be honest with you,
2:41:26 continual learning— updating weights— we already have that in different flavors.
2:41:30 If you think about how...
2:41:32 I think the distinction here is:
2:41:34 do you do that on a personalized custom model for each person,
2:41:39 or do you do it on a global model scale?
2:41:42 I think we have that already, going from GPT-5 to 5.1 and 5.2.
2:41:47 It's maybe not immediate, but it is a curated update,
2:41:50 a quick curated update where there was feedback about things they couldn't do,
2:41:54 feedback by the community.
2:41:55 They updated the weights, next model, and so forth.
2:41:58 So it is a flavor of that.
2:42:02 Another even finer-grained example is like RLVR; you run it, it updates.
2:42:08 The problem is you can't just do that for each person because
2:42:12 it would be too expensive to update the weights for each person,
2:42:15 and I think that's the problem.
2:42:16 Unless you get...
2:42:17 Even at OpenAI scale, building the data centers, it would be too expensive.
2:42:22 I think that is only feasible once you have something
2:42:25 on the device where the cost is on the consumer.
2:42:27 Like what Apple tried to do with the Apple Foundation models,
2:42:30 putting them on the phone, where they learn from experience.
2:42:34 A bit of a related topic,
2:42:37 but this kind of, maybe anthropomorphized term: memory.
2:42:42 What are different ideas for the mechanism of how
2:42:44 to add memory to these systems as we're increasingly seeing?
2:42:47 Personalized memory especially?
2:42:50 Right now, it's mostly basically stuffing things
2:42:53 into the context and then just recalling that.
2:42:56 But again, I think it's expensive because you have to—you can cache it,
2:43:03 but still you spend tokens on that.
2:43:06 And the second one is you can only do so much.
2:43:09 I think it's more like a preference or style.
2:43:11 I mean, a lot of people do that when they solve math problems.
2:43:14 You say it's way so you can add previous knowledge and stuff,
2:43:17 but you also give it certain preference prompts:
2:43:20 "do what I preferred last time," or something like that.
2:43:23 But it doesn't unlock new capabilities.
2:43:26 So for that, one thing people still use is LoRA adapters.
2:43:31 These are basically, instead of updating the whole weight matrix,
2:43:35 there are two smaller weight matrices that you kind
2:43:38 of have in parallel or overlays like the delta.
2:43:41 But yeah, you can do that to some extent, but then again, it is economics.
2:43:47 There were also papers, for example, LoRA learns less but forgets less.
2:43:53 It's like, there's no free lunch.
2:43:54 If you want to learn more,
2:43:56 you need to use more weights, but it gets more expensive.
2:43:58 And then again, if you learn more, you forget more,
2:44:01 and you have to find that Goldilocks zone basically.
2:44:05 We haven't really mentioned it much,
2:44:06 but implied in this discussion is context length also.
2:44:09 Is there a lot of innovation that's possible there?
2:44:13 I think the colloquially accepted thing is that it's
2:44:16 a compute and data problem where you can...
2:44:19 and sometimes small architecture things like attention variants.
2:44:23 We talked about hybrid attention models,
2:44:26 which is essentially if you have what looks
2:44:29 like a state space model within your transformer.
2:44:31 And those are better suited because you have
2:44:34 to spend less compute to model the furthest along token.
2:44:38 I think that, those aren't free because they have to be
2:44:43 accompanied by a lot of compute or the right data.
2:44:47 How many sequences of 100,000 tokens do you have in the world,
2:44:51 and where do you get these?
2:44:53 It just ends up being pretty expensive to scale them.
2:44:56 We've gotten pretty quickly to a million tokens of input context length.
2:45:00 I would expect it to keep increasing and get
2:45:03 to 2 million or 5 million this year, but I don't expect it to go to 100 million.
2:45:07 That would be like a true breakthrough,
2:45:09 and I think those breakthroughs are possible.
2:45:11 I think of the continual learning thing as a research problem where there could
2:45:15 be a breakthrough that just makes transformers
2:45:17 work way better at this and it's cheap.
2:45:20 These things could happen with so much scientific attention.
2:45:22 But turning the crank, it'll be consistent increases over time.
2:45:28 Looking at the extremes, I think there's, again, no free lunch.
2:45:30 So, the one extreme to make it cheap: you have, let's say,
2:45:33 an RNN that has a single state
2:45:35 where you save everything from the previous stuff.
2:45:37 It's like a specific fixed-size thing, so you never really grow the memory
2:45:43 because you are stuffing everything into one state,
2:45:46 but then the longer the context gets, the more information you forget because
2:45:50 you can't compress everything into one state.
2:45:53 Then on the other hand, you have the transformers,
2:45:56 which try to remember every token,
2:45:57 which is great sometimes if you want to look up specific information,
2:46:00 but very expensive because you have the KV cache that grows,
2:46:04 the dot product that grows.
2:46:05 But then, like you said, the Mamba layers—I mean,
2:46:08 they kind of have the same problem.
2:46:10 Like an RNN, you try to compress everything into one state;
2:46:12 you're a bit more selective there.
2:46:14 But then I think it's like this Goldilocks zone again.
2:46:17 With Nemotron 3, they found a good ratio of how many attention layers do you
2:46:22 need for the global information where everything
2:46:24 is accessible compared to having these compressed states.
2:46:27 And I think that's how we will scale more—by finding better,
2:46:32 let's say, ratios in the Goldilocks zone,
2:46:35 like between making computing cheap enough to run,
2:46:39 but then also making it powerful enough to be useful.
2:46:43 And one more plug here, the Recursive Language Model paper,
2:46:47 that is one of the papers that tries to kind of address the long context thing.
2:46:51 So what they found is essentially instead of stuffing everything
2:46:55 into this long context if you break it up into multiple smaller tasks,
2:46:59 so you save memory by having multiple smaller cores,
2:47:03 you can actually get better accuracy than
2:47:05 having the LLM try everything all at once.
2:47:08 I mean, it's a new paradigm.
2:47:10 We will see, you know, there might be other flavors of that.
2:47:13 So I think with that, we will still make improvement on long context,
2:47:17 but then also, like Nathan said, I think the problem is for pre-training itself,
2:47:20 we don't have as many long context documents as other documents.
2:47:24 So it's harder to study basically how LLMs
2:47:28 behave and stuff like that on that level.
2:47:31 There are some rules of thumb where
2:47:33 essentially you pre-train a language model, like OLMo.
2:47:35 we pre-trained at like 8K context length and then extended to 32K with training.
2:47:39 And there are some rules of thumb
2:47:41 where you're essentially doubling the training context length,
2:47:44 it takes like 2X compute,
2:47:45 and then you can normally like 2 to 4X the context length again.
2:47:50 So I think a lot of it ends up being
2:47:52 kind of compute bound at pre-training, which is in this...
2:47:55 Like we talked about,
2:47:56 everyone talks about this big increase in compute for the top labs this year,
2:47:59 and that should reflect in some longer context windows.
2:48:02 But I think on the post-training side, there are some more interesting things.
2:48:04 As we have agents, the agents are gonna manage this context on their own,
2:48:08 where now agents, people that use Claude Code a lot dread the compaction,
2:48:12 which is when Claude takes its entire full 100,000
2:48:14 tokens of work and compacts it into a bulleted list.
2:48:17 But what the next models will do—and I'm sure people are already
2:48:22 working on this—is essentially the model can control when it compacts and how.
2:48:26 So you can essentially train your RL algorithm where compaction is an action-
2:48:30 ...where it shortens the history and then the problem formulation will be,
2:48:34 "I want to keep the maximum evaluation scores that I
2:48:38 have gotten while the model compacts its history to the minimum
2:48:42 length." Because then you have the minimum amount of tokens
2:48:44 that you need to do this kind of compounding autoregressive prediction.
2:48:47 So there are actually pretty nice problem setups in this, where the...
2:48:51 Like these agentic models learn to use their context
2:48:54 in a different way than just plow forward.
2:48:57 One interesting recent example would be DeepSeek-V3.2,
2:49:00 where they had a sparse attention mechanism
2:49:03 where they have essentially a very efficient, small, lightweight indexer.
2:49:07 And instead of attending to all tokens, it selects:
2:49:10 "What tokens do I actually need?" I mean,
2:49:13 it almost comes back to the original idea of attention where you are selective,
2:49:17 but attention is always on, you have maybe zero weight on some of them,
2:49:21 but you use them all.
2:49:22 But they are even more like, "Let's just mask that out or not even do
2:49:27 that." And even with sliding window attention in OLMo,
2:49:30 that is also kind of that idea.
2:49:32 You have a rolling window where you keep it fixed,
2:49:34 because you don't need everything.
2:49:35 Occasionally, some layers you might, but it's wasteful.
2:49:38 But right now, I think, if you use everything,
2:49:40 you're on the safe side—it gives you the best
2:49:42 bang for the buck because you never miss information.
2:49:44 And I think this year will be more about figuring out,
2:49:48 like you said, how to be smarter about that.
2:49:51 Right now, people want to have the next state-of-the-art,
2:49:54 and the state-of-the-art happens to be the brute-force, expensive thing.
2:49:59 And then once you have that, as you said, keep that accuracy,
2:50:03 but let's see how we can do that cheaper now, with tricks.
2:50:07 Yeah.
2:50:07 All this scaling thing.
2:50:08 The reason we get the Claude 4.5 Sonnet model first is because you
2:50:13 can train it faster and you're not hitting these compute walls as soon.
2:50:16 They can just try a lot more things and get the model faster,
2:50:18 even though the bigger model is actually better.
2:50:22 I think we should say that there's a lot
2:50:23 of exciting stuff going on in the AI space.
2:50:25 My mind has recently been really focused on robotics.
2:50:29 Today, we almost entirely didn't talk about robotics.
2:50:33 There's a lot of stuff on image generation, video generation.
2:50:38 I think it's fair to say that the most
2:50:42 exciting research work in terms of the amount, intensity,
2:50:45 and fervor is in the LLM space, which is why I think it's justified for us
2:50:50 to focus on the LLMs that we're discussing.
2:50:53 But it'd be nice to bring in certain things that might be useful.
2:50:57 For example, world models— there's growing excitement on that.
2:51:00 Do you think there will be any use
2:51:03 in this coming year for world models in the LLM space?
2:51:06 Yes, I do think so.
2:51:09 Also with LLMs, what's interesting here is
2:51:12 that if we unlock more LLM capabilities,
2:51:14 it also automatically unlocks all the other
2:51:17 fields because it makes progress faster.
2:51:20 A lot of researchers and engineers use LLMs for coding.
2:51:24 So even if they work on robotics,
2:51:27 if you optimize these LLMs that help with coding, it pays off.
2:51:31 But then, yes, world models are interesting.
2:51:34 It's basically where you have the model run
2:51:37 a simulation of the world in a sense,
2:51:39 like a little toy thing of the real thing, which can,
2:51:43 again, unlock capabilities regarding data the LLM is not aware of.
2:51:49 It can simulate things.
2:51:50 And I think LLMs happen to work
2:51:55 well by pre-training and doing next-token prediction.
2:51:59 But we could do this even more sophisticatedly in a sense.
2:52:03 I think there was a paper by Meta, a paper called World Models.
2:52:09 So where they basically apply the concept of world models to LLMs again,
2:52:14 where instead of just having next-token prediction and verifiable rewards,
2:52:18 checking the answer correctness,
2:52:20 they also make sure the intermediate variables are correct.
2:52:23 You know, it's kind of like the model
2:52:25 is learning basically a code environment in a sense.
2:52:28 And I think this makes a lot of sense.
2:52:30 It's just expensive to do, but it is making things more sophisticated,
2:52:37 like modeling the whole thing, not just the result.
2:52:42 And so it can add more value.
2:52:45 I remember when I was a grad student, there is a...
2:52:51 competition called CASP, I think, where they do protein structure prediction.
2:52:56 They predict the structure of a protein that is not solved yet at that point.
2:53:02 So in a sense, this is actually great,
2:53:04 and I think we need something like that for LLMs also,
2:53:07 where you do the benchmark, but no one does.
2:53:09 You hand in the results, but no one knows the solution.
2:53:12 And then after the fact, someone reveals that.
2:53:14 But, AlphaFold, when it came out, it crushed this benchmark.
2:53:20 I mean, there were also multiple iterations, but I remember the first one.
2:53:25 I'm not an expert in that subject,
2:53:28 but the first one explicitly modeled the physical interactions of the...
2:53:32 You know, the physics of the molecule.
2:53:34 Also the angles, impossible angles.
2:53:35 And then in the next version,
2:53:37 I think they got rid of this, and just with brute force, scaling it up.
2:53:40 And I think with LLMs,
2:53:42 we are currently in this brute force scaling because it just happens to work.
2:53:45 But I do think at some point it might make sense to bring back this thing.
2:53:50 And I think with world models,
2:53:52 I think that is where I think that might be actually quite cool.
2:53:56 I mean, yeah.
2:53:57 And of course, also for robotics, which is completely unrelated to LLMs.
2:54:03 Yeah.
2:54:03 And robotics is very explicit.
2:54:05 So there's the problem of locomotion or manipulation.
2:54:08 Locomotion is much more solved, especially in the learning domain.
2:54:11 But there's a lot of value, just like with the initial protein folding systems,
2:54:14 bringing in the traditional model-based methods.
2:54:17 So you don't...
2:54:19 it's unlikely that you can just learn the manipulation or the whole body,
2:54:25 local manipulation problem end to end.
2:54:28 That's the dream.
2:54:28 But then you realize when you look at the magic
2:54:31 of the human hand and the complexity of the real world,
2:54:35 it's really hard to learn this all the way through,
2:54:37 the way I guess AlphaFold 2 didn't.
2:54:41 I'm excited about the robotic learning space.
2:54:43 I think it's collectively getting supercharged by all
2:54:46 the excitement and investment in language models generally,
2:54:49 where the infrastructure for training transformers,
2:54:52 which is a general modeling thing, is becoming world-class industrial tooling,
2:54:59 where wherever there was a limitation for robotics, it's just way better.
2:55:03 There's way more compute.
2:55:04 And then on top of that, they take these language models as kind
2:55:07 of central units where you can do
2:55:09 interesting explorative work around something that already works.
2:55:12 And then I see it emerging as, kind of like we talked about,
2:55:17 Hugging Face transformers and Hugging Face.
2:55:18 I think when I was at Hugging Face,
2:55:20 I was trying to get this to happen, but it was too early.
2:55:22 It's like these open robotic models on Hugging Face,
2:55:26 and having people be able to contribute data and fine-tune them.
2:55:29 I think we're much closer now that the investment in robotics and self-driving
2:55:33 cars is related and it enables this, where once you get to the point
2:55:37 where you can have this sort of ecosystem where somebody can download a robotics
2:55:41 model and maybe fine-tune it to their robot or share datasets across the world.
2:55:45 There's some work in this area like RTX,
2:55:48 I think it was a few years ago, where people are starting to do that.
2:55:52 But once they have this ecosystem, it'll look very different.
2:55:54 And then this whole post-ChatGPT boom is putting more resources
2:55:58 into that, which I think is a very good area for doing research.
2:56:02 This is also resulting in much better, more accurate,
2:56:05 and more realistic simulators being built,
2:56:07 closing the sim-to-real gap in the robotic space.
2:56:10 But, you know, you mentioned a lot of excitement
2:56:13 in the robotics space and a lot of investment.
2:56:16 The downside of that, which happens in hype cycles,
2:56:20 I personally believe, and most robotics people believe, that it's not...
2:56:24 Robotics is not going to be solved
2:56:27 at the time scale as being implicitly or explicitly promised.
2:56:32 And so what happens when there's all these robotics companies
2:56:36 that spring up and then they don't have a product that works?
2:56:41 Then there's going to be this kind of crash of excitement,
2:56:44 which is nerve-wracking.
2:56:45 Hopefully something else will come in and keep swooping in so
2:56:50 that the continued development of some of these ideas keeps going.
2:56:54 I think it's also related to the continual learning issue,
2:56:57 essentially, where the real world is so complex.
2:57:00 With LLMs, you don't really need to have something learn for the user,
2:57:05 because there are a lot of things everyone has to do.
2:57:08 Everyone maybe wants to, I don't know,
2:57:10 fix their grammar in their email or code or something like that.
2:57:14 It's more constrained, so you can kind of prepare the model for that.
2:57:17 But preparing the robot for the real world is harder.
2:57:20 I mean, you have the robotic foundation models,
2:57:22 and you can learn certain things like grasping things.
2:57:27 But then again, everyone's house is different.
2:57:30 It's so different, and that is, I think,
2:57:33 where the robot would have to learn on the job, essentially.
2:57:36 And that, I guess, is the bottleneck right now:
2:57:39 how to, customize it on the fly, essentially.
2:57:42 I don't think I can possibly understate the importance of the thing that doesn't
2:57:48 get talked about almost at all by robotics folks or anyone, which is safety.
2:57:53 All the interesting complexities we talk about learning,
2:57:55 all the failure modes and failure cases, everything we've been talking about
2:57:59 with LLMs—sometimes they fail in interesting ways.
2:58:01 All of that is fun and games in the LLM space.
2:58:06 In the robotic space, in people's homes,
2:58:09 across millions of minutes and billions of interactions,
2:58:14 you really are almost allowed to fail never.
2:58:17 When you have embodied systems that are put out there in the real world,
2:58:23 you just have to solve so many problems you never thought you'd
2:58:28 have to solve when just thinking about the general robot learning problem.
2:58:33 I'm so bearish on in-home learned robots for consumer purchase.
2:58:38 I'm very bullish on self-driving cars,
2:58:39 and I'm very bullish for robotic automation, e.g.,
2:58:43 like Amazon distribution where Amazon has built whole new
2:58:46 distribution centers designed for robots first rather than humans.
2:58:49 There's a lot of excitement in AI
2:58:51 circles about AI enabling automation and mass-scale manufacturing,
2:58:54 and I do think that the path to robots doing that is more reasonable,
2:58:57 where it's a thing that is designed and optimized to do
2:59:03 a repetitive task that a human could conceivably do but doesn't want to.
2:59:07 And then I'm much, but it's also going
2:59:10 to take a lot longer than people probably predict.
2:59:14 I think the leap from AI singularity to we
2:59:18 can now scale up mass manufacturing in the US because
2:59:21 we have a massive AI advantage is one that is
2:59:25 troubled by a lot of political and other challenging problems.
2:59:32 Let's talk about timelines, specifically timelines to AGI or ASI.
2:59:38 Is it fair, as a starting point,
2:59:40 to say that nobody really agrees on the definitions of AGI and ASI?
2:59:46 I kind of think there's a lot of disagreement,
2:59:48 but I've been getting pushback where a lot of people kind of say the same thing,
2:59:53 which is like a thing that could reproduce most digital economic work.
2:59:57 So, the remote worker is a fairly reasonable example.
3:00:01 And I think OpenAI's definition is somewhat related to that, which is like an AI
3:00:06 that can do a lot of economically valuable
3:00:08 tasks—which I don't really love as a definition,
3:00:11 but I think it could be a grounding point,
3:00:15 because language models today, while immensely powerful,
3:00:19 are not this remote worker drop-in.
3:00:21 And there are things that could be done
3:00:24 by an AI that are way harder than remote work,
3:00:27 which are like finding an unexpected
3:00:30 scientific discovery that you couldn't even posit,
3:00:32 which would be an example of something
3:00:34 that somebody says is an artificial superintelligence problem.
3:00:37 Or, taking in all medical records and finding
3:00:43 linkages across certain illnesses that people didn't know,
3:00:46 or figuring out that some common drug can treat some niche cancer.
3:00:50 They would say that that is a superintelligence thing.
3:00:52 So these are kind of natural tiers.
3:00:54 My problem with it is that it becomes deeply entwined
3:00:58 with the quest for meaning of AI and these religious aspects to it.
3:01:03 So there's different paths you can take it.
3:01:06 And I don't even know if the remote worker
3:01:09 is a good definition because what exactly is that?
3:01:12 I actually, I mean, I like...
3:01:14 I don't know if you like the originally titled AI27 report.
3:01:18 They focus more on code and research taste,
3:01:22 so the target there is the superhuman coder.
3:01:25 So they have several milestone systems:
3:01:28 Superhuman coders, superhuman AI researcher,
3:01:31 then superintelligent AI researcher,
3:01:33 and then the full ASI, artificial superintelligence.
3:01:37 But after you develop the superhuman coder, everything else follows quickly.
3:01:45 There, the task is to have fully autonomous, automated coding.
3:01:52 So any kind of coding you need to do
3:01:54 in order to perform research is fully automated.
3:01:57 And from there, humans would be doing AI research together with that system,
3:02:02 and they will quickly be able to develop
3:02:04 a system that can actually do the research for you.
3:02:07 That's the idea.
3:02:09 And initially their prediction was 2027, 2028,
3:02:12 and now they've pushed it back by three to four years to 2031 (mean prediction).
3:02:19 Probably my prediction is even beyond 2031, but at least you can,
3:02:24 in a concrete way, think about how
3:02:28 difficult it is to fully automate programming.
3:02:31 Yeah, I disagree with some of their presumptions
3:02:34 and dynamics on how it would play out,
3:02:36 but I think they did good work in the scenario-defining
3:02:39 milestones that are concrete and tell a useful story,
3:02:42 which is why the reach for this AI 2027 document transcended Silicon Valley.
3:02:47 It's because they told a good story and they
3:02:50 did a lot of rigorous work to do this.
3:02:53 I think the camp that I fall into is that AI is so-called
3:02:56 "jagged," which will be excellent at some things and really bad at some things.
3:03:00 I think that when they're close to this automated software engineer,
3:03:05 what it will be good at is traditional ML systems and frontend,
3:03:09 the model is excellent at; but distributed ML,
3:03:12 the models are actually quite bad at because there's
3:03:14 so little training data on doing large-scale distributed learning.
3:03:17 This is something we already see, and I think this will just get amplified.
3:03:21 And then it's kind of messier in these trade-offs,
3:03:24 like how you think AI research works and so on.
3:03:28 So you think basically a superhuman coder is almost unachievable,
3:03:31 because of the jagged nature of the thing,
3:03:33 you're just always going to have gaps in capabilities?
3:03:38 I think it's assigning completeness to something where
3:03:41 the models are already superhuman at some types of code.
3:03:44 I think that will continue.
3:03:46 And people are creative, so they'll utilize these incredible abilities to fill
3:03:50 in the weaknesses of the models and move really fast.
3:03:53 There'll always be this dance for a long time between
3:03:57 the humans enabling the thing that the model can't do.
3:04:00 And the best AI researchers are the ones that can enable this superpower.
3:04:04 And I think those lines lead to what we already see.
3:04:06 I think like Claude Code for building a website,
3:04:08 you can stand up a beautiful website in a few
3:04:10 hours or do data going to keep getting better,
3:04:12 and we'll pick up some new coding skills along the way.
3:04:17 Linking to what's happening in big tech,
3:04:23 this AI 2027 report leans into the singularity idea,
3:04:27 whereas I think research is messy,
3:04:30 social and largely in the data in ways that AI models can't process.
3:04:34 But what we do have today is really powerful and these tech companies
3:04:39 are all collectively buying into this with tens
3:04:41 of billions of dollars of investment.
3:04:43 So we are going to get some much better version of ChatGPT,
3:04:47 a much better version of Claude Code than we already have.
3:04:50 I think it's just hard to predict where that is going,
3:04:53 but the bright clarity of that future is why some of the most
3:04:57 powerful people in the world are putting so much money into this.
3:05:00 And I think it's just kind of small differences between
3:05:04 like—we don't actually know what a better version of ChatGPT is,
3:05:07 but also, can it automate AI research?
3:05:10 I would say probably not, at least in this timeframe.
3:05:14 Big tech is going to spend $100 billion much faster than
3:05:17 we get an automated AI researcher that enables an AI research singularity.
3:05:23 So you think your prediction would be— if this is
3:05:26 even a useful milestone or more than 10 years out?
3:05:31 I would say less than that on the software side,
3:05:33 but I think longer than that on things like research.
3:05:37 Well, let's just for fun try to imagine
3:05:40 a world where all software writing is fully automated.
3:05:43 Can you imagine that world?
3:05:46 By the end of this year,
3:05:48 the amount of software that'll be automated will be so high.
3:05:50 But it'll be things like trying to train a model with RL
3:05:54 and you need to have multiple bunches of GPUs communicating with each other.
3:05:59 That'll still be hard, but it'll be much easier.
3:06:02 One way to think about this— the full automation
3:06:04 of programming—is just thinking of lines of useful code written,
3:06:10 the fraction of that to the number of humans in the loop.
3:06:15 So presumably there'll be for a long
3:06:17 time humans in the loop of software writing.
3:06:19 It'll just be fewer and fewer relative to the amount of code written.
3:06:23 Right?
3:06:24 And the superhuman coder—I think the presumption there is it goes to zero,
3:06:29 the number of humans in the loop.
3:06:31 What does that world look like when the number
3:06:33 of humans in the loop is in the hundreds, not in the hundreds of thousands?
3:06:39 I think software engineering will be driven
3:06:41 more to system design and goals of outcomes,
3:06:44 where I do think software is largely going to be.
3:06:47 I think this has been happening over the last few weeks,
3:06:50 where people have gone from a month ago saying, "Oh yeah,
3:06:53 agents are kind of slop," which is a famous Karpathy quote,
3:06:56 to what is a little bit of a meme—the industrialization
3:07:01 of software when anyone can just create software with their fingerprints.
3:07:04 I do think we are closer to that side of things,
3:07:07 and it takes direction and understanding how systems
3:07:11 work to extract the best from the language models.
3:07:14 I think it's hard to accept the gravity of how much is going to change
3:07:18 with software development and how many more people
3:07:20 can do things without ever looking at the code.
3:07:22 I think what's interesting is to think about whether
3:07:25 these systems will be independent—completely independent in the sense
3:07:27 that, while I have no doubt that LLMs will
3:07:30 kind of at some point solve coding in a sense,
3:07:33 like calculators solve calculating, right?
3:07:35 So at some point, humans developed a tool where
3:07:38 you never need a human to calculate that number.
3:07:41 You just type it in, and it's an algorithm.
3:07:43 You can do it in that sense.
3:07:45 And I think that's the same probably for coding.
3:07:48 But the question isn't...
3:07:50 I think what will happen is, you will just say,
3:07:52 "Build that website." It will make a really good website,
3:07:55 and then you maybe refine it.
3:07:56 But will it do things independently where...
3:07:59 Will you still be having humans asking the AI to do something?
3:08:05 Like will there be a person to say, "Build that website?" Or will there be
3:08:09 AI that just builds websites or something?
3:08:12 I think talking about building websites is—- Too simple.
3:08:17 The problem with websites and the problem with the web, you know,
3:08:21 HTML and all that kind of stuff, it's very resilient to just...
3:08:25 slop.
3:08:25 It will show you slop; it's good at showing slop.
3:08:28 I would rather think of safety-critical systems,
3:08:32 like asking AI to end-to-end generate something that manages logistics,
3:08:40 or manages cars, a fleet of cars—all that kind of stuff.
3:08:43 So it end-to-end generates that for you.
3:08:46 I think a more intermediate example is
3:08:47 take something like Slack or Microsoft Word.
3:08:50 I think if organizations allow it,
3:08:52 AI could very easily implement features end-to-end and do a fairly
3:08:57 good job for like things that you want to try.
3:09:00 You want to add a new tab in Slack that you want to use,
3:09:03 and I think AI will be able to do that pretty well.
3:09:06 Actually, that's a really great example.
3:09:07 How far away are we from that?
3:09:09 Like this year.
3:09:12 See, I don't know.
3:09:13 I don't know.
3:09:15 I guess I don't know how bad production codebases are,
3:09:17 but I think that within...
3:09:18 on the order of a few years, a lot of people are going to be pushed
3:09:22 to be more of a designer and product manager,
3:09:24 where you have multiple of these agents that can try things for you and they
3:09:28 might take one to two days to implement a feature or attempt to fix a bug.
3:09:32 And you have these dashboards,
3:09:34 which I think Slack is actually a good dashboard where
3:09:36 your agents will talk to you and you'll then give feedback.
3:09:40 But things like, if I make a website, like,
3:09:43 "Do you want a passable logo?" I think
3:09:45 these cohesive design things and the style is
3:09:48 going to be very hard for models and deciding on what to add the next time.
3:09:54 I just...
3:09:54 Okay.
3:09:55 I hang out with a lot of programmers and some
3:09:57 of them are a little bit on the skeptical side in general.
3:10:03 That's just their vibe.
3:10:05 I just think there's a lot of complexity
3:10:07 involved in adding features to complex systems.
3:10:10 Like, if you look at the browser, Chrome.
3:10:13 If I wanted to add a feature,
3:10:15 if I wanted to have tabs as opposed to up top, I want them on the left side.
3:10:21 Interface-wise, right?
3:10:22 I think we're not...
3:10:23 This is not a next-year thing.
3:10:26 One of the Claude releases this year, one of their tests was:
3:10:28 we give it a piece of software and leave Claude to run to recreate it entirely,
3:10:32 and it could already almost rebuild Slack from scratch,
3:10:36 just given the parameters of the software
3:10:38 and left in a sandbox environment to do that.
3:10:41 So the "from scratch" part, I like almost better.
3:10:44 So it might be that smaller and newer companies are advantaged,
3:10:47 and they're like, "We don't have the bloat and complexity,
3:10:51 and therefore this feature exists."- And I
3:10:54 think this gets to the point you mentioned,
3:10:56 that some people you talk to are skeptical.
3:10:58 I think that's not because the LLM can't do X, Y, Z.
3:11:02 It's because people don't want it to do it this way.
3:11:05 Some of that could be a skill issue on the human side.
3:11:08 We have to be honest with ourselves.
3:11:10 And some of that could be an underspecification issue.
3:11:13 So, programming, it's like you're just assuming...
3:11:17 This is like an issue with communication in relationships and friendships.
3:11:22 You're assuming the LLM is supposed to read your mind.
3:11:26 This is where spec-driven design is really important.
3:11:28 Using natural language to specify what you want.
3:11:32 If you talk to people at the labs,
3:11:34 they use these in their training and production code.
3:11:37 Claude Code is built with Claude Code,
3:11:39 and they all use these things extensively.
3:11:41 Dario talks about how much of Claude's code...
3:11:44 It's like these people are slightly ahead in terms
3:11:48 of the capabilities they have and what they probably spend on inference.
3:11:53 They could spend 10 to 100x as much as we're
3:11:56 spending on a lowly $100 or $200 a month plan.
3:11:59 They truly let it rip.
3:12:01 And I think that, with the pace of progress that we have,
3:12:06 it seems like- a year ago we didn't have
3:12:09 Claude Code and we didn't really have reasoning models.
3:12:11 The difference between sitting here today and what
3:12:13 we can do with these models is significant,
3:12:16 and there's a lot of low-hanging fruit to improve them.
3:12:21 The failure modes are pretty dumb.
3:12:23 Like- "Claude, you tried to use a CLI command I don't have installed 14 times,
3:12:27 and then I sent you the command to run."
3:12:30 That, from a modeling perspective, is pretty fixable.
3:12:33 So I don't know.
3:12:34 I agree with you.
3:12:36 I've been becoming more and more bullish in general.
3:12:38 Speaking to what you're articulating, I think it is a human skill issue.
3:12:44 Anthropic is leading the way, along with other companies,
3:12:49 in understanding how to best use the models for programming;
3:12:53 therefore, they're effectively using them.
3:12:54 There are a lot of programmers on the outskirts who don't...
3:12:58 I mean, there's not a really good guide on how to use them.
3:13:03 People are trying to figure it out, but-- It might be very expensive.
3:13:06 The entry point might be $2,000 a month,
3:13:09 which is only for tech companies and rich people.
3:13:12 That could be it.
3:13:14 But it might be worth it.
3:13:15 If the final result is a working software system, it might be worth it.
3:13:19 By the way, it's funny how we converged from the discussion
3:13:22 of timeline to AGI to something more pragmatic and useful.
3:13:25 Is there anything concrete, interesting, useful,
3:13:29 and profound to be said about the timeline to AGI and ASI?
3:13:33 Or are these discussions a bit too detached from the day to day?
3:13:39 There are interesting bets.
3:13:40 A lot of people are trying to do
3:13:42 RLVR— Reinforcement Learning with Verifiable Rewards—in real scientific domains,
3:13:46 where startups with hundreds of millions of funding have wet labs where
3:13:50 they're having language models propose hypotheses
3:13:52 that are tested in the real world.
3:13:54 I would say that they're early, but with the pace of progress,
3:13:59 it's like- ...maybe they're early by six months
3:14:02 and they make it because they were there first,
3:14:05 or maybe they're early by eight years; you don't know.
3:14:08 That type of moonshot to branch this momentum
3:14:13 into other sciences would be very transformative.
3:14:18 If, AlphaFold moments happen in all sorts
3:14:21 of other scientific domains by a startup solving this.
3:14:24 I think there are startups—maybe Harmonic is one—where they're
3:14:27 going all in on language models plus Lean for math.
3:14:31 You had another guest where you talked about this recently,
3:14:34 and we don't know exactly what's going to fall
3:14:37 out of spending $100 million on that model.
3:14:40 Most of them will fail, but a couple might be big breakthroughs that are
3:14:45 very different than ChatGPT or Claude Code type software experiences.
3:14:50 A tool that's only good for a PhD mathematician but makes them 100X effective...
3:14:58 I agree.
3:14:58 I think this will happen in a lot of domains,
3:15:01 especially domains that have a lot of resources,
3:15:05 like finance, legal, and pharmaceutical companies.
3:15:09 But then again, is it really AGI?
3:15:12 Because we are now specializing it again.
3:15:14 Is it really that much different from back
3:15:17 in the day when we had specialized algorithms?
3:15:19 It's just the same thing, way more sophisticated,
3:15:23 but I don't know, is there a threshold for AGI?
3:15:27 I think the real cool thing here is
3:15:29 that we have foundation models we can specialize.
3:15:32 That's like the breakthrough.
3:15:33 Right now, I think we are not there yet because, first, it's too expensive,
3:15:38 but also, ChatGPT doesn't just give away their model to customize it.
3:15:42 I think once that's true...
3:15:44 And I can imagine this as a business model, where OpenAI says at some point,
3:15:51 "Hey, Bank of America,
3:15:52 for $100 million we will do your custom model," something like that.
3:15:56 I think that will be the huge economic value-add.
3:15:59 The other thing, though, is also...
3:16:02 Companies, I mean, what is the differentiating factor?
3:16:06 If everyone uses the same LLM, if everyone uses ChatGPT,
3:16:10 they will all do the same thing.
3:16:12 Well, if everyone is moving in lockstep,
3:16:15 but companies want to have a competitive advantage,
3:16:18 there is no way around using some of their private data and specializing.
3:16:24 It's gonna be interesting.
3:16:26 Seeing the pace of progress, it does feel like things are coming.
3:16:30 I don't think the AGI and ASI thresholds are particularly useful.
3:16:36 I think the real question, and this relates to the remote worker thing,
3:16:40 is: when are we going to see a big, obvious leap in economic impact?
3:16:47 Because currently there's not been an obvious leap
3:16:51 in the economic impact of LLM models, for example.
3:16:55 And that's, you know, aside from AGI or ASI, all that stuff,
3:16:59 there's a real question of, "When are
3:17:03 we going to see a GDP..." "...jump?"- Yeah, what is the GDP made up of?
3:17:08 A lot of it is financial services, so I don't know what this is.
3:17:13 Right, GDP is a-- It's just hard for me to think about the GDP bump,
3:17:17 but I would say that software development becomes valuable in a different way,
3:17:22 when you no longer have to look at the code anymore.
3:17:25 When Claude Code will make you a small business.
3:17:29 Which is essentially, Claude can set up your website,
3:17:31 your bank account, your email, and your whatever else.
3:17:34 And you just have to express what you're trying to put into the world.
3:17:39 That's not just an enterprise market, but it is hard.
3:17:42 I don't know how you get people to try doing that.
3:17:45 I guess if ChatGPT can do it—people are trying ChatGPT.
3:17:49 I think it boils down to the scientific question of, "How hard
3:17:52 is tool use to solve?" Because a lot of the stuff you're implying,
3:17:57 the remote work stuff, is tool use.
3:17:59 It's like...
3:18:00 computer use, like how you have an LLM that goes out there, this agentic system,
3:18:06 and does something in the world, and only screws up 1% of the time.
3:18:12 Computer use-- Or less.
3:18:12 ...is a good example of what labs care about
3:18:14 and we haven't seen a lot of progress on.
3:18:16 We saw multiple demos in 2025 of, like,
3:18:20 Claude can use your computer, or OpenAI had operator, and they all suck.
3:18:24 They're investing money in this, and I think that'll be a good example.
3:18:29 Whereas actually, something where it just seems
3:18:32 like taking over the whole screen seems
3:18:34 a lot harder than having an API that they can call in the back end.
3:18:39 For some of that, you have to set up
3:18:41 a different environment for them all to work in.
3:18:43 They're not working on your MacBook;
3:18:45 they are individually interfacing with Google and Amazon and Slack,
3:18:49 and they handle all these things in a very different way than humans do.
3:18:53 So some of this might be structural blockers.
3:18:56 Also, specification-wise, I think the problem is for arbitrary tasks, well,
3:19:01 you still have to specify what you want your LLM to do.
3:19:04 And how do you do that?
3:19:06 What is the environment?
3:19:07 How do you specify?
3:19:08 You can say what the end goal is, but if it can't solve the end goal...
3:19:13 with LLMs, if you ask it for text, it can always clarify or do sub-steps.
3:19:17 How do you put that information into a system that, let's say,
3:19:20 books a travel trip for you?
3:19:22 You can say, "You screwed up my credit card
3:19:24 information," but even to get it to that point,
3:19:27 even to get it to that point, how do you,
3:19:29 as a user, guide the model before it can even attempt that?
3:19:33 I think the interface is really hard.
3:19:36 Yeah, it has to learn a lot about you specifically.
3:19:39 And this goes to continual learning,
3:19:41 about the general mistakes that are made throughout,
3:19:45 and then mistakes that are made through you.
3:19:48 All the AI interfaces are getting set up to ask humans for input.
3:19:51 I think Claude Code we talked about a lot.
3:19:54 It asks feedback and questions.
3:19:55 If it doesn't have enough specification on your plan or your desired goal,
3:19:59 it starts to ask questions,
3:20:01 "Would you rather?" We talked about Memory, which saves across chats.
3:20:06 Its first implementation is kind of odd,
3:20:08 where it'll mention my dog's name or something in a chat.
3:20:11 I'm like, "You don't need to be subtle about this.
3:20:14 I don't care." But things that are emerging, ChatGPT has the Pulse feature.
3:20:19 Which is like a curated couple of paragraphs with links to something to look
3:20:24 at, and people talk about how models are going to ask you questions.
3:20:28 Which I think is a very...
3:20:31 It's probably going to work.
3:20:33 The language model knows you had a doctor appointment and asks, "Hey,
3:20:36 how are you feeling after that?" Which
3:20:38 again goes into the territory where humans
3:20:40 are very susceptible to this, and there's a lot of social change to come.
3:20:45 But also, they're experimenting with having the models engage.
3:20:48 Some people like this Pulse feature,
3:20:50 which processes your chats and automatically searches
3:20:53 for information and puts it in the app.
3:20:56 So there are a lot of things coming.
3:20:59 I used that feature before,
3:21:00 and I always feel bad because it does that every day, and I rarely check it out.
3:21:05 It's like, how much compute is burned
3:21:07 on something I don't even look at, you know?
3:21:10 It's kind of like, "Oh..."- There's also a lot of idle compute in the world,
3:21:14 so don't feel too bad.
3:21:16 Okay.
3:21:17 Do you think new ideas might be needed?
3:21:20 Is it possible that the path to AGI,
3:21:22 however we define that, to solve computer use more generally,
3:21:26 to solve biology and chemistry and physics—sort
3:21:31 of the Dario Amodei definition of AGI?
3:21:35 Do you think it's possible that totally new ideas are needed?
3:21:42 Non-LLM, non-RL ideas?
3:21:45 What might they look like?
3:21:47 We're going into philosophy land a bit.
3:21:51 For something like a singularity to happen, I would say yes.
3:21:54 The new ideas could be architectures or training algorithms,
3:21:58 fundamental deep learning things.
3:22:00 But in that nature, they're pretty hard to predict.
3:22:04 I think we won't get very far even without those advances.
3:22:08 We might get the software solution,
3:22:10 but it might stop at software and not do computer use without more innovation.
3:22:15 So I think that a lot of progress will be coming, but if you're gonna zoom out,
3:22:20 there's still ideas in the next 30 years that are gonna look like
3:22:24 that was a major scientific innovation that enabled the next chapter of this.
3:22:29 And I don't know if it comes in one year or in 15 years.
3:22:33 Yeah.
3:22:33 I wonder if the bitter lesson holds true for the next 100 years,
3:22:36 what that looks like.
3:22:38 If scaling laws are fundamental in deep learning,
3:22:40 I think the bitter lesson will always apply,
3:22:42 which is compute will become more abundant, but even within abundant compute,
3:22:48 the ones that have a steeper scaling law slope or a better offset— like,
3:22:53 this is a 2D plot of performance
3:22:55 and compute—and like even if there's more compute available,
3:22:58 the ones that get 100x out of it will win.
3:23:01 It might be something like literally
3:23:04 computer clusters orbiting Earth with solar panels.
3:23:09 The problem with that is heat dissipation.
3:23:11 You get all the radiation from the sun and don't have any air to dissipate heat.
3:23:15 But there is a lot of space to put clusters.
3:23:17 There's a lot of solar energy there
3:23:19 and you could figure out the heat dissipation,
3:23:21 as there is a lot of energy and there probably could
3:23:24 be engineering will to solve the heat problem— so there could be.
3:23:27 Is it possible—and we should say that it definitely is
3:23:30 possible— that we're basically going to be plateauing this year?
3:23:36 Not in terms of— the system capabilities,
3:23:40 but what they actually mean for human civilization.
3:23:44 So on the coding front, really nice websites will be built.
3:23:50 Very nice auto-complete.
3:23:53 Very nice way to understand code bases and maybe help debug,
3:23:59 but really just a very nice helper on the coding front.
3:24:03 It can help research mathematicians do some math.
3:24:06 It can help you with shopping.
3:24:09 It's a nice helper.
3:24:10 It's Clippy on steroids.
3:24:12 What else?
3:24:13 It may be a good education tool and all that kind of stuff,
3:24:19 but computer use turns out extremely difficult to solve.
3:24:24 So I'm trying to frame the cynical case in all
3:24:28 these domains where there's not a really huge economic impact,
3:24:32 but realize how costly it is to train these systems at every level,
3:24:37 both the pre-training and the inference,
3:24:39 how costly the inference is, the reasoning, all of that.
3:24:43 Like, is that possible?
3:24:44 And how likely is that, do you think?
3:24:47 When you look at the models,
3:24:49 there are so many obvious things to improve and it takes
3:24:52 a long time to train these models and to do this art,
3:24:55 and it'll take us with the ideas that we have multiple years
3:24:59 to actually saturate in terms of whatever
3:25:02 benchmark or performance we are searching for.
3:25:05 It might serve very narrow niches;
3:25:07 like the average ChatGPT 800 million user might not get a lot of benefit out
3:25:12 of this, but it is going to serve
3:25:14 different populations by getting better at different things.
3:25:18 But I think what everybody's chasing now
3:25:21 is a general system that's useful to everybody.
3:25:24 So, okay, so if that's not...
3:25:26 That can plateau, right?
3:25:28 I think that dream is actually kind of dying.
3:25:30 As you talked about with the specialized models where it's like...
3:25:34 And multimodal is often...
3:25:36 Video generation is a totally different thing.
3:25:38 Thing.
3:25:39 "That dream is kind of dying" is a big statement,
3:25:42 because I don't know if it's dying.
3:25:44 If you ask the actual frontier lab people, they...
3:25:46 I mean, they're still chasing it, right?
3:25:48 I do think they are still rushing to get the next model out,
3:25:52 which will be much better than the...
3:25:54 "Much" is a relative term, but it will be better than the previous one.
3:25:58 And I can't see them slowing down.
3:26:00 I just think the gains will be made or felt
3:26:03 more through not only scaling the model, but now...
3:26:08 I feel like there's a lot of tech debt.
3:26:10 It's like, "Well, let's just put the better
3:26:12 model in there." Better model, better model.
3:26:15 And now people are like, "Okay,
3:26:17 let's also at the same time improve everything around it
3:26:19 too." Like the engineering of the context and inference scaling.
3:26:23 The big labs will still keep doing that.
3:26:26 And now also the smaller labs will catch up, because now they are hiring more.
3:26:31 There will be more people and LLMs.
3:26:33 It's kind of like a circle.
3:26:35 They also make them more productive and it's just...
3:26:38 It's like amplification.
3:26:39 I think what we can expect is amplification, but not like a change of any...
3:26:43 not like a paradigm change.
3:26:45 I don't think that is true, but everything will be just amplified and amplified,
3:26:48 and I can see that continuing for a long time, you know?
3:26:53 Yeah.
3:26:53 I guess my statement that the dream is dying
3:26:55 depends on exactly what you think it's gonna be doing.
3:26:58 Like, Claude Code is a general model that can do a lot of things,
3:27:02 but it's not necessarily...
3:27:05 It depends a lot on integrations.
3:27:06 I bet Claude Code could do a fairly good job of doing your email,
3:27:10 and the hardest part is figuring out how to give information to it
3:27:13 and how to get it to be able to send your emails.
3:27:17 But that's just kind of like...
3:27:18 I think it goes back to what is the "one model to rule everything" ethos,
3:27:23 which is just like a thing in the cloud that handles
3:27:26 your entire digital life and is way smarter than everybody.
3:27:29 It's like it's operating in a...
3:27:34 So it's an interesting leap of faith to go
3:27:37 from "Claude Code becomes that," which in some ways is...
3:27:41 There are some avenues for that, but I do think
3:27:45 that the rhetoric of the industry is a little bit different.
3:27:49 I think the immediate thing we will feel next as a normal
3:27:52 person using LLMs will probably be related to something trivial,
3:27:57 like making figures.
3:27:58 Right now, LLMs are terrible at making figures.
3:28:01 Is it because we are getting served the cheap
3:28:04 models with much less inference compute than behind the scenes?
3:28:08 Maybe some.
3:28:09 Like, there are some ways to get better figures, but if you ask today,
3:28:13 ..."Draw a flowchart of X, Y, Z," it's most of the time terrible.
3:28:18 And it is a very simple task for a human.
3:28:20 I think it's almost easier sometimes to draw something than to write something.
3:28:25 Yeah, the multimodal understanding does feel like something that is odd...
3:28:28 ...that it's not better solved.
3:28:31 I think we're not saying one obvious thing that we're not realizing,
3:28:35 that's a gigantic thing that's hard to measure,
3:28:37 which is making all of human knowledge accessible— —to the entire world.
3:28:46 One thing that is hard to articulate is
3:28:49 the huge difference between Google Search and an LLM.
3:28:52 I feel like I can basically ask an LLM anything and get an answer,
3:28:59 and it's doing less and less hallucination.
3:29:04 And that means understanding my own life, figuring out a career trajectory,
3:29:09 solving the problems all around me,
3:29:11 learning about anything through human history.
3:29:16 I feel like nobody's really talking about that, because they
3:29:22 just immediately take it for granted that this is awesome.
3:29:25 That's why everybody's using it: because you get answers for stuff.
3:29:29 Think about the impact across time.
3:29:33 This is not just in the United States; it's all across the world.
3:29:37 Kids throughout the world being able to learn
3:29:40 these ideas— the impact that has across time is probably...
3:29:45 ...That's the real impact.
3:29:48 Talk about GDP; it won't be like a leap.
3:29:51 It'll be...
3:29:52 ...that's how we get to Mars, that's how we build these things,
3:29:56 that's how we have a million new OpenAIs and all the innovation from there.
3:30:00 It's this quiet force that permeates everything: human knowledge.
3:30:06 I agree with you.
3:30:08 In a sense, it makes knowledge more accessible,
3:30:10 but it also depends on what the topic is.
3:30:13 For something like math, you can ask it questions and it answers,
3:30:21 but if you want to learn a topic from scratch,
3:30:26 the sweet spot is still elsewhere.
3:30:28 There are really good math textbooks laid out linearly,
3:30:32 and that is a proven strategy to learn a topic.
3:30:36 It makes sense, if you start from zero,
3:30:39 to use information-dense text to soak it up,
3:30:43 but then you use the LLM to make infinite exercises.
3:30:47 Like, you have problems in a certain area
3:30:49 or have questions that something's- uncertain about certain things,
3:30:53 you ask it to generate example problems, you solve them,
3:30:59 and you need more background knowledge, you ask it to generate that.
3:31:03 But then...
3:31:04 it won't give you anything, let's say, that is not in the textbook.
3:31:10 It's just packaging it differently, if that makes sense.
3:31:13 But then there are things I feel like where
3:31:15 it also adds value in a more timely sense,
3:31:18 where there is no good alternative besides a human doing it on the fly.
3:31:24 For example, if you're planning to go to Disneyland and you
3:31:28 try to figure out which tickets to buy for which park when,
3:31:32 well, there is no textbook on that.
3:31:34 There is no information-dense resource.
3:31:36 There's only the sparse internet, and then there is a lot of value in the LLM.
3:31:40 You just ask it.
3:31:42 You have constraints on traveling these days.
3:31:44 I want to go there and there.
3:31:46 Please figure out what I need, when and from where, what it costs and stuff like
3:31:50 that, and it is a very customized, on-the-fly package.
3:31:56 And this is like one of a thousand examples
3:31:58 of personalized- Personalization is essentially
3:32:01 pulling information from the sparse internet,
3:32:04 the non-information-dense thing where there's no better version that exists.
3:32:09 It just doesn't exist.
3:32:10 You make it almost from scratch.
3:32:12 And if it does exist, it's full of- speaking of Disney World,
3:32:15 full of- what would you call it?
3:32:18 Ad slop.
3:32:19 It's impossible.
3:32:20 Take any city in the world, what are the top 10 things to do?
3:32:27 An LLM is just way better to ask than anything on the internet.
3:32:30 Well, for now, that's because they're subsidized
3:32:32 and they're gonna be paid for by ads.
3:32:36 Oh my goodness.
3:32:37 It's coming.
3:32:38 No.
3:32:39 No.
3:32:39 I mean, I'm hoping there's a very clear indication what's
3:32:43 an ad and what's not an ad in that context.
3:32:46 That's something I mentioned a few years ago.
3:32:49 If, I don't know, if you are looking for a new running shoe,
3:32:52 well, is it a coincidence that Nike maybe comes up first?
3:32:56 Maybe, maybe not.
3:32:58 But I think there are clear laws.
3:33:00 You have to be clear about that.
3:33:02 I think that's what everyone fears.
3:33:04 It's the subtle message in there, but that also brings us to the topic of ads,
3:33:11 where I think this was a thing.
3:33:13 Hopefully, I think for- in 2025, just because I think it's they're still
3:33:19 not making money in other ways right now.
3:33:22 Having ad spots in there...
3:33:24 but the thing is, they couldn't,
3:33:26 because there are alternatives without ads and people
3:33:30 would just flock- to the other products.
3:33:33 It's also just crazy how- yeah, how they're one-upping each other,
3:33:38 spending so much money to just get the users.
3:33:41 I think so.
3:33:42 Like, some Instagram ads— I don't use Instagram,
3:33:44 but I understand the appeal of paying a platform
3:33:48 to find users who will genuinely like your product,
3:33:52 and that is the best case of things like Instagram ads.
3:33:56 But there are also plenty of cases
3:33:58 where advertising is very awful for incentives,
3:34:00 and I think that a world where the power
3:34:04 of AI can integrate with that positive view
3:34:06 of, "I am a person and I have a small business and I want to make the best,
3:34:11 I don't know, damn steak knives in the world,
3:34:13 and I want to sell them to somebody who needs them."
3:34:16 And if AI can make that sort of advertising thing work even better,
3:34:20 that's very good for the world, especially with digital infrastructure,
3:34:24 because that's how the modern web has been built.
3:34:27 But that's not to say that addicting feeds so
3:34:31 that you can show people more content is a good thing.
3:34:35 So, I think that's even what OpenAI would say,
3:34:37 is they want to find a way that can make
3:34:40 the monetization upside of ads while still giving their users agency.
3:34:45 And I personally would think that Google is probably going
3:34:47 to be better at figuring out how to do this, because
3:34:50 they already have ad supply and if they figure out how
3:34:54 to turn this demand in their Gemini app into useful ads,
3:34:57 then they can turn it on.
3:34:59 And somebody will figure it out—I don't know if it's this year,
3:35:03 but there will be experiments with it.
3:35:06 I do think what holds companies back right now
3:35:08 is really just that the competition is not doing it.
3:35:11 It's more like a reputation thing.
3:35:13 It's just, I think people are just afraid
3:35:16 right now of ruining or losing their reputation,
3:35:19 losing users, because it would make headlines if someone launched these ads.
3:35:22 But—- Unless they were great, but the first ads won't be great because it's
3:35:26 a hard problem that we don't know how to solve.
3:35:28 Yeah, I think also the first version of that will likely be something like on X,
3:35:32 like the timeline where you have a promoted post sometimes in between.
3:35:35 It'll be something where it will say "promoted" or something small,
3:35:38 and then there will be an image.
3:35:39 I think right now the problem is: who makes the first move?
3:35:43 If we go 10 years out,
3:35:44 the proposition for ads is that you will make so much money on ads by having
3:35:49 so many users that you can use this to fund better R&D and make better models,
3:35:53 which is why YouTube is dominating the market
3:35:57 for any— Netflix is scared of YouTube.
3:36:01 They have the ads, they make—I pay $28 a month for Premium.
3:36:04 They make at least $28 a month off of me and many other people.
3:36:09 And they're just creating such a dominant position in video.
3:36:12 So I think that's the proposition:
3:36:14 that ads can make you have a sustained advantage.
3:36:17 in what you're spending per user.
3:36:20 But there's so much money in it right now that somebody
3:36:24 starting that flywheel is scary because it's a long-term bet.
3:36:29 Do you think there'll be some crazy big moves this year business-wise?
3:36:33 Like Google or Apple acquiring Anthropic or something like this?
3:36:40 Dario will never sell, but we are starting to see some types of consolidation
3:36:44 with Groq for $20 billion and Scale AI for almost $30
3:36:49 billion and countless other deals like this that are structured
3:36:52 in a way that is detrimental to the Silicon Valley ecosystem,
3:36:57 which is this licensing deal where not everybody gets brought along,
3:37:02 rather than a full acquisition that benefits
3:37:05 the rank-and-file employee by getting their stock vested.
3:37:07 That's a big issue for culture to address
3:37:10 because the startup ecosystem is the lifeblood where, if you join a startup,
3:37:16 even if it's not successful, it might get acquired on a cheap premium
3:37:21 and you'll get paid out for this equity.
3:37:24 These licensing deals are taking the top talent a lot of the time.
3:37:27 The deal for Groq to NVIDIA is rumored to be better to the employees,
3:37:32 but it is still this antitrust-avoiding thing.
3:37:35 But I think that this trend of consolidation will continue.
3:37:39 Me and many smart people I respect
3:37:41 have been expecting consolidation to have happened sooner,
3:37:44 but it seems like some of these things are starting to turn,
3:37:49 but at the same time,
3:37:51 companies are raising ridiculous amounts of money for reasons where I'm like,
3:37:56 "I don't know why you're taking that money." So it's mixed this year,
3:38:01 but some consolidation pressure is starting.
3:38:05 What kind of surprising consolidation will we see?
3:38:07 You say Anthropic is a "never." I mean, Groq is a big one.
3:38:10 Groq with a Q, by the way.
3:38:12 Yeah.
3:38:12 There's just a lot of startups and a very high premium on AI startups.
3:38:16 So there could be a lot of- that kind of stuff, yeah.
3:38:19 $10 billion range acquisitions,
3:38:21 which is really big for a startup that was maybe founded a year ago.
3:38:25 I think Manus.ai...
3:38:27 this company based in Singapore that Meta-founded was founded
3:38:30 eight months ago and then had a $2 billion exit.
3:38:33 I think there will be some
3:38:35 other multi-billion dollar acquisitions, like Perplexity.
3:38:39 Like Perplexity, right?
3:38:40 Yeah, people rumor them to Apple.
3:38:41 I think there's a lot of of pressure and liquidity in AI.
3:38:46 There's pressure on big companies to have outcomes and- I would guess that a big
3:38:52 acquisition gives people leeway to then tell the next chapter of that story.
3:38:56 I guess Cursor—we've been talking about code—somebody acquires Cursor.
3:39:00 if somebody acquires Cursor...
3:39:02 They're in such a good position by having so much user data.
3:39:05 And we talked about continual learning.
3:39:07 They had one of the most interesting sentences in a blog post,
3:39:10 which is that they had their new Composer model,
3:39:13 which was a fine-tune of one of these large Mixture of Expert models from China.
3:39:17 You can know that by asking it or because the model
3:39:21 sometimes responds in Chinese— ...which none of the American models do.
3:39:24 And they had a blog post where they're like,
3:39:26 "We're updating the model weights every 90
3:39:28 minutes based on real-world feedback from people
3:39:30 using it." Which is like the closest
3:39:32 thing to real-world RL happening on a model,
3:39:34 and it's just mentioned in one of their blog posts—- That's incredible.
3:39:37 which is super cool.
3:39:38 And by the way, I should say I use Composer a lot
3:39:40 because one of the benefits it has is that it's fast.
3:39:43 I need to try it 'cause everybody says this.
3:39:45 And there'll be some IPOs potentially.
3:39:48 You think Anthropic, OpenAI, xAI.
3:39:51 They can all raise so much money so easily that they don't feel a need to.
3:39:55 So long as fundraising is easy,
3:39:56 they're not going to IPO because public markets apply pressure.
3:40:00 I think we're seeing in China that the ecosystem's a little
3:40:02 different with both MiniMax and Z.ai applying for, filing IPO paperwork,
3:40:08 which will be interesting to see how the Chinese market reacts.
3:40:11 I actually would guess that it's going to be similarly hypey to the US,
3:40:16 so long as all this is going and not based
3:40:18 on the reality that they're both losing a ton of money.
3:40:21 I wish more of the gigantic American AI startups were public because it
3:40:25 would be very interesting to see how
3:40:26 they're spending money and have more insight.
3:40:28 And also just to give people access to investing in these, because I
3:40:33 think they're some of the most
3:40:35 formidable companies—they're the companies of the era.
3:40:38 And the tradition is now for so many
3:40:40 of the big startups in the US to not go public.
3:40:43 It's like we're still waiting for Stripe and the IPO,
3:40:46 but Databricks definitely didn't.
3:40:47 They raised like a Series G or something.
3:40:50 And I just feel like it's kind
3:40:52 of a weird equilibrium for the market where it's like,
3:40:56 I would like to see these companies go public
3:40:58 and evolve in that way that a company can.
3:41:01 Do you think 10 years from now some
3:41:03 of the frontier model companies are still around?
3:41:06 Anthropic, OpenAI?
3:41:08 I definitely don't see it as a winner-takes-all unless there truly is
3:41:12 some algorithmic secret that one of them finds that lets this flywheel.
3:41:16 Because the development path is so similar for all of them.
3:41:19 Google and OpenAI have all the same products, and then Anthropic's more focused,
3:41:24 but when you talk to people it sounds
3:41:26 like they're solving a lot of the same problems.
3:41:28 So I think...
3:41:29 and there's offerings that'll spread out.
3:41:30 There's a lot of...
3:41:31 it's a very big cake being made that people are going to take money out of.
3:41:37 I don't want to trivialize it,
3:41:39 but OpenAI and Anthropic are primarily LLM service providers.
3:41:44 And some of the other companies like Google and xAI,
3:41:48 linked to X, do other stuff too.
3:41:51 And so it's very possible, if AI becomes more commodified,
3:41:56 that the companies just providing LLMs will die.
3:42:00 I think the advantage they have is a lot of users,
3:42:03 and I think they will just pivot.
3:42:05 Like Anthropic, I think, pivoted.
3:42:08 I don't think they originally planned to work on code,
3:42:13 but they found, "Okay, this is a nice niche,
3:42:16 and now we are comfortable and we push
3:42:18 on this niche." I can see the same thing...
3:42:20 Let's say hypothetically, I'm not sure if it will be true,
3:42:24 but let's say Google takes all the market share of the general chatbot.
3:42:27 Maybe OpenAI will then focus on some other sub-topic.
3:42:31 They have too many users to go away in the foreseeable future.
3:42:37 I think Google is always ready to say, "Hold my beer," with AI models.
3:42:41 I think the question is if the companies can support the valuations.
3:42:44 I see the AI companies being looked at in some ways like AWS, Azure,
3:42:50 and GCP are, all competing in the same space and all very successful businesses.
3:42:54 There's a chance that the API market is so unprofitable
3:42:58 that they go up and down the stack to products and hardware.
3:43:01 They have so much cash that they can build power plants and data centers,
3:43:04 which is a durable advantage now.
3:43:06 But there's also a reasonable outcome that these APIs are so
3:43:10 valuable and so flexible for developers that they become something like AWS.
3:43:15 But AWS and Azure are also going to have these APIs,
3:43:20 so having five or six people competing in the API market is hard.
3:43:24 So maybe that's why they get squeezed out.
3:43:27 You mentioned "RIP Llama." Is there a path to winning for Meta?
3:43:32 I think nobody knows.
3:43:34 They're moving a lot, so they're signing licensing deals with Black Forest Labs,
3:43:40 which is an image generation company, or Midjourney.
3:43:43 So I think in some ways on the product and consumer-facing AI front,
3:43:49 it's too early to tell.
3:43:51 I think they have some people who are
3:43:53 excellent and very motivated being close to Zuckerberg.
3:43:56 So I think there's still a story to unfold there.
3:44:00 Llama is a bit different,
3:44:02 where Llama was the most focused expression of the organization.
3:44:06 And I don't see Llama being supported to that extent.
3:44:09 I think it was a very successful brand for them.
3:44:13 So they still might participate in the open ecosystem
3:44:16 or continue the Llama brand into a different service,
3:44:19 because people know what Llama is.
3:44:21 You think there's a Llama 5?
3:44:24 Not an open-weight one.
3:44:27 It's interesting.
3:44:28 I think Llama was the pioneering open-weight model.
3:44:32 With Llama 1, 2, and 3, there was a lot of love.
3:44:37 But I think then, hypothesizing or speculating,
3:44:40 I think the leaders at Meta, like the upper executives, they...
3:44:44 I think they got very excited about Llama because
3:44:47 they saw how popular it was in the community.
3:44:49 And then I think the problem was trying to, let's say,
3:44:53 monetize the open—or not monetize the open source,
3:44:55 but use it to make a bigger splash.
3:44:58 It felt almost forced,
3:45:01 like developing these very big Llama 4 models to be on top of the benchmarks.
3:45:08 But I don't think the goal of Llama models
3:45:10 is to be on top of the benchmarks beating, let's say, ChatGPT or other models.
3:45:14 I think the goal was to have a model that people can use,
3:45:18 trust, modify, and understand.
3:45:20 So that includes having smaller models.
3:45:21 They don't have to be the best models.
3:45:23 And what happened was, these models were, of course...
3:45:27 the benchmarks suggested that they were better than they were because they
3:45:31 had specific models trained on preferences
3:45:33 so that they performed well on benchmarks.
3:45:35 That's kind of, like, this overfitting thing to force it to be the best.
3:45:38 But then at the same time,
3:45:40 they didn't do the small models that people could use.
3:45:42 And I think that no one could run these big models then.
3:45:45 And then there was kind of a weird thing.
3:45:47 I think it's just because people got
3:45:49 too excited about headlines pushing the frontier.
3:45:52 I think that's it.
3:45:54 And too much on the benchmarking side.
3:45:56 It's too much work.
3:45:57 I think it imploded under internal political fighting and misaligned incentives.
3:46:03 The researchers want to build the best models,
3:46:06 but there's a layer of organization— ...and management
3:46:08 that is trying to demonstrate that they do these things.
3:46:11 And then there are rumors about how,
3:46:14 for example, some horrible technical decision was made.
3:46:19 It just seems like it got so bad that it all just crashed out.
3:46:25 Yeah, but we should also give huge props to Mark Zuckerberg.
3:46:28 I think it comes from Mark, actually, from Mark Zuckerberg,
3:46:32 from the top of the leadership, saying open source is important.
3:46:35 The fact that that leadership exists means there could be a Llama 5,
3:46:41 where they learn the lessons from benchmarking and say,
3:46:44 "We're going to be GPT-OSS—" "...and provide a really awesome library of open
3:46:51 source."- What people say is that there's
3:46:53 a debate between Mark and Alexandr Wang,
3:46:56 who is very bright, but much more against open source.
3:46:59 And to the extent that he has a lot of influence over the AI org,
3:47:02 it seems much less likely, because it seems like Mark brought him
3:47:05 in for a fresh leadership eye in directing AI.
3:47:10 And if being open or closed is no longer the defining nature of the model,
3:47:14 I don't expect that to be a defining argument between Mark and Alex.
3:47:19 They're both very bright, but I just have a hard time understanding all
3:47:23 of it because Mark wrote this piece in July of 2024,
3:47:28 which was probably the best blog post at the time,
3:47:33 saying "The Case for Open Source AI." And then July 2025 came around and it was,
3:47:38 "We're reevaluating our relationship with open source." So it's just kind of...
3:47:43 But I think also the problem...
3:47:44 Not the problem, but I think, well,
3:47:46 we may have been a bit too harsh, and that caused some of that.
3:47:50 Because I mean, we as open source developers or the community...
3:47:54 Even though the model was maybe not what
3:47:57 everyone hoped for, it got a lot of backlash.
3:48:00 And I think that was unfortunate because I can see that as a company,
3:48:04 they were hoping for positive headlines.
3:48:06 And instead of just getting no headlines or positive headlines,
3:48:11 in turn they got negative headlines.
3:48:13 And then it kind of reflected bad on the company.
3:48:17 I think that is also something where
3:48:19 it's maybe a spite reaction, almost like, "Okay, we tried to do something nice,
3:48:24 we tried to give you something cool, like an open source model,
3:48:27 and now you are kind of being negative about us,
3:48:31 even for the company." So in that sense,
3:48:34 it looks like, "Well, maybe then we'll change our mind." I guess.
3:48:37 I don't know.
3:48:39 Yeah, that's where the dynamics of discourse on X can lead us,
3:48:46 as a community, astray.
3:48:48 Because sometimes it feels random.
3:48:49 People pick the thing they like and don't like.
3:48:52 I mean, you can see the same thing with Grok 4.1 and Grok Code Fast 1.0.
3:48:59 I don't think, vibe-wise, people love it publicly.
3:49:04 But a lot of people use it.
3:49:09 So if you look to Reddit and X, they don't really give it praise
3:49:13 from the programming community, but they use it.
3:49:17 And the same thing with probably Llama.
3:49:19 I don't understand the dynamics of either positive hype or negative hype.
3:49:23 I don't understand it.
3:49:25 I mean, one of the stories of 2025 is the US filling the gap of Llama,
3:49:29 which is the rise of these Chinese open-weight models,
3:49:33 models- to the point where that was the single
3:49:35 issue I've spent a lot of energy on lately,
3:49:37 trying to do policy work to get the US to invest in this.
3:49:42 So just tell me the story of ADAM.
3:49:43 The ADAM Project started as me calling it the American DeepSeek Project,
3:49:47 which doesn't really work for DC audiences,
3:49:49 but it's the story of the most impactful thing I can do with my career,
3:49:54 which is that these Chinese open-weight models are cultivating a lot of power,
3:49:58 and there is a lot of demand for building on these open models,
3:50:02 especially in enterprises in the US that are very cagey about Chinese models.
3:50:06 The ADAM Project, American Truly Open Models,
3:50:10 is a US-based initiative to build and host high-quality,
3:50:14 genuinely open-weight AI models and supporting
3:50:16 infrastructure explicitly aimed at competing
3:50:19 with and catching up to China's rapidly advancing open-source AI ecosystem.
3:50:25 I think the one-sentence summary would be that...
3:50:28 or two sentences.
3:50:29 One is a proposition that open models are going to be
3:50:32 an engine for AI research because that is what people start with; therefore,
3:50:36 it's important to own them.
3:50:37 And the second one is, therefore, the US should be building the best models
3:50:42 so that the best research happens in the US,
3:50:45 and those US companies take the value from being
3:50:48 the home of where AI research is happening.
3:50:51 And without more investment in open models—we have
3:50:54 plots on the website where it's like, "Qwen, Qwen,
3:50:57 Qwen, Qwen"—it's all these models that are excellent
3:51:01 from these Chinese companies that are cultivating influence internationally.
3:51:05 I think the US is spending way more on AI,
3:51:10 and the ability to create open models that are a generation
3:51:13 beyond what the cutting edge of closed labs costs roughly $100 million,
3:51:18 which is a lot of money, but not a lot of money to these companies.
3:51:22 Therefore, we need a centralizing force of people who want to do this.
3:51:26 And I think we got engagement from people pretty much across the full stack,
3:51:32 whether it's policy.
3:51:34 So there has been support from the administration?
3:51:37 I don't think anyone technically in government has signed it publicly,
3:51:41 but I know people that have worked in AI policy,
3:51:45 in both the Biden and Trump administrations,
3:51:47 are very supportive of promoting open-source models in the US.
3:51:50 I think, for example,
3:51:52 AI2 got a grant from the NSF for $100 million over four years,
3:51:56 which is the biggest CS grant the NSF has ever awarded,
3:52:01 and it's for AI2 to attempt this.
3:52:04 It's a starting point.
3:52:05 But the best thing happens when
3:52:07 there are multiple organizations building models,
3:52:09 because they can cross-pollinate ideas and build this ecosystem.
3:52:13 I don't think it works if it's just Llama releasing models,
3:52:17 because Llama could go away.
3:52:19 The same thing applies for AI2; I can't be the only one building models.
3:52:25 It becomes a lot of time spent on talking to people, whether in policy...
3:52:31 I know NVIDIA is very excited about this.
3:52:34 I think Jensen Huang has been talking about the urgency
3:52:37 for this, and they've done a lot more in 2025,
3:52:40 where the Nemotron models are more of a focus.
3:52:43 They've started releasing some data along with NVIDIA's open models,
3:52:47 and very few companies do this, especially of NVIDIA's size,
3:52:51 so there are signs of progress.
3:52:54 We hear about Reflection AI, where they say their two billion dollar
3:52:58 fundraise is dedicated to building US open models,
3:53:00 and I feel and their announcement tweet reads like a blog post, right?
3:53:06 I think that cultural tide is starting to turn.
3:53:10 In July, four or five DeepSeek-caliber Chinese
3:53:14 open-weight models and and zero from the US.
3:53:18 That's the moment where I realized, like, "Oh,
3:53:20 I guess I have to spend energy on this because nobody else
3:53:23 is gonna do it." So it takes a lot of people contributing together,
3:53:26 and I don't say that, the Adam Project
3:53:28 isn't the thing that's helping to move the ecosystem,
3:53:31 but it's people like me doing this sort of thing to get the word out.
3:53:36 Do you like the 2025 America's AI Action Plan?
3:53:39 That includes open source stuff.
3:53:40 The White House AI Action Plan includes
3:53:43 a dedicated section titled "Encourage Open-Source and Open-Weight
3:53:47 AI," defining such models and arguing they
3:53:49 have unique value for innovation and startups.
3:53:52 Yeah.
3:53:53 I mean, the AI Action Plan is a plan, but largely,
3:53:56 I think it's maybe the most coherent policy
3:54:00 document that has come out of the administration,
3:54:02 and I hope that it largely succeeds.
3:54:05 I know people that have worked on the AI Action
3:54:07 Plan and the challenges of taking policy and making it real.
3:54:10 I have no idea how to do this as an AI researcher,
3:54:13 but largely a lot of things in that were very real,
3:54:16 and there's a huge build-out of AI in the country.
3:54:19 There are a lot of issues that people are hearing about,
3:54:22 from water use to whatever,
3:54:23 and we should be able to build things in this country,
3:54:26 but also, we need to not ruin places
3:54:29 in our country in the process of building it,
3:54:32 and it's worthwhile to spend energy on.
3:54:34 I think that's a role the federal government plays.
3:54:37 They set the agenda.
3:54:38 And with AI, setting the agenda
3:54:40 that open-weight should be a first consideration is
3:54:44 a large part of what they can do and then people think about it.
3:54:49 Also, for education and talent for these companies,
3:54:52 it's very important because otherwise, if there are only closed models,
3:54:56 how do you get the next generation of people contributing at some point?
3:55:01 Because otherwise, you will point only be
3:55:04 able to learn after you joined a company.
3:55:07 But at that point, how do you hire talented people?
3:55:11 How do you identify talented people?
3:55:13 I think open source is essential for a lot of things,
3:55:16 but also even just for educating the population
3:55:19 and training the next generation of researchers.
3:55:21 It's the way, or the only way.
3:55:24 The way that I could've gotten this to go more viral was
3:55:27 to tell a story of Chinese AI integrating with an authoritarian state,
3:55:31 being ASI and taking over the world,
3:55:33 and therefore we need our own American models.
3:55:35 But it's very intentional why I talk about innovation and science
3:55:38 in the US because I think it's both more realistic as an outcome,
3:55:42 but also it's a world that I would like to manifest.
3:55:48 I would say, though, also even any open-weight model,
3:55:52 I do think, is a valuable model.
3:55:55 Yeah.
3:55:55 And my argument is that we should be in a leading position.
3:55:58 But I think it's worth saying it simply because there are still voices in the AI
3:56:04 ecosystem that say we should consider banning
3:56:06 the release of open models due to safety risks.
3:56:09 And I think it's worth adding that, effectively,
3:56:12 that's impossible without making the US have its own great firewall,
3:56:16 which is also known to not work
3:56:19 that well because the cost for training these models,
3:56:22 whether it's one to a hundred million dollars,
3:56:24 is attainable to a huge amount of people
3:56:28 in the world that want to have influence,
3:56:30 so these models will be trained all over the world.
3:56:33 And we want the models, especially when,
3:56:36 like, I mean, there are safety concerns,
3:56:39 but we want this information and tools to flow freely across the world
3:56:43 and into the US so that people can use them and learn from them.
3:56:46 Stopping that would be such a restructuring
3:56:49 of our internet that it seems impossible.
3:56:51 Do you think maybe in that case the big open-weight
3:56:54 models from China are actually a good thing in a sense,
3:56:57 like, for the US companies?
3:56:58 Because maybe the US companies you
3:57:00 mentioned earlier are usually one generation behind
3:57:03 in terms of what they release open source versus what they are using?
3:57:06 For example, gpt-oss might not be the cutting-edge model.
3:57:09 Gemini 3 might not be,
3:57:11 but they do that because they know this is safe to release.
3:57:13 But then when they see, these companies see,
3:57:16 for example, there is DeepSeek-V3.2, which is really awesome,
3:57:20 and it gets used and there is no backlash, there is no security risk,
3:57:24 that could then, again, encourage them to release better models.
3:57:27 Maybe that, in a sense, is a very positive thing.
3:57:30 A hundred percent.
3:57:31 These Chinese companies have set things into motion that I think
3:57:33 would potentially not have happened if they were not all releasing models.
3:57:38 So I think it was like I'm almost
3:57:41 sure that those discussions have been had by leadership.
3:57:45 Is there a possible future where the dominant
3:57:48 AI models in the world are all open source?
3:57:51 Depends on the trajectory of progress that you predict.
3:57:53 If you think saturation in progress is coming within a few years,
3:57:57 so essentially, within the time where financial support is still very good,
3:58:01 then open models will be so optimized and so
3:58:04 much cheaper to run that they'll win out.
3:58:06 This goes back to open source ideas where so many more people will be putting
3:58:10 money into optimizing the serving of these open-weight
3:58:14 common architectures that they will become standards,
3:58:17 and then you could have chips dedicated to them,
3:58:19 and it'll be way cheaper than the offerings
3:58:22 from these closed companies that are custom.
3:58:25 We should say that the AI27 report kinda predicts one of the things it
3:58:30 does from a narrative perspective is that there will be a lot of centralization.
3:58:32 As the AI systems get smarter and smarter,
3:58:36 national security concerns will arise, and you'll centralize the labs,
3:58:40 and they'll become super secretive,
3:58:42 and there'll be this whole race- ...from a military perspective of how do you...
3:58:47 between China and the US.
3:58:48 And so all of these fun conversations we're having about LLMs...
3:58:53 the generals and the soldiers will come into the room and be like, "All right.
3:58:58 We're now in the Manhattan Project stage of this whole thing."- I think in 2025,
3:59:04 '26, '27, I don't think something like that is even remotely possible.
3:59:08 You can make the same argument for computers, right?
3:59:11 You can say, "Computers are capable and we don't
3:59:14 want the general public to get them." Or chips,
3:59:17 even AI chips, but you see how Huawei makes chips now.
3:59:22 It took a few years, but...
3:59:24 and I don't think there is a way you can contain knowledge like that.
3:59:29 I think in this day and age, it is impossible, like the internet.
3:59:34 I don't think this is a possibility.
3:59:38 On the Manhattan Project thing, I think that a Manhattan Project-like thing
3:59:42 for open models would be pretty reasonable, because it wouldn't cost that much.
3:59:46 But I think that will come.
3:59:48 It seems like culturally, the companies are changing.
3:59:51 But I agree with Sebastian on all of that.
3:59:54 I don't see it happening nor being helpful.
3:59:59 Yeah.
3:59:59 The motivating force behind the Manhattan Project was civilizational risk.
4:00:03 It's harder to motivate that for open-source models.
4:00:08 There's no civilizational risk.
4:00:10 On the hardware side, we mentioned NVIDIA a bunch of times.
4:00:15 Do you think Jensen and NVIDIA will keep winning?
4:00:19 I think they have to iterate and manufacture a lot.
4:00:22 And I think they probably...
4:00:25 what they're doing, they do innovate, but I think there's always the chance
4:00:31 that someone does something fundamentally different,
4:00:34 gets very lucky, and then does something.
4:00:37 But the problem is adoption.
4:00:39 The moat of NVIDIA is probably not just the GPU.
4:00:43 It's more like the CUDA ecosystem, and that has evolved over two decades.
4:00:47 Even back when I was a grad student,
4:00:50 I was in a lab doing biophysical simulations,
4:00:53 molecular dynamics, and we had a Tesla GPU back then just for the computations.
4:00:57 It was about 15 years ago now.
4:01:00 And they built this up for a long time, and that's the moat, I think.
4:01:05 It's not the chip itself,
4:01:07 although they have the money to iterate, build, and scale.
4:01:11 But then it's really about compatibility.
4:01:14 If you're at that scale, why would you go with something risky where there
4:01:19 are only a few chips they can make per year?
4:01:21 You go with the big one.
4:01:23 But then I do think with LLMs now,
4:01:25 it will be easier to design something like CUDA.
4:01:30 It took 15 years because it was hard,
4:01:32 but now that we have LLMs, we can maybe replicate CUDA.
4:01:36 And I wonder if there will be a separation
4:01:38 of training and inference compute as we stabilize,
4:01:42 and more compute is needed for inference.
4:01:47 That's supposed to be the point of the Groq acquisition.
4:01:50 And that's why part of what Vera Rubin is-
4:01:52 where they have a new chip with no high-bandwidth memory,
4:01:54 which is one of the- or very little, which is one of the most expensive pieces.
4:01:59 It's designed for pre-fill, which is the part of inference where
4:02:03 you essentially do a lot of matrix multiplications.
4:02:05 And then you only need the memory
4:02:07 when you're doing this autoregressive generation,
4:02:09 and you have the KV cache swaps.
4:02:11 So they have this new GPU that's designed for that specific use case,
4:02:15 and then the cost of ownership per FLOP or whatever is actually way lower.
4:02:20 But I think that NVIDIA's fate lies in the diffusion of AI still.
4:02:25 Their biggest clients are still these hyperscale companies.
4:02:29 Like, Google obviously can make TPUs.
4:02:32 Amazon is making Trainium.
4:02:34 Microsoft will try to do its own things.
4:02:37 And so long as the pace of AI progress is high,
4:02:40 NVIDIA's platform is the most flexible and people will want that.
4:02:43 But if there's stagnation, then creating bespoke chips,
4:02:47 there's more time to do it.
4:02:50 It's interesting that NVIDIA is quite active
4:02:53 in trying to develop all kinds of different products.
4:02:56 They try to create areas of commercial value that will use a lot of GPUs.
4:03:01 Mm-hmm.
4:03:02 But they keep innovating and they're doing a lot of incredible research, so...
4:03:07 Everyone says the company's super oriented around
4:03:09 Jensen and how operationally plugged in he is.
4:03:12 And it sounds so unlike many other big companies that I've heard about.
4:03:16 And so long as that's the culture,
4:03:18 I think that we can expect that to keep progress happening.
4:03:21 And it's like he's still in the Steve Jobs era of Apple.
4:03:24 So long as that is how it operates,
4:03:27 I'm pretty optimistic for their situation because it's like,
4:03:32 it is their top-order problem,
4:03:33 and I don't know if making these chips for the whole
4:03:37 ecosystem is the top goal of all these other companies.
4:03:39 They'll do a good job, but it might not be as good of a job.
4:03:43 Since you mentioned Jensen,
4:03:45 I've been reading a lot about history and about singular figures in history.
4:03:49 What do you guys think about the single man/woman view of history?
4:03:53 How important are individuals for steering
4:03:55 the direction of history in the tech sector?
4:03:58 So, you know, what's NVIDIA without Jensen?
4:04:01 You mentioned Steve Jobs.
4:04:03 What's Apple without Steve Jobs?
4:04:05 What's xAI without Elon or DeepMind without Demis?
4:04:12 People make things earlier and faster, whereas scientifically,
4:04:17 many great scientists credit being in the right place
4:04:19 at the right time and still making the innovation,
4:04:22 where eventually someone else will still have the idea.
4:04:25 So I think that in that way,
4:04:29 Jensen is helping manifest this GPU revolution much faster and much
4:04:34 more focused than it would happen without having a person there.
4:04:38 And this is making the whole AI build-out faster.
4:04:40 But I do still think that eventually, something like ChatGPT would have happened
4:04:45 and a build-out like this would have happened,
4:04:47 but it probably would not have been as fast.
4:04:50 I think that's the sort of flavor that is applied.
4:04:55 These individual people, there are people who are placing bets on something.
4:04:58 Some get lucky, some don't.
4:04:59 But if you don't have these people at the helm, it would be more diffused.
4:05:02 It's almost like investing in an ETF versus individual stocks.
4:05:06 Individual stocks might go up or down more heavily than an ETF,
4:05:11 which is more balanced.
4:05:12 It will eventually go up over time.
4:05:13 We'll get there.
4:05:14 But it's just like, you know, the focus I think is the thing.
4:05:18 Passion and focus.
4:05:20 Isn't there a real case to be made that without Jensen,
4:05:22 there's not a reinvigoration of the deep learning revolution?
4:05:27 It could've been 20 years later, is what I would say.
4:05:30 Or like another AI winter could have come if GPUs weren't around.
4:05:35 That could change history completely because you could think
4:05:37 of all the other technologies that could've come in the meantime,
4:05:42 and the focus of human civilization would get...
4:05:44 Silicon Valley would be captured by different hype.
4:05:48 But I do think there's certainly an aspect
4:05:50 where it was all planned, the GPU trajectory.
4:05:53 But on the other hand, it's also a lot of lucky coincidences or good intuition.
4:05:58 Like the investment into, let's say, biophysical simulations.
4:06:01 I mean, I think it started with video games and then it just happened
4:06:05 to be good at linear algebra because
4:06:07 video games require a lot of linear algebra.
4:06:09 And then you have the biophysical simulations.
4:06:11 But still, I don't think the master plan was AI.
4:06:16 I think it happened to be Alex Krizhevsky.
4:06:19 So someone took these GPUs and said, "Hey,
4:06:22 let's try to train a neural network on that." It happened to work really well,
4:06:26 and I think it only happened because you could purchase those GPUs.
4:06:30 Gaming would've created a demand for faster processors if
4:06:33 NVIDIA had gone out of business in the early days.
4:06:37 That's what I would think.
4:06:37 I think that the GPUs would've been different,
4:06:42 but I think GPUs would still exist at the time
4:06:46 of AlexNet and at the time of the Transformer.
4:06:48 It was just hard to know if it would be
4:06:51 one company as successful or multiple smaller companies with worse chips.
4:06:55 But I don't think that's a 100-year delay.
4:06:59 It might be a decade delay.
4:07:01 Well, it could be one, two, three, four, five-decade delay.
4:07:04 I just can't see Intel or AMD doing what NVIDIA did.
4:07:08 I don't think it would be a company that exists.
4:07:11 I think it would be a different company that would rise.
4:07:13 Like Silicon Graphics or something.
4:07:15 So yeah, some company that has died would have done it.
4:07:19 But just looking at it, it seems like these singular figures,
4:07:23 these leaders, have a huge impact on the trajectory of the world.
4:07:28 Obviously, there are incredible teams behind them.
4:07:31 But, you know, having that kind of very singular,
4:07:36 almost dogmatic focus- -is necessary to make progress.
4:07:41 Yeah, I mean, even with GPT, it wouldn't exist if there wasn't a person,
4:07:44 Ilya, who pushed for this scaling, right?
4:07:47 Yeah, Dario Amodei was also deeply involved in that.
4:07:50 If you read some of the histories from OpenAI,
4:07:52 it seems wild thinking about how early these people were like,
4:07:55 "We need to hook up 10,000 GPUs and take all of OpenAI's compute and train
4:07:58 one model." There were a lot of people who didn't want to do that.
4:08:02 Which is an insane thing to believe.
4:08:05 To believe in scaling before scaling has
4:08:07 any indication that it's going to materialize.
4:08:10 Again, singular figures.
4:08:12 Speaking of which, 100 years from now,
4:08:16 this is presumably post-singularity, whatever singularity is.
4:08:21 When historians look back at our time now,
4:08:24 what technological breakthroughs would they really emphasize
4:08:28 as the breakthroughs that led to the singularity?
4:08:32 So far we have Turing to today, 80 years.
4:08:37 I think it would still be computing,
4:08:39 like the umbrella term "computing." I don't necessarily think
4:08:42 that in 100 or 200 years it would be AI.
4:08:46 It could still very well be computers.
4:08:48 We are now taking better advantage of them, but the fact of computing remains.
4:08:54 It's basically a Moore's Law discussion.
4:08:56 Even the details of CUDA and GPUs won't even be remembered,
4:09:00 nor will all this software turmoil.
4:09:04 It'll just be, obviously, compute.
4:09:07 I generally agree, but is the connectivity
4:09:10 of the internet and compute able to be merged?
4:09:14 Or is it both of them?
4:09:18 I think the internet will probably be related to communication.
4:09:21 It could be a phone, the internet, or satellites.
4:09:25 Compute is more like the scaling aspect of it.
4:09:29 It's possible that the internet is completely forgotten-
4:09:32 -that the internet is wrapped into phone networks, like communication networks.
4:09:38 This is just another manifestation
4:09:40 of that, and the real breakthrough comes from increased compute,
4:09:44 or Moore's Law, broadly defined.
4:09:46 Well, I think the connection of people is very fundamental to it.
4:09:50 it's like, you can talk to anyone.
4:09:52 You want to find the best person in the world for something,
4:09:56 they are somewhere in the world.
4:09:57 And being able to have that flow of information—the AIs will also rely on this.
4:10:02 I've been fixating on when I said
4:10:04 the dream was dead about the one central model.
4:10:07 The thing that is evolving is people having many agents for different tasks.
4:10:11 People already started doing this with different clouds.
4:10:14 It's described as many AGIs in the data center
4:10:18 where each one manages and they talk to each other.
4:10:21 And that is reliant on networking and the free flow of information.
4:10:26 on top of compute.
4:10:27 But networking, especially with GPUs, is such a part of scaling of compute.
4:10:33 The GPUs and the data centers need to talk to each other.
4:10:36 Anything about neural networks will be remembered?
4:10:39 Like, do you think there's something very specific and singular
4:10:42 to the fact that it's neural networks that's seen as a breakthrough,
4:10:46 like a genius, that you're basically replicating,
4:10:48 in a very crude way, the human mind?
4:10:51 The structure of the human brain, the human mind?
4:10:54 I think without the human mind, we probably wouldn't have neural networks,
4:10:58 because it just was an inspiration for that.
4:11:01 But on the other end, I think it's just so, so different.
4:11:04 I mean, it's digital versus biological,
4:11:06 that I do think it will probably be more grouped as an algorithm.
4:11:12 That's massively parallelizable...
4:11:13 ...On this particular kind of compute?
4:11:15 It could have been like genetic computing; genetic algorithms just parallelized.
4:11:19 It just happens that this is more efficient and works better.
4:11:23 And it very well could be that the LLM, the neural networks,
4:11:26 the way we architect them now is just
4:11:29 a small component of the system that leads to singularity.
4:11:34 If you think of it in 100 years, I think society can be changed more
4:11:38 with more compute and intelligence because of autonomy.
4:11:41 But looking at this, what are the things
4:11:45 from the Industrial Revolution that we remember?
4:11:47 We remember the engine,
4:11:48 which is probably the equivalent of the computer in this.
4:11:51 But there's a lot of other physical transformations that people
4:11:55 are aware of, like the cotton gin and all these things,
4:11:59 these machines that are still known: air conditioning, refrigerators.
4:12:04 Some of these things from AI will still be known.
4:12:08 The word "transformer" could still be known.
4:12:11 I would guess that deep learning is definitely still known,
4:12:14 but the transformer might be evolved away
4:12:16 from in 100 years with AGI researchers everywhere.
4:12:21 But I think deep learning is likely to be a term that is remembered.
4:12:28 And I wonder what the air conditioning
4:12:30 and refrigeration of the future is that AI brings.
4:12:32 If we travel forward 100 years from now,
4:12:34 we transport there right now, what do you think is different?
4:12:37 How do you think the world looks different?
4:12:40 First of all, do you think there are humans?
4:12:42 Do you think there are robots everywhere walking around?
4:12:46 I do think specialized robots, for sure, for certain tasks.
4:12:49 Humanoid form?
4:12:51 Maybe half-humanoid.
4:12:52 We'll see.
4:12:54 I think for certain things, yes,
4:12:55 there will be humanoid robots because it's just amenable for the environment.
4:13:00 But for certain tasks, it might make sense.
4:13:03 What's harder to imagine is how we interact
4:13:05 with the devices and what humans do with devices.
4:13:08 Well, I mean, I'm pretty sure it will
4:13:11 probably not be the cellphone or the laptop.
4:13:13 Will it be implants?
4:13:16 I mean, it has to be brain-computer interfaces, right?
4:13:18 I mean, 100 years from now,
4:13:20 given the progress we're seeing now— there has to be...
4:13:25 unless there's legitimately a complete alteration
4:13:29 of how we interact with reality.
4:13:33 On the other hand, cars are older than 100 years, right?
4:13:36 And it's still the same interface.
4:13:38 We haven't replaced cars with something else.
4:13:41 We just made them better,
4:13:42 but it's still a steering wheel, still wheels, you know?
4:13:45 I think we'll still carry around a physical brick
4:13:47 of compute because people want some ability to have a private...
4:13:51 Like, you might not engage with it as much as a phone,
4:13:55 but having private information that is yours
4:13:56 as an interface between the rest of the internet, I think that will still exist.
4:14:01 It might not look like an iPhone and it might be used a lot less,
4:14:05 but I still expect people to carry things around.
4:14:08 Why do you think the smartphone is the embodiment of private?
4:14:11 There's a camera on it.
4:14:14 There's—- Private for you, like encrypted messages, encrypted photos...
4:14:19 know what your life is.
4:14:22 I guess it's a question of how optimistic on brain-machine interfaces you are.
4:14:26 Is all that just going to be stored in the cloud?
4:14:29 Your whole calendar?
4:14:30 It's hard to think about processing all the information that we can process
4:14:37 visually through brain-machine interfaces presenting something
4:14:40 like a calendar or something to you.
4:14:44 It's hard to just think about knowing, without looking, your email inbox.
4:14:49 Like you signal to a computer and then you just know your email inbox.
4:14:53 Is that something that the human brain
4:14:55 can handle being piped into it non-visually?
4:14:58 I don't know exactly how those transformations happen.
4:15:03 Humans aren't changing in 100 years.
4:15:05 I think agency and community are things that people actually want.
4:15:09 A local community, yeah.
4:15:10 People you are close to, being able to do things with them and being
4:15:15 able to ascribe meaning to your life and being able to do things.
4:15:22 In 100 years, I don't think that human biology is changing
4:15:26 away from those on a time scale that we can discuss.
4:15:30 And I think that UBI does not solve agency.
4:15:34 I do expect mass wealth, and I hope that it has spread so
4:15:38 that the average life looks very different in 100 years.
4:15:42 But that's still a lot to happen.
4:15:44 If you think about countries that are early
4:15:47 in their development process to getting access to computing and internet,
4:15:52 to build all the infrastructure and have policy
4:15:57 that shares one nation's wealth with another is...
4:16:01 I think it's an optimistic view to see all that happening in 100 years- ...while
4:16:06 they are still independent entities and not just
4:16:10 like absorbed into some international order by force.
4:16:13 But there could be just better, more elaborate, more effective...
4:16:18 social support systems that help alleviate some
4:16:22 levels of basic suffering from the world.
4:16:24 You know, the transformation of society where a lot
4:16:26 of jobs are lost in the short term, I think we have to really remember that each
4:16:31 individual job that's lost is a human being who's suffering.
4:16:35 That's like a...
4:16:38 When jobs are lost, the scale is a real tragedy.
4:16:41 You can make all kinds of arguments about
4:16:43 economics or how it's all going to be okay.
4:16:46 It's good for the GDP, there's going to be new jobs created.
4:16:50 Fundamentally at the individual level for that human being,
4:16:54 that's real suffering.
4:16:55 That's a real personal sort of tragedy.
4:16:58 And we have to not forget that as the technologies are being developed.
4:17:03 And also my hope for all the AI slop we're seeing is that there will be
4:17:10 a greater and greater premium for the fundamental
4:17:14 aspects of the human experience that are in-person.
4:17:17 The things that we all...
4:17:19 Like seeing each other, talking together in-person.
4:17:23 The next few years are definitely going to be an increased
4:17:26 value on physical goods and events— ...and even more pressure on slop.
4:17:32 So it'll be...
4:17:33 the slop is only starting.
4:17:35 The next few years will be more and more diverse ...versions of slop.
4:17:38 They would be drowning in slop.
4:17:39 Is that what—- So I'm hoping that society drowns
4:17:42 in slop enough to snap out of it and be like, "We can't deal with it.
4:17:47 It just doesn't matter." And then, the physical has such a higher premium on it.
4:17:54 Even like classic examples, I honestly think this is true,
4:17:57 and I think we will get tired of it.
4:17:59 We are already kind of tired of it.
4:18:01 I mean, even art.
4:18:02 I don't think art will go away.
4:18:04 You have paintings, physical paintings.
4:18:06 There's more value, not just monetary value,
4:18:10 but just more value appreciation for the actual
4:18:13 painting than a photocopy of that painting.
4:18:14 It could be a perfect digital reprint,
4:18:16 but there is something when you go to a museum and you look
4:18:19 at that art and you see the real thing and you just think, "Okay.
4:18:22 A human." It's like a craft.
4:18:24 You have like an appreciation for that.
4:18:26 And I think the same is true for writing,
4:18:27 for talking, for any type of experience...
4:18:32 I do unfortunately think it will be like a dichotomy,
4:18:36 like a fork where some things will be automated.
4:18:40 Like, you know, there are not as many paintings as there used to be,
4:18:42 you know, 200 years ago.
4:18:43 There are more photographs, more photocopies.
4:18:46 But at the same time, it won't go away.
4:18:49 There will be value in that.
4:18:51 I think the difference will just be, you know, what's the proportion of that.
4:18:56 But personally, I have a hard time reading
4:18:59 things where I obviously see it's obviously AI generated.
4:19:02 I'm sorry.
4:19:03 It might—it might be really good information there, but I'm just like, "Nah,
4:19:08 not for me."- I think eventually they'll fool you,
4:19:10 and it'll be on platforms that give ways of verifying or building trust.
4:19:15 So you will trust that Lex is not AI generated, having been here.
4:19:19 So then you have trust in this channel.
4:19:22 But it's harder for new people who don't have that trust.
4:19:25 Well, that will get interesting because I think fundamentally it's a solvable
4:19:30 problem by having trust in certain outlets that they won't do it,
4:19:35 but it's all going to be trust-based.
4:19:37 There will be systems to authorize, "Okay, this is real.
4:19:39 This is not real." There will be some telltale signs where
4:19:42 you can obviously tell this is AI generated and this is not.
4:19:45 But some will be so good that it's hard to tell, and then you have to trust.
4:19:50 And well, that will get interesting and a bit problematic.
4:19:54 The extreme case of this is to watermark all human content.
4:19:57 So all photos that we take on our own have
4:20:00 some watermark until they are edited or something like this.
4:20:03 And software can manage communications with the device
4:20:07 manufacturer- device manufacturer to maintain human
4:20:10 editing— which is the opposite of the discussion to try to watermark AI images.
4:20:15 And then you can make a Google image that has
4:20:18 a watermark and use a different Google tool to remove it.
4:20:21 Yep.
4:20:21 It's going to be an arms race, basically.
4:20:23 And we've been mostly focusing on the positive aspects of AI.
4:20:27 All the capabilities that we've been talking about can be used
4:20:32 to destabilize human civilization with even
4:20:34 just relatively dumb AI applied at scale,
4:20:39 and then further, superintelligent AI systems.
4:20:42 Of course, there's the sort of doomer take
4:20:45 that's important to consider as we develop these technologies.
4:20:50 What gives you hope about the future of human civilization,
4:20:53 given everything we've been talking about?
4:20:56 Are we going to be okay?
4:20:59 I think we will.
4:21:00 I'm definitely a worrier, both about AI and non-AI things.
4:21:04 But humans do tend to find a way.
4:21:08 I think that's what humans are built for: to have
4:21:11 community and find a way to figure out problems.
4:21:14 That's what has gotten us to this point.
4:21:16 And to think that the AI opportunity and related technologies is really big.
4:21:23 And I think that there's big social
4:21:25 and political problems to help everybody understand that.
4:21:29 And I think that's what we're staring at a lot of right now,
4:21:33 is like the world is a scary place, and AI is a very uncertain thing.
4:21:36 And it takes a lot of work that is not necessarily building things.
4:21:41 It's like telling people and understanding people,
4:21:44 that the people building AI are historically not motivated or wanting to do.
4:21:50 But it is something that is probably doable.
4:21:52 It just will take longer than people want.
4:21:55 And we have to go through that long period of like hard,
4:22:00 distraught AI discussions if we want to have the lasting benefits.
4:22:05 Yeah.
4:22:05 Through that process,
4:22:06 I'm especially excited that we get a chance to better understand ourselves,
4:22:12 us at the individual level as humans and at the civilization level,
4:22:16 and answer some of the big mysteries,
4:22:18 like what is this whole consciousness thing going on here?
4:22:23 It seems to be truly special.
4:22:24 Like, there's a real miracle in our mind.
4:22:27 And AI puts a mirror to ourselves and we
4:22:29 get to answer some of the big questions about like,
4:22:33 what is this whole thing going on here?
4:22:35 Well, one thing about that is also what I do think makes us
4:22:39 very different from AI and why I don't worry about AI taking over is,
4:22:44 like you said, consciousness.
4:22:45 We humans, we decide what we want to do.
4:22:47 AI in its current implementation, I can't see it changing.
4:22:51 You have to tell it what to do.
4:22:54 And so you have still the agency.
4:22:56 It doesn't take the agency from you because it becomes a tool.
4:23:00 You can think of it as a tool.
4:23:01 You tell it what to do.
4:23:03 It will be more automatic than other previous tools.
4:23:06 It's certainly more powerful than a hammer,
4:23:08 it can figure things out, but it's still you in charge, right?
4:23:12 So the AI is not in charge, you're in charge.
4:23:15 You tell the AI what to do and it's doing it for you.
4:23:18 So in the post-singularity, post-apocalyptic war between humans and machines,
4:23:22 you're saying humans are worth fighting for?
4:23:27 100%.
4:23:27 I mean, this is...
4:23:29 The movie Terminator, they made in the '80s, essentially, and I do think,
4:23:33 well, the only thing I can see going wrong is,
4:23:38 of course, if things are explicitly programmed
4:23:40 to do the thing that is harmful, basically.
4:23:43 I think actually in that, in a Terminator type of setup, I think humans win.
4:23:49 I think we're too clever.
4:23:51 It's hard to explain how we figure it out, but we do.
4:23:56 And we'll probably be using local LLMs,