Nemotron Dev Days Seoul: Sovereign AI, Hackathon Highlights & Build-a-Claw | Nemotron Labs

Nemotron Dev Days Seoul: Sovereign AI, Hackathon Highlights & Build-a-Claw | Nemotron Labs

NVIDIA Developer

1:03 Hello and welcome everybody to a very special NeMo Tron Labs

1:07 live stream where we are joining you live from Seoul, South Korea.

1:13 We are very excited for this for this and we're joined by very special guest,

1:19 our host today, Chung Won.

1:21 Chung Won, take it away.

1:22 Hi, I'm Chung Won.

1:24 Thank you for coming here.

1:25 Thank you for visiting Korea and we are

1:28 having a first live stream in Seoul, right?

1:31 First of many, I hope.

1:32 Yeah.

1:33 Uh so, over the last few days,

1:36 we have been here to do something called the NeMo Tron Developer Day,

1:42 which has been very exciting so far.

1:45 Deep dive.

1:46 Go ahead.

1:47 Yeah, deep dive sessions with from the HQ

1:51 NVIDIA researchers including Brian BPA as well.

1:55 That's right.

1:56 Brian Catanzaro made an appearance.

1:57 It was pretty epic.

1:59 We're we're going to get more to that in just a second, guys,

2:01 but in order to set up kind of the vibe

2:03 of what it was like here over the last few days,

2:05 we're going to kick off to a video that showcased

2:07 what the first day of of that developer day looked like.

2:11 So, Zach, if you want to go ahead and play that video.

2:17 [music] [music] [music] Cool.

2:48 Very cool.

2:49 Very cool.

2:51 It was it was a lot of fun to be there.

2:53 A lot of very interesting projects that we saw over the the few days.

2:59 Chung Won, tell us a little bit about, you know,

3:02 how you set this up and and kind of like

3:04 what the idea was to bring as much you know,

3:09 energy and and kind of the the vibe of NeMo

3:13 Tron to the big the good people of Seoul.

3:16 Yeah, actually Korea is one of the most

3:21 important country about sovereign AI because we are building

3:25 models from the scratch and there are a lot

3:28 of model business about fine-tuning or post-training as well.

3:32 So, and especially we have many industries including manufacturing.

3:37 So, it is there there are a lot of area we can apply our NeMo Tron ecosystem.

3:45 So, we prepared the deep dive sessions and we had

3:49 a hackathon and we had a build a claw session.

3:53 So, it was amazing full of NeMo Tron everything like

3:57 model and SDKs and training strategies and even NeMo claw, right?

4:02 That's right.

4:03 Yeah, I mean it was it was a lot of fun just

4:06 to see everybody building with NeMo Tron or the NVIDIA ecosystem of course,

4:12 but also so exciting to see how passionate and skilled,

4:18 you know, the Korean devs are when it comes to leveraging technology.

4:22 We had a lot of people join for the build a claw event, you know,

4:26 where they were getting kind of started

4:27 with claws on their own hardware or using, you know, cloud instances.

4:31 You know, at the end of the day,

4:32 I think this was a incredible experience for people joining.

4:37 Guys, don't forget if you have any questions for any

4:40 of the people you're going to see here today,

4:42 don't forget to drop them in the chat.

4:44 We're going to be having some presentations and so when we have

4:48 those, please feel free to ask the presenters any questions that you have.

4:52 Chung Won, one of the things I noticed during the deep dive sessions, right?

4:57 Is that a lot of people were not

4:59 just like listening and nodding along, but taking notes.

5:02 Really really they had really great questions afterwards.

5:05 You know, how [clears throat] do you think people are adopting AI in Korea?

5:13 How do you think that impacts how, you know,

5:16 us as NVIDIA need to communicate to to the Korean developers?

5:22 Actually, some of them are building their own model

5:26 or some of them are using AI for building agents,

5:31 but this time we had a really deep dive session.

5:35 So, during that, I think when we are just building our AI model or agents,

5:44 we are sometimes we are competing with other companies like we are developers.

5:50 We know opening our model authors or strategies is important are important,

5:57 but in company it's really difficult to open our knowledge.

6:02 So, but in this event, we shared all of our NVIDIA strategy for training models.

6:12 So, it was really really impressed by Korean developers.

6:17 I heard a lot of recommendation or reviews

6:23 from them and I really excited about that experience.

6:27 Yeah, it was it was really cool and I

6:29 mean what was also cool obviously was the hackathon.

6:33 Seeing so many people building with tools over

6:35 48 hours and the projects were extraordinarily ambitious.

6:40 And not just ambitious, but like they kind of like got her done.

6:44 We're actually going to see some of these teams who won.

6:49 So, can you outline for us a little

6:51 bit like what the what the hackathon structure was?

6:55 You know, there's this idea of three tracks.

6:57 Do you mind like outlining what that was all about?

7:00 Actually, we had a two days for the hackathon.

7:04 It was really really short time to do,

7:07 but we planned three stacks including create created creating a builder agent

7:15 creative use case and a fine-tuning models and create a data set.

7:21 So, synthesize the data set.

7:23 So, it was really short time to do that.

7:26 But everyone made really great project and they will bring

7:31 to their company or their individual project after the hackathon as well.

7:37 Yeah, it was really incredible to see the amount

7:39 of work that all of the teams were able to do.

7:41 We're going to see the winners today and we're going

7:43 to get a chance to see their presentation what they actually created,

7:47 but I do want to say that all of the projects were strong.

7:50 It was it was a hard decision to determine

7:52 the winners because there was so much talent over that 48

7:56 hours and I was just absolutely astounded by the level

8:00 of ambition and and kind of greatness of the projects,

8:03 but we've got the three best projects that we saw at the hackathon here

8:07 one from each track and so we're actually going to get our first guest.

8:11 Do you mind introducing them?

8:13 Yeah, he's from Naver and they made a creative agent.

8:19 Please join us.

8:20 Yes.

8:22 Zach, if we can get that picture up showing them holding the spark wooden.

8:30 Hi.

8:32 I'm Han Su and I'm from Naver.

8:51 Is this slide on, Zach?

8:54 You're good.

8:55 Okay.

8:57 Please introduce yourself.

8:59 Okay, my name is Han Su and just you can call me just Han Su.

9:03 I've been working at Naver as a AI engineer and I'm

9:06 very glad to share my experience in about my hackathon.

9:10 So, I'm glad to see you.

9:13 Yeah, we can just start it with your stuff.

9:17 Please introduce your projects.

9:19 Okay.

9:21 My project is the NeMo briefing.

9:23 The NeMo means the new model.

9:25 So, NeMo briefing is an AI agent to collect

9:29 the users' reactions and summarize and deliver the report about that.

9:36 So, I started.

9:40 Okay, the first I will introduce my team and my company and I

9:47 I joined the this hackathon with my awesome teammates Song Jin and In Gon.

9:51 They are very awesome AI developers and so,

9:56 our let's introduce our company, the neighbor.

9:59 The is the biggest IT company in South

10:02 Korea and the competitor to Google in search engine.

10:06 We have the rather many business models such as place or B2B or AI or cloud.

10:13 So, let's start it.

10:15 Uh Our key idea is started from our internal work project.

10:20 We built some leaderboard to evaluate the various LLMs.

10:25 They are accuracy and performance.

10:28 But, we found that there is a significant gap

10:32 between the benchmark score and the real users' reactions.

10:36 So, we thought it was really good to build an AI agent

10:41 to collect the real users' feedback

10:44 and reactions and add them to our leaderboard.

10:48 So, we started it.

10:51 Yeah, it's the architecture of Nemo briefing.

10:55 First, we build a query generator agent for the refining

11:00 the query and for the high-quality search result.

11:05 So, the second part is the collector agent and it

11:09 is our key point and inside the each collector,

11:12 we had a validator and validator filter out the unnecessary items.

11:18 And we collect various from the various source such as Reddit,

11:23 Rob Stars, or Hugging Face community.

11:25 And the final uh part is the reporter agent.

11:28 The reporter agent does aggregate the whole

11:31 data and summarize them as a markdown report.

11:35 This whole part is orchestrated by Nemo Transformer model with React prompt.

11:41 So, we used every NVIDIA steps in our project.

11:46 First, we used Nemo Transformer model.

11:48 We used super model for difficult tasks such as React or validation.

11:53 And we used nano model for simple tasks such as query generation.

11:57 And these models are deployed on Brave Cloud and with the Nemo framework.

12:03 It is very useful and easy to use.

12:05 And we developed our agent with Nemo agent toolkit and known as NAT.

12:11 I think it is really awesome framework.

12:13 I I've used other agent frameworks such as LangChain or ADK,

12:16 but it is really cool and easy to use and extendable.

12:20 I really like it this configuration.

12:24 So, let's see the demo demo video.

12:29 I searched the Nemo Transformer model

12:33 and query generator starts and I skip them.

12:37 The data collector works but in the parallel and they

12:41 collect the each data and validate them in parallel.

12:44 So, I skip it.

12:47 The final section is reporter.

12:49 It takes some long time, so I skip that.

12:52 This is a final report.

12:55 So, let's see type in the example report about the Nemo Transformer model.

13:02 The report say that some users have

13:04 some trouble in the configuring the reasoning mode,

13:08 but I think it is not the model problem.

13:10 It's just about the issue with the framework or template,

13:14 but it is very useful to catch catch catch this gap with our real time.

13:22 I think Nemo briefing can help this.

13:27 Uh Okay, so final session and we don't have much time.

13:32 It is a hackathon and 36 hour isn't enough, so we Here is the future work.

13:39 I will skip that.

13:40 It is not important.

13:42 Okay.

13:43 I'm done.

13:44 Great.

13:45 Great project.

13:47 This was the winner of the track A.

13:52 Do you have any questions?

13:56 Please shout out the text.

13:59 Hello.

14:01 Congrats Hansu and team.

14:03 Great demo.

14:10 Yeah, um during the hackathon you met uh

14:14 NVIDIA researchers and you had a good great mentoring.

14:20 Any great experience from that?

14:22 Okay, they they helped me and they helped our team to select

14:28 the topic proper topic and make right architecture at the proper time.

14:36 Their advice is really helpful for me.

14:38 Yeah, thank you.

14:40 And actually Nemo is a neural models but neural

14:44 models but you saw you saw the new model, right?

14:47 Right.

14:48 Yeah.

14:49 Why did you think like that?

14:51 Because our project is to collect the real

14:55 user reaction from the new model our new models.

14:59 We used the Nemo.

15:00 Are you going to use it for the automatically like agents?

15:04 Mhm.

15:04 Right?

15:05 Right.

15:06 For the your company research or some?

15:09 Yes, I developed some agent in our company such as AI nation or agent.

15:16 You will open to public or you will use it

15:18 on your company I it released at March in this year.

15:23 You can use it neighbor.

15:25 And everyone can use it.

15:27 It's good for the government.

15:28 Ah, I see.

15:30 It's great for researchers, I think.

15:32 What all model builders or model users.

15:36 Thank you for the great project.

15:38 Thank you.

15:39 Thank you, Hansu.

15:41 So, we are going to move to to next uh track.

15:48 Track B.

15:50 Uh the winner was from the Pyla.

15:53 Will can you join us?

15:55 Jack, please show the winner's image.

15:59 Thank you.

16:15 Oh, I didn't mention the important thing.

16:21 Actually, the prize was the Spark.

16:23 So, the [laughter] You can see the Spark on the image, right?

16:29 Yeah.

16:30 Thank you.

16:31 Thank you for coming in.

16:32 Please introduce yourself.

16:33 Uh yeah, hello.

16:34 I'm from Pyla AI and my name is Junsu Kim.

16:38 We uh did we winning the track

16:41 B about the fine-tuning open-source Nemo Transformer model.

16:45 Thank you.

16:45 Please introduce your project as well.

16:48 Yeah.

16:49 Okay.

16:50 Let me start.

16:50 Yeah, hello everyone.

16:52 I'm from Pyla AI and today I'll talk about the fine-tuning NVIDIA's

16:57 Nemo Transformer nano into a fine-grained

16:58 trustworthy video understanding model for content safety.

17:03 So, when you ask a model to classify a video for safety,

17:07 you will really have two choices.

17:09 First, you can feed the whole video it at once and get a single label.

17:14 But, in this case you lose any

17:16 sense of when the harmful content actually appears.

17:19 Or, you can split the video into chunks and analyze each one,

17:23 which give you fine-grained temporally grounded results.

17:26 We want to see how much this second approach changed

17:30 model behavior and how far we can push it with fine-tuning.

17:35 For data, we build our training set on top of Safe Watch,

17:39 which published in ICLR 2025.

17:41 That gave us about 8,000 videos and 25K chunks for training

17:46 plus a test set of 1,200 videos with video-level descriptions and labels.

17:53 So, our first step is supervised fine-tuning.

17:55 We run both LoRA and full fine-tuning on Nemo Transformer

17:58 nano 12B V2 VL model using the Megatron-Bridge SFT pipeline.

18:04 The goal is to teach the model to produce

18:07 chunk-level explanations instead of a single video-level summary.

18:14 For post-training, we wanted to do preference optimiz- preference optimizations.

18:19 So, but Nemo doesn't officially support the Nemo Transformer nano V2 VL yet,

18:23 so we implement simple and DPO ourselves directly on the top of Megatron-Bridge.

18:29 Uh due to time constraint,

18:31 the results I'm sharing today are from simple training method.

18:34 The reasoning of this preference optimization is open-source models

18:38 are ready to a reasonable job on whole video understanding,

18:42 but they struggle with temporal grounding.

18:44 So, we use segment-level temporally grounded output as the winning

18:47 data and whole video only description as the losing data.

18:53 Yeah, so for evaluation, we put together three benchmarks.

18:56 First, we split Safe Watch training set into train

19:00 and evaluation data set using TransNet for temporal grounding.

19:04 The second, because the Safe Watch test set only has video-level description,

19:08 we use NVIDIA's AI blueprint VSS

19:10 to pseudo-labeling the chunk-level explanations and classifications.

19:15 And finally, to stress test longer videos,

19:17 we take video MME clips and inject harmful segment in the middle,

19:21 so the model have to quickly temporally grounded the injected portion.

19:27 Uh we evaluated on two metrics.

19:29 First is average F1 for accuracy and second

19:31 TIOU TIOU for the temporal grounding ability.

19:36 So, our baseline pipeline is straightforward.

19:40 Nemo Transformer by a VLM takes the whole

19:42 video in one shot and outputs three things:

19:45 the description and category label and six categories and explanation.

19:49 So, the baseline results on the open source model

19:53 is like 0.25 0.5 real is like our approach 0.318.

19:59 The finally, the nano 12B V2 real model

20:02 showed the best performance on Safe for us benchmark.

20:06 So, our approach at chunk level inference on top here.

20:10 We train on Megatron Bridge expert to hugging face and service real LM.

20:15 For each temporal chunk,

20:16 the model emits a description, godly labels, and explanations.

20:23 Yeah, so on the Safe for us evaluation split,

20:25 the baseline Nemo Transit as 0.5 of about 0.5 F1,

20:30 the chunk level prompting alone jumped at that to 0.75.

20:34 The Laura's SFT more pushed up performance of this model, and finally,

20:40 the fine-tuning with the full fine-tuning with chunk

20:43 level approach the best performance on temporal grounding.

20:47 However, about the accuracy, we it degrade the performance,

20:51 but we should find more accurate hyperparameter.

20:56 And for temporal grounding results,

20:58 uh when uh for full top fine-tuning and Laura

21:02 also increased the performance of T I O U,

21:05 and also reinforcement learning also increased our performance, too.

21:12 And this is our qualitative example.

21:14 So, we cannot show the exact Safe for us sample directly.

21:17 The content is genuinely harmful.

21:20 So, you can see our model localize

21:22 the harmful segment much more precisely than the baseline.

21:27 Yeah, finally, our everything source is open source now.

21:30 The supervised fine-tuning pipeline and Megatron

21:33 Bridge are a post training code, and the full toolkit,

21:36 data creation, real LM inference, evaluation, and visualizing web demo,

21:40 all is on the our GitHub repo.

21:43 So, if you're interested in here, please visit our web page.

21:47 So, thank you.

21:49 Oh, thank you.

21:50 Wow, you shared all codes.

21:52 I think it's really valuable for all communities.

21:56 Uh If you have any questions,

21:59 please uh write the questions in the comments, please.

22:03 Um I will start with my questions first.

22:07 Um Do you have any uh tried new thing um

22:12 what have you have not done in your company before?

22:17 Uh in during hackathon.

22:18 Yeah, so in our uh actually,

22:21 temporal grounding ability is hard to inject in MLM model.

22:26 So, we should divide the chunk level approach first

22:29 like the Transit or other architecture for dividing the scene,

22:33 but uh in this case, we used a single forward for chunking level approach

22:39 in the like a chunk level description and the chunk level godly.

22:43 So, uh we think like firstly,

22:46 we think like we exactly MLM can do this only in single forward pass,

22:52 but when we fine-tuning with the appropriate data,

22:56 we can achieve some temporal grounding accuracy and godly performance, too.

23:01 Yeah.

23:02 Thank you.

23:02 And you analyzed about the video.

23:05 So, is there any interesting video you filtered or something like that?

23:10 Uh yeah.

23:12 So, there are very many harmful video, and uh I will talk about like uh like

23:18 Megatron Bridge mode uh when we train on Megatron Bridge,

23:21 if it is not trained with Megatron Bridge,

23:23 then the training the model and also like

23:27 a loading data and all kind of pipeline

23:29 is uh like there are a lot of bottleneck on the loading model or loading data,

23:34 but when we train on the Megatron Bridge

23:37 with the MV codec or many MV Nvidia's stack,

23:40 then we can use really efficient pipeline for training

23:43 video and also on the inference pipeline, too.

23:45 Great.

23:47 So, you have only two days Yeah.

23:49 to have any strategy to train a model?

23:52 Because training a model during two days is big, right?

23:56 Yeah, yeah, yeah.

23:57 [laughter] It's very big, but like very simply,

24:00 like Nvidia give like a little more GPU for track B participants.

24:06 So, we should fine-tune the model.

24:08 So, yeah, we used the one node of a H100 for two days.

24:12 That's a really expensive and also very good our resource,

24:16 so we can fine-tune our model,

24:18 and we should find more uh appropriate hyperparameter, too.

24:23 But like yeah, we can do some kind of baseline approach here.

24:27 Yeah.

24:27 Yeah, thank you.

24:28 Uh if you have one more day or two more days, uh what you want to do?

24:34 Yeah, we should uh find the appropriate parameter

24:38 for achieve uh temporal grounding and accuracy at once,

24:42 because now our parameter it can

24:44 achieve only temporal grounding uh performance now.

24:47 So, yeah, we should find that more now.

24:50 Thank you.

24:51 He was the the winner of the track B of the Nemo Tron hackathon.

24:57 Thank you.

24:59 And we are going to invite the winner of the track C,

25:03 and especially the winner is the all of the hackathon as uh one thing.

25:10 So, please join us.

25:15 And please show the uh image of the winners.

25:19 Yeah.

25:29 Yeah, winner is here.

25:30 Uh the team is from the Nota AI.

25:34 Please introduce yourself.

25:35 Uh hello everyone.

25:37 I'm Hancho Park from Nota AI,

25:39 the world best AI model optimization company, hopefully.

25:43 Uh thank you for inviting me.

25:45 Uh I'm work as uh tech lead in our company.

25:52 So, okay.

25:55 Uh in this hackathon, we uh proposed a novel uh MOE quantization

26:01 method in the perspective of the data set.

26:05 Uh we call our pipeline as Pascal MOE,

26:08 which stands for pipeline for activation balanced

26:11 and sensitivity aware calibration set generation for MOE quantization.

26:16 Uh as you know, modern LLMs increasingly adopt

26:20 MOE architecture because of its uh computational efficiency.

26:24 However, the memory usage is remaining high,

26:27 making the lower precision quantization method is required.

26:32 However, existing quantization method do not

26:35 consider the following MS MOE specific characteristics.

26:40 First, some experts are lowly activated during calibration,

26:44 resulting in poor activation statistics,

26:47 which naturally lead to large uh quantization errors.

26:51 The second is uh quantization sensitivity

26:54 experts must also be sufficiently covered.

26:57 The figure below uh describe the problem.

27:04 So, we hypothesize that more balanced experts activate during calibration may

27:09 reduce the quantization error by improving

27:12 activation statistics for most experts.

27:15 So, naturally, our goal is to develop the pipeline

27:19 to synthesize samples that better cover underutilized experts.

27:25 We also hypothesize that more

27:27 frequent activation of quantization sensitivity experts

27:31 may reduce quantization error by improving

27:35 the accuracy of their activation statistics.

27:38 Uh our pipeline will be uh generate the samples

27:42 that frequently activate quantization sensitivity uh sensitive experts.

27:49 Uh this slide shows our proposed method.

27:52 The proposed method seems quite complicated,

27:55 but uh it's quite uh straightforward.

27:58 First, our pipeline investigated the target experts,

28:03 which uh either frequently activated or sensitive ones,

28:08 uh using the inference stage.

28:11 Uh And then, uh using the those the target target experts,

28:17 we can uh our system can identify the target tokens.

28:20 The target token will be a set of context and token pairs.

28:24 And then, we assign the label to the target

28:27 tokens in the initial uh the data set like this.

28:32 Uh With the uh in the next page,

28:35 with the predefined predefined prompts, with the sample annotated samples,

28:40 we can generate it the patterns,

28:43 which just describe what kinds of the samples the system has to generate.

28:49 Uh And then, we uh use the Nemo uh Nemo data designer

28:54 to synthesize the samples that we want uh with the generated Pascal guideline,

29:02 and then we replace it undesirable samples with the newly generated samples.

29:09 Uh this slide shows an example of the output from each case.

29:14 Uh interestingly, we observed that intuitively less common

29:16 uh less natural tokens tended to activate underutilized experts,

29:22 such as backslash respect or double backslash.

29:25 And uh if as you can see this picture uh below,

29:31 the same shading indicate the rules

29:33 and the synthesized expression that follow the rule.

29:37 Uh we observed that the generated guideline was very interesting for us.

29:43 Uh in in In to uh evaluate uh the our proposed pipeline,

29:50 we contest the QN3 MOE model using

29:55 the various kinds of the NVIDIA software stacks.

29:59 The quantization scheme was into four.

30:03 We used the GPTQ algorithm.

30:08 So, first, as you can see in this slide,

30:12 our proposed pipeline generates the more target target samples

30:19 in terms of the expert balancing or sensitivity experts.

30:26 This slide shows the very impressive research.

30:30 Our proposed pipeline outperforms the the NeMo

30:37 training data that has contain undesirable samples.

30:45 Although I wanted to show more things, but this is last page.

30:51 Thank you for your attention.

30:52 Yeah, thank you for your presentation

30:54 and congratulations again for the first prize.

30:58 And actually, this is what we expect during the hackathon exactly

31:03 because we can generate something easily in this era with AI's,

31:09 but for the data set, it must be useful, right?

31:14 But you made a data set and it is applied to your product as well.

31:19 So, it was really amazing.

31:22 I want to ask about the experience of NeMo data designer.

31:26 Yeah.

31:29 I think that NeMo data designer is a very

31:33 impressive for us because everything has been kept.

31:36 We just gave the prompt.

31:38 We just gave the guideline,

31:40 and then the NeMo NeMo designer generated so many things

31:43 so many samples which fit our which meet our expectation.

31:49 So, we will use this design this this software later,

31:55 and we hopefully we wanted to give our message to the I mean I

32:01 mean contribute to our proposed pipeline

32:05 to the to this designer design designer later.

32:09 Yeah.

32:10 Great.

32:11 Yeah.

32:11 And there is a good question.

32:14 Sky World Horizon asked about how do you define quantization sensitive experts?

32:21 Yeah, it's very important, but I missed it because of the time constraint.

32:25 Actually, the Yeah, quantization sensitive during the hackathon actually

32:30 there are a lot of definition of the sensitive experts,

32:33 but we simply define the sensitive expert

32:38 as the the largest uh magnitude of the weight.

32:43 I mean the we sorted the maximum magnitude of each weight,

32:49 and then we define the oh,

32:52 this is this this expert very sensitive because the weight magnitude is so high.

32:58 Yeah.

32:58 This will be Yeah.

33:01 I mean we we in the future work, we will propose a method to select

33:07 automatically select the the sensitive or expert.

33:11 Yeah, really great to hear the details.

33:14 During the hackathon,

33:15 there was a great experience one more thing because was demo session because

33:21 we we shared our knowledge and experience during the hackathon in the demo.

33:26 So, it was really great.

33:28 And one more good questions, username Hobby 75 and 24.

33:34 Do you think this method will work for other quantization algorithms as well?

33:40 Yeah, of course because as I mentioned before,

33:43 the our method is not just for the algorithm.

33:46 It's our our method is related to the data set itself.

33:51 So, our method can be used any kind of the quantization method.

33:57 Yeah.

33:58 Thank you.

33:58 Thank you for your attention.

34:01 Please visit GitHub and you can experience the code as well.

34:05 So, please turn the our sketch video of day two, Zach, please.

34:16 We had a great day.

34:51 Yay!

34:55 [laughter] Absolutely incredible presentations.

34:57 I mean that that again, that's uh Imagine that all of the projects that we

35:03 watched were close to that caliber of of presentation.

35:07 Obviously, these guys won for a reason,

35:09 but there there wasn't a big split between the top

35:15 the top projects in the in the next best projects, right?

35:18 So, we watched uh 20 or so presentations.

35:23 Yeah, right.

35:24 [laughter] All of them so good.

35:28 Yeah.

35:28 Yeah.

35:29 We had a we had a bunch of you you know

35:32 uh time to hear about the the projects in the the hackathon.

35:37 May- maybe we'll spend some time though

35:39 answering some questions from the the chat.

35:42 I saw a few earlier.

35:44 I also saw some legends from our community like Cody's Guide to AI.

35:49 Thanks for coming to the special live stream.

35:54 Guys, we have uh also some comments talking about how good the food is in Korea.

36:00 I can confirm that the food is extremely good.

36:04 I have had the pleasure of eating it all week,

36:06 and I'm a little bit sad to have to go back to Toronto,

36:11 but I'm going to make sure to find a good Korean spot there.

36:13 So, maybe let's let's scroll up a little bit here in the chat,

36:18 and let's see I know we had a question early on.

36:23 Maybe we can uh talk to that one.

36:29 Yeah, okay.

36:29 So, at Hamad Raza 584 asked, "How can you use AI agents created

36:36 with NeMo agent toolkit with actual physical AI agents?

36:40 Is there any integration with ROS 2 operating system?" So,

36:44 this is a interesting question, Chuanlu.

36:46 I don't know your your familiarity,

36:49 but do you know of any way that we can actually

36:52 hook these LLM brains up to maybe some arms and legs?

36:56 Maybe we can add some skills, and actually agents will build new skills as well.

37:05 So, it is possible, I think, and there are a lot of documents there.

37:09 So, you can easily add some skills with AI.

37:13 Yeah, and I think, you know,

37:15 we obviously we do have a lot of this is a NeMo Tron Labs live stream,

37:19 so we're we're we're kind of focused on the NeMo Tron series of models,

37:23 but NVIDIA has a a huge array of physical AI models.

37:28 Yeah.

37:28 We do have a physical AI stream as well that you can check

37:33 out if you wanted to look at that where they discuss more about robotics,

37:37 but things like Groot, Isaac Labs.

37:40 Yeah, Cosmos.

37:41 Cosmos, of course.

37:43 You know, if you want to get started

37:44 with physical AI or transitioning from LLM's to physical AI,

37:48 lots of great technology for you there.

37:51 And of course, it's only going to get better.

37:53 This is the worst that robots will ever be, which is which is crazy.

37:57 Yeah.

37:58 Uh we have another few questions.

38:02 NeMo Squid from Angel Romero.

38:04 [laughter] Someone asks Ad Board Tomato from YouTube asked,

38:09 "What do you think about the NeMo Tron challenge?

38:11 How do I get past 0.84 without training all the linear layers?" So,

38:15 this is actually in reference to our NeMo

38:17 Tron reasoning challenge that we're running on Kaggle.

38:20 Oh, really?

38:21 Yes.

38:21 Uh I mean, I would there was a midpoint

38:26 that just happened where the team actually you know,

38:32 chose the best project so far,

38:34 and I think that that that write-up obviously has some great ideas.

38:38 If you wanted to look at that to get a little bit of inspiration.

38:41 So, what what what are your approaches when you're trying to squeeze

38:46 the the last few drops of juice out of training a model, Chuanlu?

38:51 For me, just check the data set.

38:57 Just make a more validate and good quality of data set.

39:02 So, we I should check about the details from there.

39:05 I analyze it and filtering some bad cases, or I should find the failure cases,

39:12 and add some data set for that as well.

39:15 So, I think the problem is just data set.

39:18 I mean, a classic and very good answer.

39:22 Yeah.

39:22 Look at your data.

39:23 Make it better.

39:25 If you don't think you can make your data better,

39:29 you haven't looked at your data enough, right?

39:31 So, I think it's also like similar to K food.

39:34 Like if there is a good ingredient and good mind, yeah, it it comes out.

39:40 There you go.

39:41 That's [laughter] the secret.

39:42 Good ingredients and good minds make good models and good [laughter] food.

39:48 That's good.

39:49 That's good wisdom for for Friday.

39:53 Cody's got AI asked, how do we open this box up and add the NVIDIA agent

39:57 toolkit so it refines itself the whole time and stops failing tools.

40:00 What is the best local model to use for Nemo claw and how to make agents?

40:04 Cody's got many great questions there.

40:07 Community legend here.

40:09 I mean, maybe let's just think about the first one.

40:12 So, you know, we saw in the first project from Naver this idea

40:17 that we could build an agentic system to kind of collect feedback.

40:21 I think Cody's got AI is asking,

40:23 how do we take that feedback and then roll it back

40:26 into the the technology to to make it better over over time.

40:30 What are your thoughts on that, Chuan?

40:35 [sighs] Keep trying actually,

40:35 but based on the model there is ability of the two callings metric as well.

40:44 So, choose good great model first and try keep make a great loop for that.

40:52 If there is a failure,

40:53 try paint it place made it try again for the great loop, I think.

41:01 Yeah, I mean yeah, I think you hit the nail on the head.

41:05 I mean, at the end of the day, right, we we are enter we're seeing more agents

41:10 do this kind of recursive self-improvement thing, right?

41:13 Where like Hermes agent or something like that is going to you know,

41:16 edit its own skills, its own config, its own even sometimes, you know,

41:22 you can you can work with it to enable

41:25 it to edit the way it's implemented, right?

41:27 So, it improves over time.

41:28 I think though for the first project, the the the Naver you know,

41:34 kind of collect this feedback project,

41:36 it's really important for companies, right?

41:37 To know what people think about

41:40 their products and where they're finding problems

41:42 so that those companies can take

41:43 those insights and make the product better, right?

41:46 I remember on day one talking to the to the Naver team and say,

41:50 "Hey guys, you know, like this is very useful even for us.

41:55 What the community thinks about [laughter] our technology, you know,

41:58 and if we saw there that someone's struggling with the the reasoning feature,

42:03 you know, we should do?

42:04 We should make it easier, more clear to use the reasoning feature of the model.

42:07 There you go." We're just scrolling through the chat here, guys.

42:12 So many different Oh, yes.

42:16 AI career path RLM or bust.

42:20 RLM, of course, recursive language model.

42:24 Some work done by Alex in the MIT SAIL team.

42:28 Really great research that a lot of people are very excited about.

42:32 If you haven't heard of it, definitely look into it.

42:35 Yes.

42:37 Simon Falk, interesting quantization shows

42:39 where efficiency meets fragility like biology.

42:42 It's not just energy, it's how it's structured.

42:44 That's where intelligence holds.

42:45 What a comment.

42:47 Absolutely true though.

42:49 Quantization is a is a big thing.

42:56 [laughter] You know, as as as the as the team described, right?

43:02 But not not an easy task, especially when it comes to MOE.

43:06 I see you have another question.

43:08 Chuan, maybe you'd like to read this one.

43:10 This one?

43:11 Yes.

43:13 This this one.

43:16 Oh, have you tested Nematron Cascade 2 or Nematron 3 Nano with Nemo claw?

43:24 Yes.

43:25 Yeah, so Nemo claw is the NVIDIA runtime for Open claw.

43:31 We've done a few live streams about it recently,

43:34 which is why we're getting some questions about it.

43:36 I would say yesterday and the day before actually we

43:40 were showcasing the model using both of course Nematron 3 Super.

43:47 I'm very biased, but that's my favorite model.

43:50 And and the Gemma 4, the new Gemma model.

43:55 Both running in through Open claw and then

43:59 as well you can run them through Nemo claw.

44:01 So, I would say those are great models to start with.

44:03 In terms of a stacking the models, right,

44:07 we we're we're actually going to have some live stream

44:11 coming up in a few weeks where we're going to get

44:13 a little bit into more the weeds on how you

44:16 might stack models or sub-agents with the kind of claw technology.

44:23 So, Cody's got AI, make sure to tune in the future,

44:26 but yeah, Nematron Cascade 2 and 3 Nano are both great models.

44:31 I would tend to say that something like

44:34 Super Caliber is better for the main agent of your claws and then we're going

44:40 to use things like Nano and Cascade as sub-agents.

44:44 Are you big claw guy, Chuan?

44:46 Are you you know, are you are you out there clawing it up?

44:51 [laughter] Yeah, you said.

44:52 I I think we are adopt with elements agents nowadays.

44:57 So, we are really experiencing the self-healing agents as well.

45:03 But we are having struggle with tokens, right?

45:09 Yeah, because that's going to burning out.

45:12 So, we need our own own models sometimes.

45:15 So, we need Spark or some our own cloud stuff.

45:21 So, you mentioned about Gemma 4 or our models as well.

45:25 But I want to mention one more stuff for the quantization because there

45:29 is a model quantized with MBF 16 for the Gemma model as well.

45:35 So, you should try with that.

45:37 Yeah, famously DGX Spark, right, has has works great with NVFP4.

45:46 So, any model that's available in NVFP4 format is

45:50 going to run a little bit better on the Spark.

45:53 I do think at the end of the day, you know,

45:57 there's a lot of times that your agent's going to do work

46:00 that isn't you don't need to use a lot of expensive tokens.

46:06 Grepping through the file, you know, to to to find some snippet or you know,

46:12 creating a directory or you know, making a small edit to a to a code file.

46:17 We don't need we don't need Claude Opus 4.7 for that, right?

46:21 We can get away with [laughter] we can get away with Gemma for that.

46:25 Yeah.

46:27 Excellent stuff, guys.

46:29 We have NVFP4 is dope.

46:32 Yes, true.

46:35 What's the best open source to use with Nemo claw?

46:37 What tool fits best with Quantum?

46:40 Well, I've never experienced about that actually.

46:43 [laughter] Well, what I can tell you is that Nemo claw

46:47 is designed to work well with all kinds of different things.

46:50 So, obviously it should work with NVIDIA technology,

46:53 but should work with most everything else as well.

46:56 The idea is to you know,

46:58 to have a platform that works well with whatever you're using.

47:01 And if you don't know what you want to use

47:04 or you have an idea that we haven't put in yet,

47:07 please contribute to the repository, contribute to the project.

47:11 And if you don't know how to contribute to the project,

47:13 you should check out Tuesday's live stream where we

47:16 talked all about how you can contribute to open source.

47:19 You know, if you if you're like, I have a great idea.

47:21 I wish it worked with this or that, please contribute it.

47:24 And if you have a Quantum model that isn't

47:27 yet supported by Nemo claw, please submit a PR.

47:33 [laughter] It would be that would be great.

47:35 I I think we have some some exciting Quantum stuff coming up,

47:40 but for right now that support, if you need it, put it in.

47:49 Yes.

47:49 We have some time for waiting for the great great comments.

47:53 We I want to mention one more thing during the hackathon.

47:57 We had a great session with a meditation by Nematron.

48:02 Zach, can you please show the image from the for the meditation session?

48:10 So, [laughter] we had a meditation with Nematron.

48:14 That's right.

48:15 Yeah, Nematron Super generated a text for the meditation,

48:20 especially for the developers.

48:22 So, the text was like please forget the error messages or the stacked PR.

48:32 Just relaxing.

48:34 We had a second day of the morning that we had a meditation.

48:39 So, it was really great,

48:41 relaxing and we can focus on the hackathon after the meditation as well.

48:47 That's right.

48:48 I mean, at the end of the day, Nemotation.

48:52 [laughter] Yeah.

48:53 That's from IT Sucks.

48:56 I mean, listen, hackathons are long.

49:00 You have to stretch your mind a lot.

49:02 It's nice to take a little break, little breather, refocus,

49:05 recenter as guided by Nematron 3 Super who's a pretty chill guy.

49:12 There you go.

49:14 Also, I think you you had a voice to or text to speech.

49:18 So, you can just close your eyes.

49:20 And if you bring the image back up, Zach,

49:22 little Easter egg for people who know the developer community team,

49:26 you'll see in the back left corner there, that's Mark Heat.

49:30 [laughter]

49:32 An absolute legend in the developer community effort here at uh NVIDIA.

49:38 So, uh Nemo Zen is great.

49:41 Raphael, uh absolutely great Nemo Zen.

49:44 He was really busy with uh Build a Claw session, right?

49:48 yeah.

49:48 Build a Claw was popping off, guys.

49:50 I Uh if you're going to uh an NVIDIA event

49:54 in the future and you're wondering all about claws, Mhm.

49:57 you can probably find a Build a Claw tent near you uh that will

50:02 help you understand this technology a little

50:04 bit better and get hands-on with it, see some sweet demos.

50:07 Mhm.

50:08 One of the things that I was really excited about, Uh-huh.

50:10 Cha- Chan- Lon, is, you know,

50:12 here in in Korea you've built a really impressive community of AI practitioners

50:18 Yeah.

50:19 called Pseudo Labs.

50:20 And they actually showed up and they helped out

50:22 at the Build a Claw event showing demos, helping people.

50:26 So, it wasn't just like NVIDIA came to town and all from HQ like, "Oh,

50:29 blah blah blah." It was actually local devs from your community

50:33 that were helping people use un- understand this technology.

50:37 If someone wants to learn more about Pseudo Labs,

50:39 so they're excited by what we're doing and they're

50:42 excited by maybe digging deeper on nights and weekends,

50:45 what where should they go?

50:46 How do they find Pseudo Labs?

50:48 Uh actually, we are based on Discord and we

50:52 are non-profit AI and data science research uh community.

50:56 So, you can just come on in and you can

50:59 visit our uh there are several rooms and our timetables.

51:04 So, you can just come in and you can

51:06 have experience about uh agents or researching AI models.

51:10 So, physical AI as well.

51:13 Just there are a lot of projects there.

51:16 And actually, we invite them into our Build a Claw session,

51:21 but they're all uh employees from other companies like big enterprise,

51:27 but they had a a PTO for uh uh for visiting our session and not just visiting

51:36 uh helping our uh session and they they

51:40 shared about their own agent Claw experience for us.

51:45 It was amazing, I think.

51:47 Yeah, it really was amazing.

51:49 I mean, these are working professionals that took time off of work

51:52 to come help out uh to help people understand and access this technology better.

51:58 Uh that's the vibe.

51:58 That's been the vibe the whole time

52:00 that we've had the pleasure of being in Korea.

52:02 Uh it is just absolutely amazing.

52:04 I can't wait to see what you, Chan- Lon,

52:06 and the team that you don't see behind the camera here, guys,

52:09 uh get up to uh to to keep growing that that community and keep you know,

52:15 helping the the the kind of AI you know,

52:29 work being done in These live streams are fun, Mhm.

52:33 but there's a whole ecosystem of things you can learn from.

52:36 So, uh if you're hearing in say in the in the track A project, right?

52:42 You're hearing about agents, but you're not quite sure what they are,

52:45 what they mean, uh then uh you can check out these learning

52:49 paths at developer.nvidia.com/topic/ai/how-to-build-an-AI-agent.

52:54 We've got some updates coming that should uh make

52:57 this more uh related to claws in the near future.

53:00 Zach, if you go to the next slide, please.

53:03 Uh we also have, of course, our NVIDIA Nemo NeMo Tron repository.

53:08 This is where you can find how to run the model,

53:10 how to build a model from scratch yourself if you wanted to.

53:14 Uh you might need more than a 8100 node and 36 hours though,

53:19 but you you could build your own NeMo Tron 3 Super if you wanted to.

53:23 Uh and then, of course, we've got another slide.

53:26 Uh if you want to join us at discord, we've got discord.gg/nvidiadeveloper.

53:30 Once you join that, you can look for the hashtag NeMo Tron models channel.

53:34 Uh you'll be able to find me in there uh 24/7.

53:38 Uh even if you send me a message in uh in Korea time,

53:41 I'll probably respond pretty quickly uh because I'm I'm up all night.

53:45 So, uh but leave any questions you have in there for us.

53:49 Uh if you have any questions for the teams,

53:51 you can also put those in that channel uh and I'll

53:54 make sure that they get forwarded to uh Chan- Lon,

53:58 who will make sure that it gets forwarded

53:59 to our uh to our our incredible winners.

54:01 Yeah.

54:02 Next slide, please, Zach.

54:04 Uh if you have any questions for the NVIDIA developer community team,

54:08 so this is people like Mark Heeps and Zach,

54:10 please reach out to community@nvidia.com.

54:13 Uh Zach does live in a cave.

54:15 It's always dark and the only screen he has is this email,

54:19 so he'll answer your question if you send him an email.

54:21 I'm just kidding.

54:22 Uh we let him out from time to time.

54:24 Uh [laughter] yeah.

54:26 He's not an intern either.

54:27 He's a real full-time employee, guys.

54:29 Uh otherwise, I've one last question for you,

54:33 Chan- Lon, well we've got uh about 5 minutes left.

54:35 Yeah.

54:36 Uh I've I've enjoyed my time so much here.

54:38 Uh I've enjoyed getting to meet all of these incredible uh developers.

54:43 Uh you know, uh every time I get to visit somewhere new,

54:46 I'm so excited to see how passionate uh people are about AI,

54:49 but how did you get into AI?

54:51 What what made you decide what you want to do

54:54 with your life is AI and not just learning about it,

54:59 but then teaching people and sharing it with a community of people.

55:02 What What uh gets Chan- Lon out of bed for teaching AI every morning?

55:12 [laughter] Actually, I started with uh experience of the Kaggle.

55:16 Actually, I I grew from that community.

55:20 So, every day I see the leaderboard from there and then

55:25 I see the improvement me and the community as well.

55:29 So, nowadays there are a lot of AI stuff in all communities,

55:34 all interesting stuff.

55:36 And so, when I woke up, I just see what's going on today like,

55:43 you know, uh today also there are big news about uh AI.

55:47 So, GPT-4.5 is released.

55:51 As like that, if something new AI is coming out,

55:55 we experience we want to experience all things, right?

55:59 Yeah, I started with that and new experience gives me new life, I think.

56:05 Yeah, that's coming from what I do every day, every morning.

56:09 And I start from that experience from uh since Alpha Alpha Alpha Go.

56:14 So, I every day every morning I experience new AI stuff.

56:19 So, with agent, it helps me a lot to experience more

56:25 and more and a lot of experience I knowledge it as well.

56:28 So, it's incredible life and new new life and with NVIDIA

56:34 community and NVIDIA models and supported by our GPUs.

56:40 So, it was amazing.

56:42 Incredible stuff.

56:44 Incredible answer.

56:45 Just what you'd expect from someone who's uh who's who's helping

56:49 so many developers uh across the the country and the world.

56:53 Uh one last question in the chat,

56:55 when is NeMo Tron Ultra releasing approximately?

56:57 Thanks so much.

56:58 It'll come out when it comes out, guys.

57:01 I know every every stream you want to know.

57:04 It will release when it releases and uh I'm sure that we'll tell

57:08 you uh you should expect to see uh that we we announce it.

57:13 So, uh it won't sneak up on you.

57:15 Uh you'll know.

57:16 We'll we'll we'll tell you.

57:17 Uh thank you so much for your time today, Chan- Lon.

57:20 Thank you for sharing uh the office with us.

57:23 Thank you for coming to Korea here.

57:25 Of course.

57:26 Absolutely amazing.

57:28 Kamsahamnida.

57:29 Kamsahamnida.

57:29 Uh Kaja.

57:32 Kaja.

57:33 Kaja.

57:33 Okay, thank you, everybody.

57:35 We will see you next time.

57:36 Bye.

57:36 Bye.

57:36 Bye.

57:37 Bye, everybody.

Study with Looplines Download Captions Watch on YouTube