Nemotron Dev Days Seoul: Sovereign AI, Hackathon Highlights & Build-a-Claw | Nemotron Labs
NVIDIA Developer
1:03 Hello and welcome everybody to a very special NeMo Tron Labs
1:07 live stream where we are joining you live from Seoul, South Korea.
1:13 We are very excited for this for this and we're joined by very special guest,
1:19 our host today, Chung Won.
1:21 Chung Won, take it away.
1:22 Hi, I'm Chung Won.
1:24 Thank you for coming here.
1:25 Thank you for visiting Korea and we are
1:28 having a first live stream in Seoul, right?
1:31 First of many, I hope.
1:32 Yeah.
1:33 Uh so, over the last few days,
1:36 we have been here to do something called the NeMo Tron Developer Day,
1:42 which has been very exciting so far.
1:45 Deep dive.
1:46 Go ahead.
1:47 Yeah, deep dive sessions with from the HQ
1:51 NVIDIA researchers including Brian BPA as well.
1:55 That's right.
1:56 Brian Catanzaro made an appearance.
1:57 It was pretty epic.
1:59 We're we're going to get more to that in just a second, guys,
2:01 but in order to set up kind of the vibe
2:03 of what it was like here over the last few days,
2:05 we're going to kick off to a video that showcased
2:07 what the first day of of that developer day looked like.
2:11 So, Zach, if you want to go ahead and play that video.
2:17 [music] [music] [music] Cool.
2:48 Very cool.
2:49 Very cool.
2:51 It was it was a lot of fun to be there.
2:53 A lot of very interesting projects that we saw over the the few days.
2:59 Chung Won, tell us a little bit about, you know,
3:02 how you set this up and and kind of like
3:04 what the idea was to bring as much you know,
3:09 energy and and kind of the the vibe of NeMo
3:13 Tron to the big the good people of Seoul.
3:16 Yeah, actually Korea is one of the most
3:21 important country about sovereign AI because we are building
3:25 models from the scratch and there are a lot
3:28 of model business about fine-tuning or post-training as well.
3:32 So, and especially we have many industries including manufacturing.
3:37 So, it is there there are a lot of area we can apply our NeMo Tron ecosystem.
3:45 So, we prepared the deep dive sessions and we had
3:49 a hackathon and we had a build a claw session.
3:53 So, it was amazing full of NeMo Tron everything like
3:57 model and SDKs and training strategies and even NeMo claw, right?
4:02 That's right.
4:03 Yeah, I mean it was it was a lot of fun just
4:06 to see everybody building with NeMo Tron or the NVIDIA ecosystem of course,
4:12 but also so exciting to see how passionate and skilled,
4:18 you know, the Korean devs are when it comes to leveraging technology.
4:22 We had a lot of people join for the build a claw event, you know,
4:26 where they were getting kind of started
4:27 with claws on their own hardware or using, you know, cloud instances.
4:31 You know, at the end of the day,
4:32 I think this was a incredible experience for people joining.
4:37 Guys, don't forget if you have any questions for any
4:40 of the people you're going to see here today,
4:42 don't forget to drop them in the chat.
4:44 We're going to be having some presentations and so when we have
4:48 those, please feel free to ask the presenters any questions that you have.
4:52 Chung Won, one of the things I noticed during the deep dive sessions, right?
4:57 Is that a lot of people were not
4:59 just like listening and nodding along, but taking notes.
5:02 Really really they had really great questions afterwards.
5:05 You know, how [clears throat] do you think people are adopting AI in Korea?
5:13 How do you think that impacts how, you know,
5:16 us as NVIDIA need to communicate to to the Korean developers?
5:22 Actually, some of them are building their own model
5:26 or some of them are using AI for building agents,
5:31 but this time we had a really deep dive session.
5:35 So, during that, I think when we are just building our AI model or agents,
5:44 we are sometimes we are competing with other companies like we are developers.
5:50 We know opening our model authors or strategies is important are important,
5:57 but in company it's really difficult to open our knowledge.
6:02 So, but in this event, we shared all of our NVIDIA strategy for training models.
6:12 So, it was really really impressed by Korean developers.
6:17 I heard a lot of recommendation or reviews
6:23 from them and I really excited about that experience.
6:27 Yeah, it was it was really cool and I
6:29 mean what was also cool obviously was the hackathon.
6:33 Seeing so many people building with tools over
6:35 48 hours and the projects were extraordinarily ambitious.
6:40 And not just ambitious, but like they kind of like got her done.
6:44 We're actually going to see some of these teams who won.
6:49 So, can you outline for us a little
6:51 bit like what the what the hackathon structure was?
6:55 You know, there's this idea of three tracks.
6:57 Do you mind like outlining what that was all about?
7:00 Actually, we had a two days for the hackathon.
7:04 It was really really short time to do,
7:07 but we planned three stacks including create created creating a builder agent
7:15 creative use case and a fine-tuning models and create a data set.
7:21 So, synthesize the data set.
7:23 So, it was really short time to do that.
7:26 But everyone made really great project and they will bring
7:31 to their company or their individual project after the hackathon as well.
7:37 Yeah, it was really incredible to see the amount
7:39 of work that all of the teams were able to do.
7:41 We're going to see the winners today and we're going
7:43 to get a chance to see their presentation what they actually created,
7:47 but I do want to say that all of the projects were strong.
7:50 It was it was a hard decision to determine
7:52 the winners because there was so much talent over that 48
7:56 hours and I was just absolutely astounded by the level
8:00 of ambition and and kind of greatness of the projects,
8:03 but we've got the three best projects that we saw at the hackathon here
8:07 one from each track and so we're actually going to get our first guest.
8:11 Do you mind introducing them?
8:13 Yeah, he's from Naver and they made a creative agent.
8:19 Please join us.
8:20 Yes.
8:22 Zach, if we can get that picture up showing them holding the spark wooden.
8:30 Hi.
8:32 I'm Han Su and I'm from Naver.
8:51 Is this slide on, Zach?
8:54 You're good.
8:55 Okay.
8:57 Please introduce yourself.
8:59 Okay, my name is Han Su and just you can call me just Han Su.
9:03 I've been working at Naver as a AI engineer and I'm
9:06 very glad to share my experience in about my hackathon.
9:10 So, I'm glad to see you.
9:13 Yeah, we can just start it with your stuff.
9:17 Please introduce your projects.
9:19 Okay.
9:21 My project is the NeMo briefing.
9:23 The NeMo means the new model.
9:25 So, NeMo briefing is an AI agent to collect
9:29 the users' reactions and summarize and deliver the report about that.
9:36 So, I started.
9:40 Okay, the first I will introduce my team and my company and I
9:47 I joined the this hackathon with my awesome teammates Song Jin and In Gon.
9:51 They are very awesome AI developers and so,
9:56 our let's introduce our company, the neighbor.
9:59 The is the biggest IT company in South
10:02 Korea and the competitor to Google in search engine.
10:06 We have the rather many business models such as place or B2B or AI or cloud.
10:13 So, let's start it.
10:15 Uh Our key idea is started from our internal work project.
10:20 We built some leaderboard to evaluate the various LLMs.
10:25 They are accuracy and performance.
10:28 But, we found that there is a significant gap
10:32 between the benchmark score and the real users' reactions.
10:36 So, we thought it was really good to build an AI agent
10:41 to collect the real users' feedback
10:44 and reactions and add them to our leaderboard.
10:48 So, we started it.
10:51 Yeah, it's the architecture of Nemo briefing.
10:55 First, we build a query generator agent for the refining
11:00 the query and for the high-quality search result.
11:05 So, the second part is the collector agent and it
11:09 is our key point and inside the each collector,
11:12 we had a validator and validator filter out the unnecessary items.
11:18 And we collect various from the various source such as Reddit,
11:23 Rob Stars, or Hugging Face community.
11:25 And the final uh part is the reporter agent.
11:28 The reporter agent does aggregate the whole
11:31 data and summarize them as a markdown report.
11:35 This whole part is orchestrated by Nemo Transformer model with React prompt.
11:41 So, we used every NVIDIA steps in our project.
11:46 First, we used Nemo Transformer model.
11:48 We used super model for difficult tasks such as React or validation.
11:53 And we used nano model for simple tasks such as query generation.
11:57 And these models are deployed on Brave Cloud and with the Nemo framework.
12:03 It is very useful and easy to use.
12:05 And we developed our agent with Nemo agent toolkit and known as NAT.
12:11 I think it is really awesome framework.
12:13 I I've used other agent frameworks such as LangChain or ADK,
12:16 but it is really cool and easy to use and extendable.
12:20 I really like it this configuration.
12:24 So, let's see the demo demo video.
12:29 I searched the Nemo Transformer model
12:33 and query generator starts and I skip them.
12:37 The data collector works but in the parallel and they
12:41 collect the each data and validate them in parallel.
12:44 So, I skip it.
12:47 The final section is reporter.
12:49 It takes some long time, so I skip that.
12:52 This is a final report.
12:55 So, let's see type in the example report about the Nemo Transformer model.
13:02 The report say that some users have
13:04 some trouble in the configuring the reasoning mode,
13:08 but I think it is not the model problem.
13:10 It's just about the issue with the framework or template,
13:14 but it is very useful to catch catch catch this gap with our real time.
13:22 I think Nemo briefing can help this.
13:27 Uh Okay, so final session and we don't have much time.
13:32 It is a hackathon and 36 hour isn't enough, so we Here is the future work.
13:39 I will skip that.
13:40 It is not important.
13:42 Okay.
13:43 I'm done.
13:44 Great.
13:45 Great project.
13:47 This was the winner of the track A.
13:52 Do you have any questions?
13:56 Please shout out the text.
13:59 Hello.
14:01 Congrats Hansu and team.
14:03 Great demo.
14:10 Yeah, um during the hackathon you met uh
14:14 NVIDIA researchers and you had a good great mentoring.
14:20 Any great experience from that?
14:22 Okay, they they helped me and they helped our team to select
14:28 the topic proper topic and make right architecture at the proper time.
14:36 Their advice is really helpful for me.
14:38 Yeah, thank you.
14:40 And actually Nemo is a neural models but neural
14:44 models but you saw you saw the new model, right?
14:47 Right.
14:48 Yeah.
14:49 Why did you think like that?
14:51 Because our project is to collect the real
14:55 user reaction from the new model our new models.
14:59 We used the Nemo.
15:00 Are you going to use it for the automatically like agents?
15:04 Mhm.
15:04 Right?
15:05 Right.
15:06 For the your company research or some?
15:09 Yes, I developed some agent in our company such as AI nation or agent.
15:16 You will open to public or you will use it
15:18 on your company I it released at March in this year.
15:23 You can use it neighbor.
15:25 And everyone can use it.
15:27 It's good for the government.
15:28 Ah, I see.
15:30 It's great for researchers, I think.
15:32 What all model builders or model users.
15:36 Thank you for the great project.
15:38 Thank you.
15:39 Thank you, Hansu.
15:41 So, we are going to move to to next uh track.
15:48 Track B.
15:50 Uh the winner was from the Pyla.
15:53 Will can you join us?
15:55 Jack, please show the winner's image.
15:59 Thank you.
16:15 Oh, I didn't mention the important thing.
16:21 Actually, the prize was the Spark.
16:23 So, the [laughter] You can see the Spark on the image, right?
16:29 Yeah.
16:30 Thank you.
16:31 Thank you for coming in.
16:32 Please introduce yourself.
16:33 Uh yeah, hello.
16:34 I'm from Pyla AI and my name is Junsu Kim.
16:38 We uh did we winning the track
16:41 B about the fine-tuning open-source Nemo Transformer model.
16:45 Thank you.
16:45 Please introduce your project as well.
16:48 Yeah.
16:49 Okay.
16:50 Let me start.
16:50 Yeah, hello everyone.
16:52 I'm from Pyla AI and today I'll talk about the fine-tuning NVIDIA's
16:57 Nemo Transformer nano into a fine-grained
16:58 trustworthy video understanding model for content safety.
17:03 So, when you ask a model to classify a video for safety,
17:07 you will really have two choices.
17:09 First, you can feed the whole video it at once and get a single label.
17:14 But, in this case you lose any
17:16 sense of when the harmful content actually appears.
17:19 Or, you can split the video into chunks and analyze each one,
17:23 which give you fine-grained temporally grounded results.
17:26 We want to see how much this second approach changed
17:30 model behavior and how far we can push it with fine-tuning.
17:35 For data, we build our training set on top of Safe Watch,
17:39 which published in ICLR 2025.
17:41 That gave us about 8,000 videos and 25K chunks for training
17:46 plus a test set of 1,200 videos with video-level descriptions and labels.
17:53 So, our first step is supervised fine-tuning.
17:55 We run both LoRA and full fine-tuning on Nemo Transformer
17:58 nano 12B V2 VL model using the Megatron-Bridge SFT pipeline.
18:04 The goal is to teach the model to produce
18:07 chunk-level explanations instead of a single video-level summary.
18:14 For post-training, we wanted to do preference optimiz- preference optimizations.
18:19 So, but Nemo doesn't officially support the Nemo Transformer nano V2 VL yet,
18:23 so we implement simple and DPO ourselves directly on the top of Megatron-Bridge.
18:29 Uh due to time constraint,
18:31 the results I'm sharing today are from simple training method.
18:34 The reasoning of this preference optimization is open-source models
18:38 are ready to a reasonable job on whole video understanding,
18:42 but they struggle with temporal grounding.
18:44 So, we use segment-level temporally grounded output as the winning
18:47 data and whole video only description as the losing data.
18:53 Yeah, so for evaluation, we put together three benchmarks.
18:56 First, we split Safe Watch training set into train
19:00 and evaluation data set using TransNet for temporal grounding.
19:04 The second, because the Safe Watch test set only has video-level description,
19:08 we use NVIDIA's AI blueprint VSS
19:10 to pseudo-labeling the chunk-level explanations and classifications.
19:15 And finally, to stress test longer videos,
19:17 we take video MME clips and inject harmful segment in the middle,
19:21 so the model have to quickly temporally grounded the injected portion.
19:27 Uh we evaluated on two metrics.
19:29 First is average F1 for accuracy and second
19:31 TIOU TIOU for the temporal grounding ability.
19:36 So, our baseline pipeline is straightforward.
19:40 Nemo Transformer by a VLM takes the whole
19:42 video in one shot and outputs three things:
19:45 the description and category label and six categories and explanation.
19:49 So, the baseline results on the open source model
19:53 is like 0.25 0.5 real is like our approach 0.318.
19:59 The finally, the nano 12B V2 real model
20:02 showed the best performance on Safe for us benchmark.
20:06 So, our approach at chunk level inference on top here.
20:10 We train on Megatron Bridge expert to hugging face and service real LM.
20:15 For each temporal chunk,
20:16 the model emits a description, godly labels, and explanations.
20:23 Yeah, so on the Safe for us evaluation split,
20:25 the baseline Nemo Transit as 0.5 of about 0.5 F1,
20:30 the chunk level prompting alone jumped at that to 0.75.
20:34 The Laura's SFT more pushed up performance of this model, and finally,
20:40 the fine-tuning with the full fine-tuning with chunk
20:43 level approach the best performance on temporal grounding.
20:47 However, about the accuracy, we it degrade the performance,
20:51 but we should find more accurate hyperparameter.
20:56 And for temporal grounding results,
20:58 uh when uh for full top fine-tuning and Laura
21:02 also increased the performance of T I O U,
21:05 and also reinforcement learning also increased our performance, too.
21:12 And this is our qualitative example.
21:14 So, we cannot show the exact Safe for us sample directly.
21:17 The content is genuinely harmful.
21:20 So, you can see our model localize
21:22 the harmful segment much more precisely than the baseline.
21:27 Yeah, finally, our everything source is open source now.
21:30 The supervised fine-tuning pipeline and Megatron
21:33 Bridge are a post training code, and the full toolkit,
21:36 data creation, real LM inference, evaluation, and visualizing web demo,
21:40 all is on the our GitHub repo.
21:43 So, if you're interested in here, please visit our web page.
21:47 So, thank you.
21:49 Oh, thank you.
21:50 Wow, you shared all codes.
21:52 I think it's really valuable for all communities.
21:56 Uh If you have any questions,
21:59 please uh write the questions in the comments, please.
22:03 Um I will start with my questions first.
22:07 Um Do you have any uh tried new thing um
22:12 what have you have not done in your company before?
22:17 Uh in during hackathon.
22:18 Yeah, so in our uh actually,
22:21 temporal grounding ability is hard to inject in MLM model.
22:26 So, we should divide the chunk level approach first
22:29 like the Transit or other architecture for dividing the scene,
22:33 but uh in this case, we used a single forward for chunking level approach
22:39 in the like a chunk level description and the chunk level godly.
22:43 So, uh we think like firstly,
22:46 we think like we exactly MLM can do this only in single forward pass,
22:52 but when we fine-tuning with the appropriate data,
22:56 we can achieve some temporal grounding accuracy and godly performance, too.
23:01 Yeah.
23:02 Thank you.
23:02 And you analyzed about the video.
23:05 So, is there any interesting video you filtered or something like that?
23:10 Uh yeah.
23:12 So, there are very many harmful video, and uh I will talk about like uh like
23:18 Megatron Bridge mode uh when we train on Megatron Bridge,
23:21 if it is not trained with Megatron Bridge,
23:23 then the training the model and also like
23:27 a loading data and all kind of pipeline
23:29 is uh like there are a lot of bottleneck on the loading model or loading data,
23:34 but when we train on the Megatron Bridge
23:37 with the MV codec or many MV Nvidia's stack,
23:40 then we can use really efficient pipeline for training
23:43 video and also on the inference pipeline, too.
23:45 Great.
23:47 So, you have only two days Yeah.
23:49 to have any strategy to train a model?
23:52 Because training a model during two days is big, right?
23:56 Yeah, yeah, yeah.
23:57 [laughter] It's very big, but like very simply,
24:00 like Nvidia give like a little more GPU for track B participants.
24:06 So, we should fine-tune the model.
24:08 So, yeah, we used the one node of a H100 for two days.
24:12 That's a really expensive and also very good our resource,
24:16 so we can fine-tune our model,
24:18 and we should find more uh appropriate hyperparameter, too.
24:23 But like yeah, we can do some kind of baseline approach here.
24:27 Yeah.
24:27 Yeah, thank you.
24:28 Uh if you have one more day or two more days, uh what you want to do?
24:34 Yeah, we should uh find the appropriate parameter
24:38 for achieve uh temporal grounding and accuracy at once,
24:42 because now our parameter it can
24:44 achieve only temporal grounding uh performance now.
24:47 So, yeah, we should find that more now.
24:50 Thank you.
24:51 He was the the winner of the track B of the Nemo Tron hackathon.
24:57 Thank you.
24:59 And we are going to invite the winner of the track C,
25:03 and especially the winner is the all of the hackathon as uh one thing.
25:10 So, please join us.
25:15 And please show the uh image of the winners.
25:19 Yeah.
25:29 Yeah, winner is here.
25:30 Uh the team is from the Nota AI.
25:34 Please introduce yourself.
25:35 Uh hello everyone.
25:37 I'm Hancho Park from Nota AI,
25:39 the world best AI model optimization company, hopefully.
25:43 Uh thank you for inviting me.
25:45 Uh I'm work as uh tech lead in our company.
25:52 So, okay.
25:55 Uh in this hackathon, we uh proposed a novel uh MOE quantization
26:01 method in the perspective of the data set.
26:05 Uh we call our pipeline as Pascal MOE,
26:08 which stands for pipeline for activation balanced
26:11 and sensitivity aware calibration set generation for MOE quantization.
26:16 Uh as you know, modern LLMs increasingly adopt
26:20 MOE architecture because of its uh computational efficiency.
26:24 However, the memory usage is remaining high,
26:27 making the lower precision quantization method is required.
26:32 However, existing quantization method do not
26:35 consider the following MS MOE specific characteristics.
26:40 First, some experts are lowly activated during calibration,
26:44 resulting in poor activation statistics,
26:47 which naturally lead to large uh quantization errors.
26:51 The second is uh quantization sensitivity
26:54 experts must also be sufficiently covered.
26:57 The figure below uh describe the problem.
27:04 So, we hypothesize that more balanced experts activate during calibration may
27:09 reduce the quantization error by improving
27:12 activation statistics for most experts.
27:15 So, naturally, our goal is to develop the pipeline
27:19 to synthesize samples that better cover underutilized experts.
27:25 We also hypothesize that more
27:27 frequent activation of quantization sensitivity experts
27:31 may reduce quantization error by improving
27:35 the accuracy of their activation statistics.
27:38 Uh our pipeline will be uh generate the samples
27:42 that frequently activate quantization sensitivity uh sensitive experts.
27:49 Uh this slide shows our proposed method.
27:52 The proposed method seems quite complicated,
27:55 but uh it's quite uh straightforward.
27:58 First, our pipeline investigated the target experts,
28:03 which uh either frequently activated or sensitive ones,
28:08 uh using the inference stage.
28:11 Uh And then, uh using the those the target target experts,
28:17 we can uh our system can identify the target tokens.
28:20 The target token will be a set of context and token pairs.
28:24 And then, we assign the label to the target
28:27 tokens in the initial uh the data set like this.
28:32 Uh With the uh in the next page,
28:35 with the predefined predefined prompts, with the sample annotated samples,
28:40 we can generate it the patterns,
28:43 which just describe what kinds of the samples the system has to generate.
28:49 Uh And then, we uh use the Nemo uh Nemo data designer
28:54 to synthesize the samples that we want uh with the generated Pascal guideline,
29:02 and then we replace it undesirable samples with the newly generated samples.
29:09 Uh this slide shows an example of the output from each case.
29:14 Uh interestingly, we observed that intuitively less common
29:16 uh less natural tokens tended to activate underutilized experts,
29:22 such as backslash respect or double backslash.
29:25 And uh if as you can see this picture uh below,
29:31 the same shading indicate the rules
29:33 and the synthesized expression that follow the rule.
29:37 Uh we observed that the generated guideline was very interesting for us.
29:43 Uh in in In to uh evaluate uh the our proposed pipeline,
29:50 we contest the QN3 MOE model using
29:55 the various kinds of the NVIDIA software stacks.
29:59 The quantization scheme was into four.
30:03 We used the GPTQ algorithm.
30:08 So, first, as you can see in this slide,
30:12 our proposed pipeline generates the more target target samples
30:19 in terms of the expert balancing or sensitivity experts.
30:26 This slide shows the very impressive research.
30:30 Our proposed pipeline outperforms the the NeMo
30:37 training data that has contain undesirable samples.
30:45 Although I wanted to show more things, but this is last page.
30:51 Thank you for your attention.
30:52 Yeah, thank you for your presentation
30:54 and congratulations again for the first prize.
30:58 And actually, this is what we expect during the hackathon exactly
31:03 because we can generate something easily in this era with AI's,
31:09 but for the data set, it must be useful, right?
31:14 But you made a data set and it is applied to your product as well.
31:19 So, it was really amazing.
31:22 I want to ask about the experience of NeMo data designer.
31:26 Yeah.
31:29 I think that NeMo data designer is a very
31:33 impressive for us because everything has been kept.
31:36 We just gave the prompt.
31:38 We just gave the guideline,
31:40 and then the NeMo NeMo designer generated so many things
31:43 so many samples which fit our which meet our expectation.
31:49 So, we will use this design this this software later,
31:55 and we hopefully we wanted to give our message to the I mean I
32:01 mean contribute to our proposed pipeline
32:05 to the to this designer design designer later.
32:09 Yeah.
32:10 Great.
32:11 Yeah.
32:11 And there is a good question.
32:14 Sky World Horizon asked about how do you define quantization sensitive experts?
32:21 Yeah, it's very important, but I missed it because of the time constraint.
32:25 Actually, the Yeah, quantization sensitive during the hackathon actually
32:30 there are a lot of definition of the sensitive experts,
32:33 but we simply define the sensitive expert
32:38 as the the largest uh magnitude of the weight.
32:43 I mean the we sorted the maximum magnitude of each weight,
32:49 and then we define the oh,
32:52 this is this this expert very sensitive because the weight magnitude is so high.
32:58 Yeah.
32:58 This will be Yeah.
33:01 I mean we we in the future work, we will propose a method to select
33:07 automatically select the the sensitive or expert.
33:11 Yeah, really great to hear the details.
33:14 During the hackathon,
33:15 there was a great experience one more thing because was demo session because
33:21 we we shared our knowledge and experience during the hackathon in the demo.
33:26 So, it was really great.
33:28 And one more good questions, username Hobby 75 and 24.
33:34 Do you think this method will work for other quantization algorithms as well?
33:40 Yeah, of course because as I mentioned before,
33:43 the our method is not just for the algorithm.
33:46 It's our our method is related to the data set itself.
33:51 So, our method can be used any kind of the quantization method.
33:57 Yeah.
33:58 Thank you.
33:58 Thank you for your attention.
34:01 Please visit GitHub and you can experience the code as well.
34:05 So, please turn the our sketch video of day two, Zach, please.
34:16 We had a great day.
34:51 Yay!
34:55 [laughter] Absolutely incredible presentations.
34:57 I mean that that again, that's uh Imagine that all of the projects that we
35:03 watched were close to that caliber of of presentation.
35:07 Obviously, these guys won for a reason,
35:09 but there there wasn't a big split between the top
35:15 the top projects in the in the next best projects, right?
35:18 So, we watched uh 20 or so presentations.
35:23 Yeah, right.
35:24 [laughter] All of them so good.
35:28 Yeah.
35:28 Yeah.
35:29 We had a we had a bunch of you you know
35:32 uh time to hear about the the projects in the the hackathon.
35:37 May- maybe we'll spend some time though
35:39 answering some questions from the the chat.
35:42 I saw a few earlier.
35:44 I also saw some legends from our community like Cody's Guide to AI.
35:49 Thanks for coming to the special live stream.
35:54 Guys, we have uh also some comments talking about how good the food is in Korea.
36:00 I can confirm that the food is extremely good.
36:04 I have had the pleasure of eating it all week,
36:06 and I'm a little bit sad to have to go back to Toronto,
36:11 but I'm going to make sure to find a good Korean spot there.
36:13 So, maybe let's let's scroll up a little bit here in the chat,
36:18 and let's see I know we had a question early on.
36:23 Maybe we can uh talk to that one.
36:29 Yeah, okay.
36:29 So, at Hamad Raza 584 asked, "How can you use AI agents created
36:36 with NeMo agent toolkit with actual physical AI agents?
36:40 Is there any integration with ROS 2 operating system?" So,
36:44 this is a interesting question, Chuanlu.
36:46 I don't know your your familiarity,
36:49 but do you know of any way that we can actually
36:52 hook these LLM brains up to maybe some arms and legs?
36:56 Maybe we can add some skills, and actually agents will build new skills as well.
37:05 So, it is possible, I think, and there are a lot of documents there.
37:09 So, you can easily add some skills with AI.
37:13 Yeah, and I think, you know,
37:15 we obviously we do have a lot of this is a NeMo Tron Labs live stream,
37:19 so we're we're we're kind of focused on the NeMo Tron series of models,
37:23 but NVIDIA has a a huge array of physical AI models.
37:28 Yeah.
37:28 We do have a physical AI stream as well that you can check
37:33 out if you wanted to look at that where they discuss more about robotics,
37:37 but things like Groot, Isaac Labs.
37:40 Yeah, Cosmos.
37:41 Cosmos, of course.
37:43 You know, if you want to get started
37:44 with physical AI or transitioning from LLM's to physical AI,
37:48 lots of great technology for you there.
37:51 And of course, it's only going to get better.
37:53 This is the worst that robots will ever be, which is which is crazy.
37:57 Yeah.
37:58 Uh we have another few questions.
38:02 NeMo Squid from Angel Romero.
38:04 [laughter] Someone asks Ad Board Tomato from YouTube asked,
38:09 "What do you think about the NeMo Tron challenge?
38:11 How do I get past 0.84 without training all the linear layers?" So,
38:15 this is actually in reference to our NeMo
38:17 Tron reasoning challenge that we're running on Kaggle.
38:20 Oh, really?
38:21 Yes.
38:21 Uh I mean, I would there was a midpoint
38:26 that just happened where the team actually you know,
38:32 chose the best project so far,
38:34 and I think that that that write-up obviously has some great ideas.
38:38 If you wanted to look at that to get a little bit of inspiration.
38:41 So, what what what are your approaches when you're trying to squeeze
38:46 the the last few drops of juice out of training a model, Chuanlu?
38:51 For me, just check the data set.
38:57 Just make a more validate and good quality of data set.
39:02 So, we I should check about the details from there.
39:05 I analyze it and filtering some bad cases, or I should find the failure cases,
39:12 and add some data set for that as well.
39:15 So, I think the problem is just data set.
39:18 I mean, a classic and very good answer.
39:22 Yeah.
39:22 Look at your data.
39:23 Make it better.
39:25 If you don't think you can make your data better,
39:29 you haven't looked at your data enough, right?
39:31 So, I think it's also like similar to K food.
39:34 Like if there is a good ingredient and good mind, yeah, it it comes out.
39:40 There you go.
39:41 That's [laughter] the secret.
39:42 Good ingredients and good minds make good models and good [laughter] food.
39:48 That's good.
39:49 That's good wisdom for for Friday.
39:53 Cody's got AI asked, how do we open this box up and add the NVIDIA agent
39:57 toolkit so it refines itself the whole time and stops failing tools.
40:00 What is the best local model to use for Nemo claw and how to make agents?
40:04 Cody's got many great questions there.
40:07 Community legend here.
40:09 I mean, maybe let's just think about the first one.
40:12 So, you know, we saw in the first project from Naver this idea
40:17 that we could build an agentic system to kind of collect feedback.
40:21 I think Cody's got AI is asking,
40:23 how do we take that feedback and then roll it back
40:26 into the the technology to to make it better over over time.
40:30 What are your thoughts on that, Chuan?
40:35 [sighs] Keep trying actually,
40:35 but based on the model there is ability of the two callings metric as well.
40:44 So, choose good great model first and try keep make a great loop for that.
40:52 If there is a failure,
40:53 try paint it place made it try again for the great loop, I think.
41:01 Yeah, I mean yeah, I think you hit the nail on the head.
41:05 I mean, at the end of the day, right, we we are enter we're seeing more agents
41:10 do this kind of recursive self-improvement thing, right?
41:13 Where like Hermes agent or something like that is going to you know,
41:16 edit its own skills, its own config, its own even sometimes, you know,
41:22 you can you can work with it to enable
41:25 it to edit the way it's implemented, right?
41:27 So, it improves over time.
41:28 I think though for the first project, the the the Naver you know,
41:34 kind of collect this feedback project,
41:36 it's really important for companies, right?
41:37 To know what people think about
41:40 their products and where they're finding problems
41:42 so that those companies can take
41:43 those insights and make the product better, right?
41:46 I remember on day one talking to the to the Naver team and say,
41:50 "Hey guys, you know, like this is very useful even for us.
41:55 What the community thinks about [laughter] our technology, you know,
41:58 and if we saw there that someone's struggling with the the reasoning feature,
42:03 you know, we should do?
42:04 We should make it easier, more clear to use the reasoning feature of the model.
42:07 There you go." We're just scrolling through the chat here, guys.
42:12 So many different Oh, yes.
42:16 AI career path RLM or bust.
42:20 RLM, of course, recursive language model.
42:24 Some work done by Alex in the MIT SAIL team.
42:28 Really great research that a lot of people are very excited about.
42:32 If you haven't heard of it, definitely look into it.
42:35 Yes.
42:37 Simon Falk, interesting quantization shows
42:39 where efficiency meets fragility like biology.
42:42 It's not just energy, it's how it's structured.
42:44 That's where intelligence holds.
42:45 What a comment.
42:47 Absolutely true though.
42:49 Quantization is a is a big thing.
42:56 [laughter] You know, as as as the as the team described, right?
43:02 But not not an easy task, especially when it comes to MOE.
43:06 I see you have another question.
43:08 Chuan, maybe you'd like to read this one.
43:10 This one?
43:11 Yes.
43:13 This this one.
43:16 Oh, have you tested Nematron Cascade 2 or Nematron 3 Nano with Nemo claw?
43:24 Yes.
43:25 Yeah, so Nemo claw is the NVIDIA runtime for Open claw.
43:31 We've done a few live streams about it recently,
43:34 which is why we're getting some questions about it.
43:36 I would say yesterday and the day before actually we
43:40 were showcasing the model using both of course Nematron 3 Super.
43:47 I'm very biased, but that's my favorite model.
43:50 And and the Gemma 4, the new Gemma model.
43:55 Both running in through Open claw and then
43:59 as well you can run them through Nemo claw.
44:01 So, I would say those are great models to start with.
44:03 In terms of a stacking the models, right,
44:07 we we're we're actually going to have some live stream
44:11 coming up in a few weeks where we're going to get
44:13 a little bit into more the weeds on how you
44:16 might stack models or sub-agents with the kind of claw technology.
44:23 So, Cody's got AI, make sure to tune in the future,
44:26 but yeah, Nematron Cascade 2 and 3 Nano are both great models.
44:31 I would tend to say that something like
44:34 Super Caliber is better for the main agent of your claws and then we're going
44:40 to use things like Nano and Cascade as sub-agents.
44:44 Are you big claw guy, Chuan?
44:46 Are you you know, are you are you out there clawing it up?
44:51 [laughter] Yeah, you said.
44:52 I I think we are adopt with elements agents nowadays.
44:57 So, we are really experiencing the self-healing agents as well.
45:03 But we are having struggle with tokens, right?
45:09 Yeah, because that's going to burning out.
45:12 So, we need our own own models sometimes.
45:15 So, we need Spark or some our own cloud stuff.
45:21 So, you mentioned about Gemma 4 or our models as well.
45:25 But I want to mention one more stuff for the quantization because there
45:29 is a model quantized with MBF 16 for the Gemma model as well.
45:35 So, you should try with that.
45:37 Yeah, famously DGX Spark, right, has has works great with NVFP4.
45:46 So, any model that's available in NVFP4 format is
45:50 going to run a little bit better on the Spark.
45:53 I do think at the end of the day, you know,
45:57 there's a lot of times that your agent's going to do work
46:00 that isn't you don't need to use a lot of expensive tokens.
46:06 Grepping through the file, you know, to to to find some snippet or you know,
46:12 creating a directory or you know, making a small edit to a to a code file.
46:17 We don't need we don't need Claude Opus 4.7 for that, right?
46:21 We can get away with [laughter] we can get away with Gemma for that.
46:25 Yeah.
46:27 Excellent stuff, guys.
46:29 We have NVFP4 is dope.
46:32 Yes, true.
46:35 What's the best open source to use with Nemo claw?
46:37 What tool fits best with Quantum?
46:40 Well, I've never experienced about that actually.
46:43 [laughter] Well, what I can tell you is that Nemo claw
46:47 is designed to work well with all kinds of different things.
46:50 So, obviously it should work with NVIDIA technology,
46:53 but should work with most everything else as well.
46:56 The idea is to you know,
46:58 to have a platform that works well with whatever you're using.
47:01 And if you don't know what you want to use
47:04 or you have an idea that we haven't put in yet,
47:07 please contribute to the repository, contribute to the project.
47:11 And if you don't know how to contribute to the project,
47:13 you should check out Tuesday's live stream where we
47:16 talked all about how you can contribute to open source.
47:19 You know, if you if you're like, I have a great idea.
47:21 I wish it worked with this or that, please contribute it.
47:24 And if you have a Quantum model that isn't
47:27 yet supported by Nemo claw, please submit a PR.
47:33 [laughter] It would be that would be great.
47:35 I I think we have some some exciting Quantum stuff coming up,
47:40 but for right now that support, if you need it, put it in.
47:49 Yes.
47:49 We have some time for waiting for the great great comments.
47:53 We I want to mention one more thing during the hackathon.
47:57 We had a great session with a meditation by Nematron.
48:02 Zach, can you please show the image from the for the meditation session?
48:10 So, [laughter] we had a meditation with Nematron.
48:14 That's right.
48:15 Yeah, Nematron Super generated a text for the meditation,
48:20 especially for the developers.
48:22 So, the text was like please forget the error messages or the stacked PR.
48:32 Just relaxing.
48:34 We had a second day of the morning that we had a meditation.
48:39 So, it was really great,
48:41 relaxing and we can focus on the hackathon after the meditation as well.
48:47 That's right.
48:48 I mean, at the end of the day, Nemotation.
48:52 [laughter] Yeah.
48:53 That's from IT Sucks.
48:56 I mean, listen, hackathons are long.
49:00 You have to stretch your mind a lot.
49:02 It's nice to take a little break, little breather, refocus,
49:05 recenter as guided by Nematron 3 Super who's a pretty chill guy.
49:12 There you go.
49:14 Also, I think you you had a voice to or text to speech.
49:18 So, you can just close your eyes.
49:20 And if you bring the image back up, Zach,
49:22 little Easter egg for people who know the developer community team,
49:26 you'll see in the back left corner there, that's Mark Heat.
49:30 [laughter]
49:32 An absolute legend in the developer community effort here at uh NVIDIA.
49:38 So, uh Nemo Zen is great.
49:41 Raphael, uh absolutely great Nemo Zen.
49:44 He was really busy with uh Build a Claw session, right?
49:48 yeah.
49:48 Build a Claw was popping off, guys.
49:50 I Uh if you're going to uh an NVIDIA event
49:54 in the future and you're wondering all about claws, Mhm.
49:57 you can probably find a Build a Claw tent near you uh that will
50:02 help you understand this technology a little
50:04 bit better and get hands-on with it, see some sweet demos.
50:07 Mhm.
50:08 One of the things that I was really excited about, Uh-huh.
50:10 Cha- Chan- Lon, is, you know,
50:12 here in in Korea you've built a really impressive community of AI practitioners
50:18 Yeah.
50:19 called Pseudo Labs.
50:20 And they actually showed up and they helped out
50:22 at the Build a Claw event showing demos, helping people.
50:26 So, it wasn't just like NVIDIA came to town and all from HQ like, "Oh,
50:29 blah blah blah." It was actually local devs from your community
50:33 that were helping people use un- understand this technology.
50:37 If someone wants to learn more about Pseudo Labs,
50:39 so they're excited by what we're doing and they're
50:42 excited by maybe digging deeper on nights and weekends,
50:45 what where should they go?
50:46 How do they find Pseudo Labs?
50:48 Uh actually, we are based on Discord and we
50:52 are non-profit AI and data science research uh community.
50:56 So, you can just come on in and you can
50:59 visit our uh there are several rooms and our timetables.
51:04 So, you can just come in and you can
51:06 have experience about uh agents or researching AI models.
51:10 So, physical AI as well.
51:13 Just there are a lot of projects there.
51:16 And actually, we invite them into our Build a Claw session,
51:21 but they're all uh employees from other companies like big enterprise,
51:27 but they had a a PTO for uh uh for visiting our session and not just visiting
51:36 uh helping our uh session and they they
51:40 shared about their own agent Claw experience for us.
51:45 It was amazing, I think.
51:47 Yeah, it really was amazing.
51:49 I mean, these are working professionals that took time off of work
51:52 to come help out uh to help people understand and access this technology better.
51:58 Uh that's the vibe.
51:58 That's been the vibe the whole time
52:00 that we've had the pleasure of being in Korea.
52:02 Uh it is just absolutely amazing.
52:04 I can't wait to see what you, Chan- Lon,
52:06 and the team that you don't see behind the camera here, guys,
52:09 uh get up to uh to to keep growing that that community and keep you know,
52:15 helping the the the kind of AI you know,
52:29 work being done in These live streams are fun, Mhm.
52:33 but there's a whole ecosystem of things you can learn from.
52:36 So, uh if you're hearing in say in the in the track A project, right?
52:42 You're hearing about agents, but you're not quite sure what they are,
52:45 what they mean, uh then uh you can check out these learning
52:49 paths at developer.nvidia.com/topic/ai/how-to-build-an-AI-agent.
52:54 We've got some updates coming that should uh make
52:57 this more uh related to claws in the near future.
53:00 Zach, if you go to the next slide, please.
53:03 Uh we also have, of course, our NVIDIA Nemo NeMo Tron repository.
53:08 This is where you can find how to run the model,
53:10 how to build a model from scratch yourself if you wanted to.
53:14 Uh you might need more than a 8100 node and 36 hours though,
53:19 but you you could build your own NeMo Tron 3 Super if you wanted to.
53:23 Uh and then, of course, we've got another slide.
53:26 Uh if you want to join us at discord, we've got discord.gg/nvidiadeveloper.
53:30 Once you join that, you can look for the hashtag NeMo Tron models channel.
53:34 Uh you'll be able to find me in there uh 24/7.
53:38 Uh even if you send me a message in uh in Korea time,
53:41 I'll probably respond pretty quickly uh because I'm I'm up all night.
53:45 So, uh but leave any questions you have in there for us.
53:49 Uh if you have any questions for the teams,
53:51 you can also put those in that channel uh and I'll
53:54 make sure that they get forwarded to uh Chan- Lon,
53:58 who will make sure that it gets forwarded
53:59 to our uh to our our incredible winners.
54:01 Yeah.
54:02 Next slide, please, Zach.
54:04 Uh if you have any questions for the NVIDIA developer community team,
54:08 so this is people like Mark Heeps and Zach,
54:10 please reach out to community@nvidia.com.
54:13 Uh Zach does live in a cave.
54:15 It's always dark and the only screen he has is this email,
54:19 so he'll answer your question if you send him an email.
54:21 I'm just kidding.
54:22 Uh we let him out from time to time.
54:24 Uh [laughter] yeah.
54:26 He's not an intern either.
54:27 He's a real full-time employee, guys.
54:29 Uh otherwise, I've one last question for you,
54:33 Chan- Lon, well we've got uh about 5 minutes left.
54:35 Yeah.
54:36 Uh I've I've enjoyed my time so much here.
54:38 Uh I've enjoyed getting to meet all of these incredible uh developers.
54:43 Uh you know, uh every time I get to visit somewhere new,
54:46 I'm so excited to see how passionate uh people are about AI,
54:49 but how did you get into AI?
54:51 What what made you decide what you want to do
54:54 with your life is AI and not just learning about it,
54:59 but then teaching people and sharing it with a community of people.
55:02 What What uh gets Chan- Lon out of bed for teaching AI every morning?
55:12 [laughter] Actually, I started with uh experience of the Kaggle.
55:16 Actually, I I grew from that community.
55:20 So, every day I see the leaderboard from there and then
55:25 I see the improvement me and the community as well.
55:29 So, nowadays there are a lot of AI stuff in all communities,
55:34 all interesting stuff.
55:36 And so, when I woke up, I just see what's going on today like,
55:43 you know, uh today also there are big news about uh AI.
55:47 So, GPT-4.5 is released.
55:51 As like that, if something new AI is coming out,
55:55 we experience we want to experience all things, right?
55:59 Yeah, I started with that and new experience gives me new life, I think.
56:05 Yeah, that's coming from what I do every day, every morning.
56:09 And I start from that experience from uh since Alpha Alpha Alpha Go.
56:14 So, I every day every morning I experience new AI stuff.
56:19 So, with agent, it helps me a lot to experience more
56:25 and more and a lot of experience I knowledge it as well.
56:28 So, it's incredible life and new new life and with NVIDIA
56:34 community and NVIDIA models and supported by our GPUs.
56:40 So, it was amazing.
56:42 Incredible stuff.
56:44 Incredible answer.
56:45 Just what you'd expect from someone who's uh who's who's helping
56:49 so many developers uh across the the country and the world.
56:53 Uh one last question in the chat,
56:55 when is NeMo Tron Ultra releasing approximately?
56:57 Thanks so much.
56:58 It'll come out when it comes out, guys.
57:01 I know every every stream you want to know.
57:04 It will release when it releases and uh I'm sure that we'll tell
57:08 you uh you should expect to see uh that we we announce it.
57:13 So, uh it won't sneak up on you.
57:15 Uh you'll know.
57:16 We'll we'll we'll tell you.
57:17 Uh thank you so much for your time today, Chan- Lon.
57:20 Thank you for sharing uh the office with us.
57:23 Thank you for coming to Korea here.
57:25 Of course.
57:26 Absolutely amazing.
57:28 Kamsahamnida.
57:29 Kamsahamnida.
57:29 Uh Kaja.
57:32 Kaja.
57:33 Kaja.
57:33 Okay, thank you, everybody.
57:35 We will see you next time.
57:36 Bye.
57:36 Bye.
57:36 Bye.
57:37 Bye, everybody.