Azure AI Infra updates to power frontier and enterprise workloads | BRK179

Azure AI Infra updates to power frontier and enterprise workloads | BRK179

Microsoft Events

0:00 Alright.

0:00 Hey folks, welcome.

0:02 Thanks for coming to Good Night.

0:05 My name is Matt Vegas.

0:06 I'm a principal Product Manager for the ND series.

0:09 This is our AI infrastructure virtual machine.

0:13 I'm here today with Saksham from Black Forest Labs

0:16 and Param who's on the software side of our AI infrastructure.

0:21 And today we're going to talk about our Azure

0:24 AI info updates that Power Frontier and enterprise workloads.

0:30 So why are we here?

0:32 You know, like, why has Microsoft for the last

0:35 several years led the way in AI infrastructure?

0:39 You know, why have we been building

0:41 these supercomputers long before other CSPS were interested,

0:44 you know, even long before ChatGPT existed?

0:48 These kinds of systems are really expensive

0:50 and complicated and they're tough to deal with.

0:53 And there's really two reasons that it starts with, you know,

0:58 Microsoft ultimately is a user of supercomputers,

1:01 whether it's GPU compute or CPU compute.

1:05 It's been a pretty integral part of our design process,

1:08 you know, whether it's hardware or silicon software.

1:11 The second one is, is, you know, we're in it now,

1:14 but this AI boom, you could have seen it coming.

1:16 It was not a question of when, but just it was not a question of if,

1:20 but like when was it actually going to happen.

1:23 And Microsoft has just this vast portfolio of software

1:27 products that add value to all of our daily lives.

1:31 And it's, I see it as the perfect

1:33 vehicle for delivering AI and delivering value through AI.

1:38 So this need for AI and really high end

1:42 infrastructure has always been intrinsic to what motivates us.

1:47 It's what drove us to build this competency.

1:49 And really it's now it's,

1:51 we get this wonderful privilege to kind of pass on the harder knowledge,

1:56 expertise, and really our infrastructure to our customers.

2:00 And so, you know, Microsoft, we're a provider of AI, but we're, you know,

2:03 at the heart of it is,

2:05 is we're really a user of all of these, of this really advanced infrastructure.

2:10 So and the, the need for this compute doesn't end with Microsoft.

2:15 Obviously the need for compute is

2:18 basically exponential across our customer base.

2:21 You know, over the last several years,

2:23 model size in terms of parameters has grown massively.

2:28 You know, each new model growing exponentially over the prior generation.

2:33 We've seen, you know, pure text models evolve into multi model.

2:37 Now we're seeing the rise of like agentic workflows,

2:40 long chain of thought these, you know,

2:43 like really complicated AI processes that are capable of doing,

2:47 you know, full workflows and deep research on behalf of the user.

2:52 And so it's, it's really advanced, you know,

2:55 as we've kind of followed this trajectory

2:58 and maintaining this path of innovation

3:00 for a customer is really what drives us to just build larger,

3:04 higher performance systems on the latest GPUs.

3:08 And so kind of Needless to say, we've been very,

3:11 very busy over the last few years just building out

3:14 a huge footprint of supercomputing infrastructure all around the world.

3:21 So I'm sure you guys read the news, it seems like everyday there's, you know,

3:26 a new massive GPU supercomputer being announced by one of the many

3:31 competitors out there that are building these types of things.

3:36 And I feel pretty strongly that Microsoft has always been,

3:40 you know, several steps ahead.

3:43 You know, like if you read the news recently, we announced Fairwater,

3:47 which is the world's most powerful data centre for AI training.

3:52 Fairwater is multiple whole regions of AI compute.

3:57 If you typically think of a region

3:59 as something that hosts all types of services,

4:01 but these are whole regions basically just designed to do this.

4:05 We're talking hundreds of thousands of GP us connected

4:08 in a massive network and all of this is real.

4:11 Like it's on the ground.

4:13 It's basically in production and it's not some future plan

4:17 on paper and an announcement like it's this is real.

4:22 Over the last few years, we've grown our worldwide fleet to support the daily

4:27 needs of inferencing that we all now use.

4:29 And it's like pivotal and our, our, our work,

4:32 whether it's copilot or things like ChatGPT.

4:36 An example of this is from mid 23 to 24, we grew our fleet by 30X.

4:42 Like so for less than a year, just the amount of supercomputers that we

4:45 host around the world has grown tremendously.

4:49 And all of this, you know,

4:51 basically comes back to just our key principles of building everything with HPC,

4:55 you know, high performance computing, it's top of mind.

4:59 We build custom platforms like fine-tuned for performance.

5:03 We deliver all of our products basically first to market,

5:06 so right when they're announced and released and we take,

5:09 you know these really advanced designs that are used

5:12 in the highest end supercomputers and we give everyone access to them.

5:16 Like all of these things are built the same

5:19 with the same like bar for quality and bar for performance.

5:27 Our investments on the product side span the entire

5:30 stack and this is what enables our customers to succeed.

5:34 You know, this year we introduced GB 200 GB 300 RTX 6000 Pro based VMS.

5:40 We continue to sell out our massive fleet of A1 hundreds

5:44 and H1 hundreds and AMD based MI300X based systems,

5:48 but it takes more than just, you know,

5:51 advanced GPUs to deliver a good cloud service.

5:54 We've invested heavily in CPU compute for pre processing.

5:58 Also 800 Gigabit XDR back end networking that was released with GB 300.

6:03 We have managed luster and object storage services,

6:07 cycle cloud and Kubernetes optimizations.

6:11 All of this is basically to ensure that, you know,

6:14 we have what we need to efficiently, that we have what we need for our customers

6:19 to efficiently support and run their AI applications.

6:26 We also continue to invest in the future of AI.

6:30 We're really excited this week at Ignite we're we've

6:35 launched the GA of ND GB 300 virtual machines.

6:39 These are for AI training and inferencing.

6:42 It's now generally available, which is a huge milestone for us.

6:45 We're also excited to announce Azure NCRTX Pro

6:49 6000 BSE is now available for public preview.

6:54 This is a cool platform.

6:55 It's a single platform for both AI and visualisation, visual compute, basically,

7:00 so that those two things can work together to accelerate enterprise workloads.

7:06 In addition, this week we're excited

7:09 to announce our continued partnership with NVIDIA

7:12 and we're that we're basically dedicated to building the future of AI

7:17 on Verirubin and we're also dedicated to our continuing our partnership with AMD

7:23 to offer more performance and flexibility

7:26 in the cloud with their upcoming MI40455X GPU.

7:32 So whether it's investments in our GPU systems

7:36 or our infrastructures stack or offerings of full stack AI solutions,

7:42 AI innovators choose to run on Azure.

7:47 This includes open AI copilot.

7:50 Just today Anthropic was announced to be a customer of ours.

7:55 Copilot runs on it Figure 8 byte dance meter,

7:58 open AI and Black Forest labs who are so,

8:01 you know, we're we're really excited to have just these phenomenal customers.

8:07 So a little back story before we kind of go to our next segment in this slide,

8:13 but 2025 has marked a pivotal shift in AI infrastructure.

8:17 It's really where we went from X86 based surfers to grace Blackwell,

8:23 which is a completely new architecture

8:26 that with advancements basically across across the design.

8:32 But you know, it has a arm processors,

8:34 it has a new envy link that connects the entire rack and enable

8:39 and basically enabling the whole rack to act as a massive supercomputer.

8:44 And so today with us, we have Black Forest Labs,

8:47 who are renowned for their ground breaking text to image generation models.

8:51 And we're one of the first people to get

8:55 access to a really large GB 200 cluster.

8:58 And so today I'm excited to have them here.

9:01 So thank you.

9:05 Hi everyone.

9:05 A bit of introduction about myself.

9:08 I'm Saksham, Consul and the lead infrastructure efforts at Black Forest Labs.

9:12 I joined about a year ago from a start up in the LLM training and serving space,

9:16 which data got acquired by AMD.

9:19 So I'm responsible for the training and inference infrastructure,

9:22 so to train and also server models at scale.

9:26 So a brief description of what Black Forest Labs is.

9:30 So Black Forest Labs is a leading AI research lab for visual intelligence.

9:34 We, we're not just building models in my opinion,

9:37 we're actually building foundation technologies that shift

9:40 how we see and understand the world.

9:43 So our latest offerings already are integrated

9:46 into previous domains and are built into production workflows,

9:49 not just in the creative domain,

9:50 but also in design software as also social media experiences.

9:54 And I truly believe that we are turning imagination

9:57 into reality for with three things that actually matter in production,

10:00 speed, precision as well as creative control.

10:04 So we launched into August 2024,

10:05 we were backed by Anderson Horowitz and General Catalyst.

10:09 And right now we have about 50 full time employees based in two headquarters,

10:13 one in Fryeburg, Germany as well as San Francisco.

10:17 So before BFL, our founding team pioneered

10:21 technologies that shifted the define generative AI,

10:24 things like VQ games, latent diffusion, stable diffusion.

10:28 These models fundamentally changed the the AI Gen.

10:32 AI landscape.

10:33 They made text to image generation accessible to millions

10:36 and set a new standard for what what's possible.

10:39 So in the last four years,

10:41 our work has been started over 100 and 5000 and 50,000 times.

10:46 So yeah, in terms of the founding team, there's Robin Rombach,

10:49 he's our CEO and he's leads the company's strategic

10:52 directions as well and they're doing building the company itself.

10:56 That's Andreas Blackman.

10:58 He leads a go to market.

11:00 He's involved for all our growth initiatives,

11:02 operations as well as partnerships.

11:04 And finally we have Patrick Asser, the man, the myth, the legend.

11:08 He leads a research at Black First Labs and he's

11:10 responsible for most of the cutting edge research that we do.

11:14 So in terms of like it's I'll be remiss to talk

11:17 about Black First Labs without talking about our flagship family.

11:21 So the, the Flux series of models are one

11:23 of the most advanced sets of image generation models out there.

11:28 Since launch, each and everyone of our models have topped

11:31 the benchmarks of both text to image and image editing benchmarks.

11:35 We the, our models are available both open weight,

11:38 so you could deploy this on premise and we also have

11:42 our high performance models available on our API and also on Azure Foundry.

11:48 So in our end, this gives us gives customers flexibility that you can

11:51 deploy them to match your security costs and integration needs that are there.

11:55 So in terms of the models that we have, we have the Flux 11,

11:58 which is like one of the high quality text image generation models out there.

12:02 And this is a workhouse that powers

12:04 a lot of production workflows out there today.

12:06 There's the ultra model, which is ultra fast and for ultra high res content.

12:11 So you can, if you want to get 2K images generated in seconds and not minutes,

12:15 that's the model you should go for.

12:18 And then we have the Flux 1 context model,

12:20 which is the state-of-the-art model in both

12:22 for in context editing as well as generation.

12:25 So this combines text and images

12:27 to generate outputs which are coherent and precise.

12:30 And finally, small teaser that we are releasing,

12:32 we have a new series of models coming out, which is the Flux 2,

12:36 which is setting a new standard for image editing and generation.

12:39 And this is this, this model will be something

12:41 that you can deploy to productions and not just for demos.

12:45 So like at BFL, we just don't talk about what we do, but we actually ship it.

12:50 So every image you see out there

12:52 on a slide right now have been generated by flux.

12:55 And feel free to scan the QR code and try out.

12:57 Or you can go to BFL dot AI slash play, generate the images yourself,

13:00 test the control and see where thousands of people use flux.

13:05 Thank you all.

13:09 Right.

13:09 So we're going to switch things up here.

13:10 We have a little bit of AQ and a that we're going to do

13:14 so so we can learn more about how BFL uses our GB 200 clusters.

13:18 So you know, Sacha, I'm working with large clusters

13:22 of GPUs is obviously not the same as, you know,

13:25 a general purpose compute infrastructure.

13:28 Walk us through your experience of onboarding an Azure

13:32 AI cluster and how you scaled to train Flex models.

13:37 Sure.

13:37 Yeah, you're absolutely right that like we were working with large scale GPUs,

13:42 it's completely different from like standard operations and infrastructure.

13:46 And it's not just on the, you know,

13:48 like when especially in training, it stresses the entire system to its limit.

13:52 So it's not just the compute plane that we normally talk about,

13:55 but also the networking, the memory and the and also the computer obviously.

14:00 So like with these challenges, like it's been really nice that Azure has

14:03 been a cloud cloud partner since day one.

14:06 So from the first customer,

14:07 have you been one of the first customers on the GB 200 NVL 70 twos?

14:11 And Azure has supported us throughout.

14:13 So in terms of training, one of the big steps is there is like data and you,

14:17 you know, when you're pre training,

14:19 there's so many processes that are touching the, there's so many

14:22 processes that are accessing the data and there's so there's so much,

14:26 the scale of data is such high that you

14:28 started the data infrastructure problem becomes pretty intense.

14:32 And if not handled properly,

14:33 you're going to have these weird issues like cache thrashing,

14:37 what else partition thresholds being met.

14:40 So in TLDR, like the performance gets really impacted.

14:44 So we had to work really closely with Azure,

14:46 the Azure support, the Azure Storage team, for example, to get to really improve

14:50 the performance and make these workflows work.

14:53 That's awesome.

14:54 So what are some other considerate consideration?

14:59 What are some other considerations you made throughout your experience on Azure?

15:03 You know, beyond just the GPU,

15:04 how did you get the infrastructure and services right?

15:08 You know how has Azure Infra basically helped you achieve your goals?

15:15 Right.

15:16 So yeah, in terms of like things apart from the compute,

15:19 I think that three things that I could talk about 1 is,

15:23 for example, like, you know, we talk about GB 200,

15:26 brand new, flashy new infrastructure available.

15:29 But it's not just the hardware.

15:30 There's, there's a lot of software that's

15:32 also involved and the Azure team and NVIDIA,

15:35 we worked closely together to to get new drivers,

15:37 validate on optimizations that we can do

15:39 on this new system to actually get the best performance.

15:43 So that's one.

15:45 Secondly, I would say networking and networking and storage.

15:50 So like pre training, like it's not it's intensive throughout the entire system.

15:54 So you know, you're you hit into these weird issues like IB,

15:58 IB flapping, you're reaching a storage thresholds, your switch are failing.

16:03 So, you know, working with the HPC and storage team,

16:05 that's really important to get, you know, to identify these common hot parts.

16:09 Yeah, one solution was like finding

16:11 these like preemptive load balancing, for example,

16:13 so that we can actually push through

16:15 the network without having to basically DDoS ourselves.

16:18 And lastly would be monitoring.

16:21 Yeah.

16:21 So when you're running pre training jobs,

16:23 these jobs are not just like running in seconds or minutes.

16:25 They run for days, in fact for months.

16:27 And then so it becomes anything can go wrong in this.

16:30 So identifying when it goes wrong,

16:32 what went wrong and actually fixing that as fast

16:35 as possible becomes essential for actually meeting our target velocity.

16:39 So that that's been something that, you know,

16:42 really helps with that has helped us out so far.

16:44 Oh.

16:45 That's very interesting.

16:46 And then how important is orchestration in managing a large cluster like this?

16:52 Can you talk a little bit about, you know,

16:54 how you landed on the stack that you're using right now, the benefits of Azure,

16:58 your experience running Slurm on Azure with Cycle

17:01 Cloud as kind of an orchestration engine?

17:04 Right.

17:05 I mean, yeah, without orchestration there is now point of a research

17:09 cluster and the orchestrator doesn't just do like, you know,

17:12 you're not just putting jobs and does the job allocation,

17:15 but also does a lot of optimizations

17:17 and getting the rewrite resources to run your jobs.

17:20 So to answer the question why we use slum,

17:23 like Slum is better tested for HPC workloads and with Cycloud's help,

17:27 we were, you had an AI, we had an Azure native solution,

17:30 which is really closely tied with Azure's infrastructure,

17:33 which helped us get to get to use the, for example,

17:37 the GB 200 rack topology system really quickly.

17:40 So traditionally, for example, Slum does not really understand what a rack is,

17:44 but for the GP2 hundreds, it becomes really important that you understand

17:48 what RAC is like that's really tight integration

17:50 of what what where the node is in the RAC and using the RAC accordingly.

17:54 So cyclic Cloud worked really closely with SCAD MD to design this RAC topology,

17:59 a block topology in cyclic cloud in slum itself,

18:02 which allows that we can allocate jobs

18:04 without having to change any other application layer.

18:08 So the jobs get you're doing RAC level job selection,

18:11 not node level job selection, which then really,

18:14 I mean, it really improves performance, yeah.

18:18 Now the, the pace of innovation in the GPU space has been crazy.

18:22 It's where like it was yesterday,

18:24 we're still installing A1 hundreds and now we've gone

18:29 through H100H200GB200GB300 and now we're talking about Vera Rubin.

18:34 I mean, it's it's just going so fast.

18:37 You know, how what was the experience like

18:40 collaborating with Azure to get kind of, you know,

18:43 first or early access to these GB 2 hundreds?

18:47 Also, in what ways did it help you having, you know,

18:50 access to just the absolute cutting edge, highest performing infrastructure?

18:56 Right.

18:56 I mean, it still blows my mind that when you talk about the GB 200 stack,

19:01 we are getting what 13 terabytes of VRAM and what

19:05 576 terabytes per second of high bandwidth memory and the like,

19:08 we get double the TT flops compressed, the H 100.

19:12 So you're right, the scale in which computer is moving is really,

19:15 really, really fast.

19:16 But and this really helps with training.

19:18 So training for both transformer models as well

19:21 as autoregressive transform transform models or diffusion models,

19:24 both of them are compute bound,

19:27 which means that as compute accelerates, training accelerates.

19:31 And things like higher band, higher memory,

19:33 for example, helps in like multiple fronts.

19:36 You're able to have able to put more weights

19:38 in, in a smaller amount of GPUs so they become more efficient.

19:41 That's less likely to break, much easier to work with.

19:45 You're still able to have higher batch sizes, which people normally think about,

19:48 but that helps with training performance quite a bit.

19:50 So those are a lot of things that really improves with like,

19:53 you know, having better infrastructure helps with training.

19:57 So the, the advantage is that working with Azure

20:00 and you guys have has helped that, you know, we've been such a small team,

20:04 we're able to leverage a lot of the support that you guys have

20:07 given us to actually use the cutting edge to, to, to, to win.

20:11 That's wonderful to hear.

20:12 And then for the audience out there who's maybe

20:15 get just getting started with Azure infrastructure, you know,

20:19 what are your like, say top three lessons learned, You know,

20:23 don't hold back the gory details if you need to.

20:27 All right.

20:29 So yeah, I mean this is I think relevant

20:31 in my opinion for all infrastructure, not just Azure.

20:36 But one thing that for me was some a really good

20:38 lesson was like learning how to Co design with your cloud provider.

20:42 So it's not just like finding issues which you will have and something like,

20:45 you know, complaining about it,

20:47 but actually sitting down with the team and discussing

20:49 on like what's on my memory access pattern,

20:52 what's my data access pattern and finding out

20:54 these, finding out these solutions that we can do.

20:57 So we were obsessed with collecting metrics and logs

20:59 and sending them to your sending them over,

21:01 not just to understand the system better,

21:04 but also to accelerate the improvement.

21:06 And to Azure's credit, you guys were really quick in getting those fixes in.

21:10 So I believe like, you know, working together in tandem really helped us there.

21:15 Secondly, I think what's really important is like, you know,

21:19 focus on the focus on the entire system and not just on the on the GPUs.

21:24 Like we always come like we always think that, you know,

21:26 training that's GPUs and that's about it.

21:27 But you're, you're, it's an entire system that's going to be stressed out,

21:31 networking, storage, memory, all of those things.

21:34 So thinking of in a holistic manner

21:36 and finding out where your actual actual bottlenecks are,

21:39 it's typically not sometimes a GPU, not always.

21:42 So like using ARM valves principle, finding it out, it's really useful.

21:45 And secondly, when you're working with like really high performance,

21:49 like when you're using something new like bleeding edge,

21:52 you put put in some time or allocate some time

21:55 for actually that there might be some engineering effort required.

22:00 So you're putting some work to get the work out to be to be honest.

22:03 So like if you're getting 2X computer, computer,

22:07 computer improvement thing some for engineering development and time.

22:12 But yeah, those are three things I will think about, yeah.

22:15 That's wonderful.

22:17 Thank you so much for joining us.

22:19 It's been an honor having you here on stage with us.

22:22 You're phenomenal customer.

22:23 Thank you.

22:24 Thank you all.

22:25 Right.

22:25 So we're going to kind of switch things up a little bit now.

22:27 Sorry, Param Shah is going to present on different side

22:32 of our stack which is the software side of our infrastructure.

22:36 Yeah, awesome.

22:37 Thank you, Matt.

22:37 Thanks, Saksham.

22:39 So now that we've seen the hardware side of Azure's AI infrastructure,

22:42 let's shift gears to the software side and orchestration,

22:46 because that's where the GP power really comes to life, right?

22:50 Because at the end of the day, you can have all the latest and greatest GPU's,

22:54 but it's also about how you can deploy,

22:57 manage and scale them for your workloads.

23:02 So across Azure's AI software stack,

23:04 we've designed our offerings around how you build AI.

23:08 And at a high level there's two main paths, right?

23:11 If you want to start with pre trained models, you should go to Azure AI Foundry.

23:15 Right?

23:16 Here you can focus on fine tuning and agent creation.

23:20 If you want to train your own custom models,

23:22 you can use Ask if you're already on Kubernetes or Cycle Cloud if you

23:26 want to bring in your own scheduler

23:27 and get your training infrastructure up quickly.

23:30 So on the left, as you can see, you get a managed experience right where you can

23:34 build use built in models and tools for fine tuning.

23:38 And on the middle, you can see the Kubernetes our ask service,

23:41 where if you have containerized pipelines with DevOps integrations,

23:45 then that's your natural fit.

23:48 At the end of this presentation,

23:49 I've also linked some fun ask and AI Foundry sessions

23:52 that you guys should feel free to go look at.

23:55 And at the very right went too far.

23:59 Let me go 1 slide back.

24:00 Cool.

24:04 Yeah.

24:04 So on the right side, if you look, we have Cycle Cloud where if you're bringing

24:09 in your own scheduler or workflow manager like Slurm

24:11 or PBS Pro and you want that fine

24:14 grain control over your distributed AI or HPC workloads,

24:17 then Cycle Cloud is built for that.

24:19 So today we're going to focus on that right side, right,

24:22 the customizable side of the spectrum where teams can orchestrate compute

24:25 at massive scale with end to end observability and deep insights.

24:31 So before we get into that, it's going to give

24:33 a high level overview of what ask is Azure Kubernetes Service.

24:36 So that is our fully managed Kubernetes platform for AI.

24:40 So it provides GPU optimized infrastructure, it gives you dev friendly tooling,

24:44 you get intelligence, scheduling,

24:46 built in observability, and finally secure networking.

24:49 So you get all of that in one Kubernetes native stack.

24:54 And moving on to Cyclecloud.

24:56 So what exactly is that?

24:58 Cyclecloud is Azure's enterprise friendly tool for orchestrating

25:01 and managing HPC and AI workloads in the cloud.

25:05 So it allows you to easily deploy,

25:08 manage and scale your HPC or AI clusters with the scheduler of your choice,

25:12 whether it's SLURM, PBS Pro, LSFHD, Condor, you name it.

25:17 And you can also provision infrastructure, attach storage, deploy your schedule,

25:20 like I said, and scale based on the demand you have.

25:23 So the beauty of this really is that it meets you where you are, right?

25:27 You don't need to rework your existing tool chains or work flows.

25:30 So in other words, Cycle Cloud gives you

25:32 that flexibility or flexibility without any of the friction.

25:37 And we also like to take it one step further.

25:40 So we have Azure Cycle Cloud Workspace for Slur or CCWS for short,

25:44 which builds on top of Cycle Cloud and it

25:47 provides a single pane for deployment management and observability.

25:51 So the big take away here is that it gives you instant cluster

25:54 provisioning so that you can deploy a slim cluster in minutes instead of weeks.

25:59 It's also built on enterprise ready architecture,

26:01 so it takes care of your secure virtual networking.

26:04 It gives you flexible storage options like NFS, net app files,

26:07 luster and it also gives you that unified operations and insights.

26:11 So you get end to end health monitoring, Grafana,

26:14 dashboards for cluster analytics all while being integrated

26:16 with Linux VDI and open on demand for seamless management.

26:20 So in short, it gives you, it takes away the complex cluster management that you

26:24 would have and makes it into a simple managed

26:27 experience from deployment to monitoring so that you can

26:29 get your cluster up and running just like BFL did.

26:34 And so now we have a quick demo.

26:36 I'll play it, I'll play like a maybe like 1/4 of it and just talk through.

26:41 So basically what we show in this video

26:44 is the HPC and AI conversion story, right?

26:47 So there's no audio on, but it's in cycle cloud workspace.

26:50 What we do is like you can see open

26:52 up and open on demand session, get into VS Code.

26:55 And what we do here is in that VS Code session, we're setting up an AI agent.

27:00 And what that AI agent does is it's able to build the slim script for open foam,

27:05 but then also be able to debug that.

27:08 And so throughout this demo,

27:09 what we're doing is we're defining the agent profile, right?

27:12 We're giving it a prompt on what it should exactly do, what are the parameters?

27:16 Then we run into an error.

27:17 So with that error, we tell it to debug that.

27:19 We give it the error file, tell it to fix it.

27:22 And then finally, after it fixes it, we, we let it submit the job.

27:27 So the analogy I like to think of in this HPCAI

27:30 convergence story is to think of the agent like a car.

27:34 So it's very powerful,

27:34 but you still need a driver's licence at the end of the day, right?

27:37 You need good prompts, you need good tools,

27:39 and you got to give it the proper context.

27:44 Now, another neat capability we have is interconnect groups or ICG.

27:49 You might have heard Sachem talk about these.

27:50 And what it does is it brings

27:53 RAC aware allocation to Azure's newest GPU systems.

27:56 So with the hardware like GB 2 hundreds and GB three hundreds,

27:59 the real unit of compute is an entire rack.

28:02 It's not just a single VM.

28:04 And now ICG ensures that your workloads land on the right

28:07 rack to unlock that full NV link and InfiniBand performance.

28:13 And if you have rack reservation, you can get visibility into your rack,

28:16 into the racks that your VMS are placed

28:18 on and you get health aware empty node queries.

28:22 And finally today we are excited to announce

28:25 the GA of Cyclecloud 8.8 and CCWS 1.2.

28:29 So in Cyclecloud 8.8, some of the big things we've added is Ubuntu

28:33 24 dot O four and enterprise Linux 9 support.

28:36 So we've basically expanded our OX flexibility across Red Hat,

28:40 OMA, Linux and Rocky.

28:42 And now next is one of the most important additions,

28:44 which is the node health agent.

28:46 Now this is important because in large AI and HPC clusters,

28:50 jobs can run for days and once they're

28:52 running you absolutely don't want to interrupt them.

28:55 And so the node health agent subscribes to these health

28:58 events and allows you to monitor them in real time.

29:02 It also brings in two different classes of checks

29:04 designed specifically for these large long running GPU workloads.

29:08 1 is these non invasive checks which run while your jobs are running.

29:12 They're very lightweight.

29:13 What they do is they monitor things like the GPU presence,

29:16 and if a problem is detected,

29:18 they'll automatically drain it without killing the active job.

29:21 The other one we have is invasive checks, right?

29:24 And so these invasive checks are deeper stress style

29:27 tests that only run when the node is idle.

29:30 They never run during a user's job.

29:33 What they do is that they validate the GPUs and interconnects,

29:36 make sure that they're fully healthy before the next job lands.

29:40 And so together these two give a customer

29:42 much more of a safer environment altogether,

29:45 especially for those long running AI jobs,

29:47 because it catches these hardware failures early and then

29:50 also protects the jobs that are already in flight.

29:54 We also have ARM 64 and HPV Vive support.

29:57 So for ARM 64 that includes the GB 200 and GB 300 architectures.

30:02 And we also have support for the HPV Vive

30:04 series for your next generation of AI and HPC systems.

30:08 And then finally, for the cycle cloud side, we have topology aware scheduling.

30:12 So the cycle cloud integration with Slurm ensures

30:15 that these distributed AI workloads align with the GPU cluster

30:18 and that includes the cluster back end topology so

30:20 that you can get the best training performance that you want.

30:26 Moving on to the right side.

30:27 So CCWS 1.21 of the biggest features

30:29 we have is the managed monitoring integration,

30:32 which visualises everything in an Azure managed Graphon

30:36 instance for a unified GPU and InfiniBand telemetry.

30:40 So it does this by installing the NVIDIA DCGM exporter and Slurm exporter,

30:45 which pushes end to end metrics into a self hosted Prometheus database.

30:50 And what this says is it visualises it in an Azure managed Grafondo.

30:53 So you get that job to GPU level, drill down that you want.

30:57 We also have availability zone support.

30:59 So what this does is it allows Co location of your GPU compute and storage

31:03 within the same availability zone so

31:05 that you can minimise your storage access latency.

31:09 Next, we have on Tri de authentication.

31:11 So that's supported right now with Cycle Cloud UI and Open on Demand.

31:16 And finally, we have the Open on Demand integration itself.

31:19 So for those of you who may not know,

31:22 Open on Demand gives you a simple browser based access to access shells,

31:27 files, interactive apps, VS Code like we saw on the demo.

31:32 And finally, we have the Linux VDI capability.

31:35 So we've done this and through a partnership with Sandy Withinlink,

31:37 they've built on top of an open source software.

31:40 And what that does is it enables interactive GPU

31:43 accelerated desktop workflows directly in the cluster environment itself.

31:51 And that is the end of the session.

31:53 If you enjoyed it and you want to learn

31:55 more about AI Foundry and Azure Kubernetes Service,

31:57 please feel free to check out these shessions on the screen.

32:01 So thank you all for coming and thank you BFL for being a valued customer.

Study with Looplines Download Captions Watch on YouTube