Azure AI Infra updates to power frontier and enterprise workloads | BRK179
Microsoft Events
0:00 Alright.
0:00 Hey folks, welcome.
0:02 Thanks for coming to Good Night.
0:05 My name is Matt Vegas.
0:06 I'm a principal Product Manager for the ND series.
0:09 This is our AI infrastructure virtual machine.
0:13 I'm here today with Saksham from Black Forest Labs
0:16 and Param who's on the software side of our AI infrastructure.
0:21 And today we're going to talk about our Azure
0:24 AI info updates that Power Frontier and enterprise workloads.
0:30 So why are we here?
0:32 You know, like, why has Microsoft for the last
0:35 several years led the way in AI infrastructure?
0:39 You know, why have we been building
0:41 these supercomputers long before other CSPS were interested,
0:44 you know, even long before ChatGPT existed?
0:48 These kinds of systems are really expensive
0:50 and complicated and they're tough to deal with.
0:53 And there's really two reasons that it starts with, you know,
0:58 Microsoft ultimately is a user of supercomputers,
1:01 whether it's GPU compute or CPU compute.
1:05 It's been a pretty integral part of our design process,
1:08 you know, whether it's hardware or silicon software.
1:11 The second one is, is, you know, we're in it now,
1:14 but this AI boom, you could have seen it coming.
1:16 It was not a question of when, but just it was not a question of if,
1:20 but like when was it actually going to happen.
1:23 And Microsoft has just this vast portfolio of software
1:27 products that add value to all of our daily lives.
1:31 And it's, I see it as the perfect
1:33 vehicle for delivering AI and delivering value through AI.
1:38 So this need for AI and really high end
1:42 infrastructure has always been intrinsic to what motivates us.
1:47 It's what drove us to build this competency.
1:49 And really it's now it's,
1:51 we get this wonderful privilege to kind of pass on the harder knowledge,
1:56 expertise, and really our infrastructure to our customers.
2:00 And so, you know, Microsoft, we're a provider of AI, but we're, you know,
2:03 at the heart of it is,
2:05 is we're really a user of all of these, of this really advanced infrastructure.
2:10 So and the, the need for this compute doesn't end with Microsoft.
2:15 Obviously the need for compute is
2:18 basically exponential across our customer base.
2:21 You know, over the last several years,
2:23 model size in terms of parameters has grown massively.
2:28 You know, each new model growing exponentially over the prior generation.
2:33 We've seen, you know, pure text models evolve into multi model.
2:37 Now we're seeing the rise of like agentic workflows,
2:40 long chain of thought these, you know,
2:43 like really complicated AI processes that are capable of doing,
2:47 you know, full workflows and deep research on behalf of the user.
2:52 And so it's, it's really advanced, you know,
2:55 as we've kind of followed this trajectory
2:58 and maintaining this path of innovation
3:00 for a customer is really what drives us to just build larger,
3:04 higher performance systems on the latest GPUs.
3:08 And so kind of Needless to say, we've been very,
3:11 very busy over the last few years just building out
3:14 a huge footprint of supercomputing infrastructure all around the world.
3:21 So I'm sure you guys read the news, it seems like everyday there's, you know,
3:26 a new massive GPU supercomputer being announced by one of the many
3:31 competitors out there that are building these types of things.
3:36 And I feel pretty strongly that Microsoft has always been,
3:40 you know, several steps ahead.
3:43 You know, like if you read the news recently, we announced Fairwater,
3:47 which is the world's most powerful data centre for AI training.
3:52 Fairwater is multiple whole regions of AI compute.
3:57 If you typically think of a region
3:59 as something that hosts all types of services,
4:01 but these are whole regions basically just designed to do this.
4:05 We're talking hundreds of thousands of GP us connected
4:08 in a massive network and all of this is real.
4:11 Like it's on the ground.
4:13 It's basically in production and it's not some future plan
4:17 on paper and an announcement like it's this is real.
4:22 Over the last few years, we've grown our worldwide fleet to support the daily
4:27 needs of inferencing that we all now use.
4:29 And it's like pivotal and our, our, our work,
4:32 whether it's copilot or things like ChatGPT.
4:36 An example of this is from mid 23 to 24, we grew our fleet by 30X.
4:42 Like so for less than a year, just the amount of supercomputers that we
4:45 host around the world has grown tremendously.
4:49 And all of this, you know,
4:51 basically comes back to just our key principles of building everything with HPC,
4:55 you know, high performance computing, it's top of mind.
4:59 We build custom platforms like fine-tuned for performance.
5:03 We deliver all of our products basically first to market,
5:06 so right when they're announced and released and we take,
5:09 you know these really advanced designs that are used
5:12 in the highest end supercomputers and we give everyone access to them.
5:16 Like all of these things are built the same
5:19 with the same like bar for quality and bar for performance.
5:27 Our investments on the product side span the entire
5:30 stack and this is what enables our customers to succeed.
5:34 You know, this year we introduced GB 200 GB 300 RTX 6000 Pro based VMS.
5:40 We continue to sell out our massive fleet of A1 hundreds
5:44 and H1 hundreds and AMD based MI300X based systems,
5:48 but it takes more than just, you know,
5:51 advanced GPUs to deliver a good cloud service.
5:54 We've invested heavily in CPU compute for pre processing.
5:58 Also 800 Gigabit XDR back end networking that was released with GB 300.
6:03 We have managed luster and object storage services,
6:07 cycle cloud and Kubernetes optimizations.
6:11 All of this is basically to ensure that, you know,
6:14 we have what we need to efficiently, that we have what we need for our customers
6:19 to efficiently support and run their AI applications.
6:26 We also continue to invest in the future of AI.
6:30 We're really excited this week at Ignite we're we've
6:35 launched the GA of ND GB 300 virtual machines.
6:39 These are for AI training and inferencing.
6:42 It's now generally available, which is a huge milestone for us.
6:45 We're also excited to announce Azure NCRTX Pro
6:49 6000 BSE is now available for public preview.
6:54 This is a cool platform.
6:55 It's a single platform for both AI and visualisation, visual compute, basically,
7:00 so that those two things can work together to accelerate enterprise workloads.
7:06 In addition, this week we're excited
7:09 to announce our continued partnership with NVIDIA
7:12 and we're that we're basically dedicated to building the future of AI
7:17 on Verirubin and we're also dedicated to our continuing our partnership with AMD
7:23 to offer more performance and flexibility
7:26 in the cloud with their upcoming MI40455X GPU.
7:32 So whether it's investments in our GPU systems
7:36 or our infrastructures stack or offerings of full stack AI solutions,
7:42 AI innovators choose to run on Azure.
7:47 This includes open AI copilot.
7:50 Just today Anthropic was announced to be a customer of ours.
7:55 Copilot runs on it Figure 8 byte dance meter,
7:58 open AI and Black Forest labs who are so,
8:01 you know, we're we're really excited to have just these phenomenal customers.
8:07 So a little back story before we kind of go to our next segment in this slide,
8:13 but 2025 has marked a pivotal shift in AI infrastructure.
8:17 It's really where we went from X86 based surfers to grace Blackwell,
8:23 which is a completely new architecture
8:26 that with advancements basically across across the design.
8:32 But you know, it has a arm processors,
8:34 it has a new envy link that connects the entire rack and enable
8:39 and basically enabling the whole rack to act as a massive supercomputer.
8:44 And so today with us, we have Black Forest Labs,
8:47 who are renowned for their ground breaking text to image generation models.
8:51 And we're one of the first people to get
8:55 access to a really large GB 200 cluster.
8:58 And so today I'm excited to have them here.
9:01 So thank you.
9:05 Hi everyone.
9:05 A bit of introduction about myself.
9:08 I'm Saksham, Consul and the lead infrastructure efforts at Black Forest Labs.
9:12 I joined about a year ago from a start up in the LLM training and serving space,
9:16 which data got acquired by AMD.
9:19 So I'm responsible for the training and inference infrastructure,
9:22 so to train and also server models at scale.
9:26 So a brief description of what Black Forest Labs is.
9:30 So Black Forest Labs is a leading AI research lab for visual intelligence.
9:34 We, we're not just building models in my opinion,
9:37 we're actually building foundation technologies that shift
9:40 how we see and understand the world.
9:43 So our latest offerings already are integrated
9:46 into previous domains and are built into production workflows,
9:49 not just in the creative domain,
9:50 but also in design software as also social media experiences.
9:54 And I truly believe that we are turning imagination
9:57 into reality for with three things that actually matter in production,
10:00 speed, precision as well as creative control.
10:04 So we launched into August 2024,
10:05 we were backed by Anderson Horowitz and General Catalyst.
10:09 And right now we have about 50 full time employees based in two headquarters,
10:13 one in Fryeburg, Germany as well as San Francisco.
10:17 So before BFL, our founding team pioneered
10:21 technologies that shifted the define generative AI,
10:24 things like VQ games, latent diffusion, stable diffusion.
10:28 These models fundamentally changed the the AI Gen.
10:32 AI landscape.
10:33 They made text to image generation accessible to millions
10:36 and set a new standard for what what's possible.
10:39 So in the last four years,
10:41 our work has been started over 100 and 5000 and 50,000 times.
10:46 So yeah, in terms of the founding team, there's Robin Rombach,
10:49 he's our CEO and he's leads the company's strategic
10:52 directions as well and they're doing building the company itself.
10:56 That's Andreas Blackman.
10:58 He leads a go to market.
11:00 He's involved for all our growth initiatives,
11:02 operations as well as partnerships.
11:04 And finally we have Patrick Asser, the man, the myth, the legend.
11:08 He leads a research at Black First Labs and he's
11:10 responsible for most of the cutting edge research that we do.
11:14 So in terms of like it's I'll be remiss to talk
11:17 about Black First Labs without talking about our flagship family.
11:21 So the, the Flux series of models are one
11:23 of the most advanced sets of image generation models out there.
11:28 Since launch, each and everyone of our models have topped
11:31 the benchmarks of both text to image and image editing benchmarks.
11:35 We the, our models are available both open weight,
11:38 so you could deploy this on premise and we also have
11:42 our high performance models available on our API and also on Azure Foundry.
11:48 So in our end, this gives us gives customers flexibility that you can
11:51 deploy them to match your security costs and integration needs that are there.
11:55 So in terms of the models that we have, we have the Flux 11,
11:58 which is like one of the high quality text image generation models out there.
12:02 And this is a workhouse that powers
12:04 a lot of production workflows out there today.
12:06 There's the ultra model, which is ultra fast and for ultra high res content.
12:11 So you can, if you want to get 2K images generated in seconds and not minutes,
12:15 that's the model you should go for.
12:18 And then we have the Flux 1 context model,
12:20 which is the state-of-the-art model in both
12:22 for in context editing as well as generation.
12:25 So this combines text and images
12:27 to generate outputs which are coherent and precise.
12:30 And finally, small teaser that we are releasing,
12:32 we have a new series of models coming out, which is the Flux 2,
12:36 which is setting a new standard for image editing and generation.
12:39 And this is this, this model will be something
12:41 that you can deploy to productions and not just for demos.
12:45 So like at BFL, we just don't talk about what we do, but we actually ship it.
12:50 So every image you see out there
12:52 on a slide right now have been generated by flux.
12:55 And feel free to scan the QR code and try out.
12:57 Or you can go to BFL dot AI slash play, generate the images yourself,
13:00 test the control and see where thousands of people use flux.
13:05 Thank you all.
13:09 Right.
13:09 So we're going to switch things up here.
13:10 We have a little bit of AQ and a that we're going to do
13:14 so so we can learn more about how BFL uses our GB 200 clusters.
13:18 So you know, Sacha, I'm working with large clusters
13:22 of GPUs is obviously not the same as, you know,
13:25 a general purpose compute infrastructure.
13:28 Walk us through your experience of onboarding an Azure
13:32 AI cluster and how you scaled to train Flex models.
13:37 Sure.
13:37 Yeah, you're absolutely right that like we were working with large scale GPUs,
13:42 it's completely different from like standard operations and infrastructure.
13:46 And it's not just on the, you know,
13:48 like when especially in training, it stresses the entire system to its limit.
13:52 So it's not just the compute plane that we normally talk about,
13:55 but also the networking, the memory and the and also the computer obviously.
14:00 So like with these challenges, like it's been really nice that Azure has
14:03 been a cloud cloud partner since day one.
14:06 So from the first customer,
14:07 have you been one of the first customers on the GB 200 NVL 70 twos?
14:11 And Azure has supported us throughout.
14:13 So in terms of training, one of the big steps is there is like data and you,
14:17 you know, when you're pre training,
14:19 there's so many processes that are touching the, there's so many
14:22 processes that are accessing the data and there's so there's so much,
14:26 the scale of data is such high that you
14:28 started the data infrastructure problem becomes pretty intense.
14:32 And if not handled properly,
14:33 you're going to have these weird issues like cache thrashing,
14:37 what else partition thresholds being met.
14:40 So in TLDR, like the performance gets really impacted.
14:44 So we had to work really closely with Azure,
14:46 the Azure support, the Azure Storage team, for example, to get to really improve
14:50 the performance and make these workflows work.
14:53 That's awesome.
14:54 So what are some other considerate consideration?
14:59 What are some other considerations you made throughout your experience on Azure?
15:03 You know, beyond just the GPU,
15:04 how did you get the infrastructure and services right?
15:08 You know how has Azure Infra basically helped you achieve your goals?
15:15 Right.
15:16 So yeah, in terms of like things apart from the compute,
15:19 I think that three things that I could talk about 1 is,
15:23 for example, like, you know, we talk about GB 200,
15:26 brand new, flashy new infrastructure available.
15:29 But it's not just the hardware.
15:30 There's, there's a lot of software that's
15:32 also involved and the Azure team and NVIDIA,
15:35 we worked closely together to to get new drivers,
15:37 validate on optimizations that we can do
15:39 on this new system to actually get the best performance.
15:43 So that's one.
15:45 Secondly, I would say networking and networking and storage.
15:50 So like pre training, like it's not it's intensive throughout the entire system.
15:54 So you know, you're you hit into these weird issues like IB,
15:58 IB flapping, you're reaching a storage thresholds, your switch are failing.
16:03 So, you know, working with the HPC and storage team,
16:05 that's really important to get, you know, to identify these common hot parts.
16:09 Yeah, one solution was like finding
16:11 these like preemptive load balancing, for example,
16:13 so that we can actually push through
16:15 the network without having to basically DDoS ourselves.
16:18 And lastly would be monitoring.
16:21 Yeah.
16:21 So when you're running pre training jobs,
16:23 these jobs are not just like running in seconds or minutes.
16:25 They run for days, in fact for months.
16:27 And then so it becomes anything can go wrong in this.
16:30 So identifying when it goes wrong,
16:32 what went wrong and actually fixing that as fast
16:35 as possible becomes essential for actually meeting our target velocity.
16:39 So that that's been something that, you know,
16:42 really helps with that has helped us out so far.
16:44 Oh.
16:45 That's very interesting.
16:46 And then how important is orchestration in managing a large cluster like this?
16:52 Can you talk a little bit about, you know,
16:54 how you landed on the stack that you're using right now, the benefits of Azure,
16:58 your experience running Slurm on Azure with Cycle
17:01 Cloud as kind of an orchestration engine?
17:04 Right.
17:05 I mean, yeah, without orchestration there is now point of a research
17:09 cluster and the orchestrator doesn't just do like, you know,
17:12 you're not just putting jobs and does the job allocation,
17:15 but also does a lot of optimizations
17:17 and getting the rewrite resources to run your jobs.
17:20 So to answer the question why we use slum,
17:23 like Slum is better tested for HPC workloads and with Cycloud's help,
17:27 we were, you had an AI, we had an Azure native solution,
17:30 which is really closely tied with Azure's infrastructure,
17:33 which helped us get to get to use the, for example,
17:37 the GB 200 rack topology system really quickly.
17:40 So traditionally, for example, Slum does not really understand what a rack is,
17:44 but for the GP2 hundreds, it becomes really important that you understand
17:48 what RAC is like that's really tight integration
17:50 of what what where the node is in the RAC and using the RAC accordingly.
17:54 So cyclic Cloud worked really closely with SCAD MD to design this RAC topology,
17:59 a block topology in cyclic cloud in slum itself,
18:02 which allows that we can allocate jobs
18:04 without having to change any other application layer.
18:08 So the jobs get you're doing RAC level job selection,
18:11 not node level job selection, which then really,
18:14 I mean, it really improves performance, yeah.
18:18 Now the, the pace of innovation in the GPU space has been crazy.
18:22 It's where like it was yesterday,
18:24 we're still installing A1 hundreds and now we've gone
18:29 through H100H200GB200GB300 and now we're talking about Vera Rubin.
18:34 I mean, it's it's just going so fast.
18:37 You know, how what was the experience like
18:40 collaborating with Azure to get kind of, you know,
18:43 first or early access to these GB 2 hundreds?
18:47 Also, in what ways did it help you having, you know,
18:50 access to just the absolute cutting edge, highest performing infrastructure?
18:56 Right.
18:56 I mean, it still blows my mind that when you talk about the GB 200 stack,
19:01 we are getting what 13 terabytes of VRAM and what
19:05 576 terabytes per second of high bandwidth memory and the like,
19:08 we get double the TT flops compressed, the H 100.
19:12 So you're right, the scale in which computer is moving is really,
19:15 really, really fast.
19:16 But and this really helps with training.
19:18 So training for both transformer models as well
19:21 as autoregressive transform transform models or diffusion models,
19:24 both of them are compute bound,
19:27 which means that as compute accelerates, training accelerates.
19:31 And things like higher band, higher memory,
19:33 for example, helps in like multiple fronts.
19:36 You're able to have able to put more weights
19:38 in, in a smaller amount of GPUs so they become more efficient.
19:41 That's less likely to break, much easier to work with.
19:45 You're still able to have higher batch sizes, which people normally think about,
19:48 but that helps with training performance quite a bit.
19:50 So those are a lot of things that really improves with like,
19:53 you know, having better infrastructure helps with training.
19:57 So the, the advantage is that working with Azure
20:00 and you guys have has helped that, you know, we've been such a small team,
20:04 we're able to leverage a lot of the support that you guys have
20:07 given us to actually use the cutting edge to, to, to, to win.
20:11 That's wonderful to hear.
20:12 And then for the audience out there who's maybe
20:15 get just getting started with Azure infrastructure, you know,
20:19 what are your like, say top three lessons learned, You know,
20:23 don't hold back the gory details if you need to.
20:27 All right.
20:29 So yeah, I mean this is I think relevant
20:31 in my opinion for all infrastructure, not just Azure.
20:36 But one thing that for me was some a really good
20:38 lesson was like learning how to Co design with your cloud provider.
20:42 So it's not just like finding issues which you will have and something like,
20:45 you know, complaining about it,
20:47 but actually sitting down with the team and discussing
20:49 on like what's on my memory access pattern,
20:52 what's my data access pattern and finding out
20:54 these, finding out these solutions that we can do.
20:57 So we were obsessed with collecting metrics and logs
20:59 and sending them to your sending them over,
21:01 not just to understand the system better,
21:04 but also to accelerate the improvement.
21:06 And to Azure's credit, you guys were really quick in getting those fixes in.
21:10 So I believe like, you know, working together in tandem really helped us there.
21:15 Secondly, I think what's really important is like, you know,
21:19 focus on the focus on the entire system and not just on the on the GPUs.
21:24 Like we always come like we always think that, you know,
21:26 training that's GPUs and that's about it.
21:27 But you're, you're, it's an entire system that's going to be stressed out,
21:31 networking, storage, memory, all of those things.
21:34 So thinking of in a holistic manner
21:36 and finding out where your actual actual bottlenecks are,
21:39 it's typically not sometimes a GPU, not always.
21:42 So like using ARM valves principle, finding it out, it's really useful.
21:45 And secondly, when you're working with like really high performance,
21:49 like when you're using something new like bleeding edge,
21:52 you put put in some time or allocate some time
21:55 for actually that there might be some engineering effort required.
22:00 So you're putting some work to get the work out to be to be honest.
22:03 So like if you're getting 2X computer, computer,
22:07 computer improvement thing some for engineering development and time.
22:12 But yeah, those are three things I will think about, yeah.
22:15 That's wonderful.
22:17 Thank you so much for joining us.
22:19 It's been an honor having you here on stage with us.
22:22 You're phenomenal customer.
22:23 Thank you.
22:24 Thank you all.
22:25 Right.
22:25 So we're going to kind of switch things up a little bit now.
22:27 Sorry, Param Shah is going to present on different side
22:32 of our stack which is the software side of our infrastructure.
22:36 Yeah, awesome.
22:37 Thank you, Matt.
22:37 Thanks, Saksham.
22:39 So now that we've seen the hardware side of Azure's AI infrastructure,
22:42 let's shift gears to the software side and orchestration,
22:46 because that's where the GP power really comes to life, right?
22:50 Because at the end of the day, you can have all the latest and greatest GPU's,
22:54 but it's also about how you can deploy,
22:57 manage and scale them for your workloads.
23:02 So across Azure's AI software stack,
23:04 we've designed our offerings around how you build AI.
23:08 And at a high level there's two main paths, right?
23:11 If you want to start with pre trained models, you should go to Azure AI Foundry.
23:15 Right?
23:16 Here you can focus on fine tuning and agent creation.
23:20 If you want to train your own custom models,
23:22 you can use Ask if you're already on Kubernetes or Cycle Cloud if you
23:26 want to bring in your own scheduler
23:27 and get your training infrastructure up quickly.
23:30 So on the left, as you can see, you get a managed experience right where you can
23:34 build use built in models and tools for fine tuning.
23:38 And on the middle, you can see the Kubernetes our ask service,
23:41 where if you have containerized pipelines with DevOps integrations,
23:45 then that's your natural fit.
23:48 At the end of this presentation,
23:49 I've also linked some fun ask and AI Foundry sessions
23:52 that you guys should feel free to go look at.
23:55 And at the very right went too far.
23:59 Let me go 1 slide back.
24:00 Cool.
24:04 Yeah.
24:04 So on the right side, if you look, we have Cycle Cloud where if you're bringing
24:09 in your own scheduler or workflow manager like Slurm
24:11 or PBS Pro and you want that fine
24:14 grain control over your distributed AI or HPC workloads,
24:17 then Cycle Cloud is built for that.
24:19 So today we're going to focus on that right side, right,
24:22 the customizable side of the spectrum where teams can orchestrate compute
24:25 at massive scale with end to end observability and deep insights.
24:31 So before we get into that, it's going to give
24:33 a high level overview of what ask is Azure Kubernetes Service.
24:36 So that is our fully managed Kubernetes platform for AI.
24:40 So it provides GPU optimized infrastructure, it gives you dev friendly tooling,
24:44 you get intelligence, scheduling,
24:46 built in observability, and finally secure networking.
24:49 So you get all of that in one Kubernetes native stack.
24:54 And moving on to Cyclecloud.
24:56 So what exactly is that?
24:58 Cyclecloud is Azure's enterprise friendly tool for orchestrating
25:01 and managing HPC and AI workloads in the cloud.
25:05 So it allows you to easily deploy,
25:08 manage and scale your HPC or AI clusters with the scheduler of your choice,
25:12 whether it's SLURM, PBS Pro, LSFHD, Condor, you name it.
25:17 And you can also provision infrastructure, attach storage, deploy your schedule,
25:20 like I said, and scale based on the demand you have.
25:23 So the beauty of this really is that it meets you where you are, right?
25:27 You don't need to rework your existing tool chains or work flows.
25:30 So in other words, Cycle Cloud gives you
25:32 that flexibility or flexibility without any of the friction.
25:37 And we also like to take it one step further.
25:40 So we have Azure Cycle Cloud Workspace for Slur or CCWS for short,
25:44 which builds on top of Cycle Cloud and it
25:47 provides a single pane for deployment management and observability.
25:51 So the big take away here is that it gives you instant cluster
25:54 provisioning so that you can deploy a slim cluster in minutes instead of weeks.
25:59 It's also built on enterprise ready architecture,
26:01 so it takes care of your secure virtual networking.
26:04 It gives you flexible storage options like NFS, net app files,
26:07 luster and it also gives you that unified operations and insights.
26:11 So you get end to end health monitoring, Grafana,
26:14 dashboards for cluster analytics all while being integrated
26:16 with Linux VDI and open on demand for seamless management.
26:20 So in short, it gives you, it takes away the complex cluster management that you
26:24 would have and makes it into a simple managed
26:27 experience from deployment to monitoring so that you can
26:29 get your cluster up and running just like BFL did.
26:34 And so now we have a quick demo.
26:36 I'll play it, I'll play like a maybe like 1/4 of it and just talk through.
26:41 So basically what we show in this video
26:44 is the HPC and AI conversion story, right?
26:47 So there's no audio on, but it's in cycle cloud workspace.
26:50 What we do is like you can see open
26:52 up and open on demand session, get into VS Code.
26:55 And what we do here is in that VS Code session, we're setting up an AI agent.
27:00 And what that AI agent does is it's able to build the slim script for open foam,
27:05 but then also be able to debug that.
27:08 And so throughout this demo,
27:09 what we're doing is we're defining the agent profile, right?
27:12 We're giving it a prompt on what it should exactly do, what are the parameters?
27:16 Then we run into an error.
27:17 So with that error, we tell it to debug that.
27:19 We give it the error file, tell it to fix it.
27:22 And then finally, after it fixes it, we, we let it submit the job.
27:27 So the analogy I like to think of in this HPCAI
27:30 convergence story is to think of the agent like a car.
27:34 So it's very powerful,
27:34 but you still need a driver's licence at the end of the day, right?
27:37 You need good prompts, you need good tools,
27:39 and you got to give it the proper context.
27:44 Now, another neat capability we have is interconnect groups or ICG.
27:49 You might have heard Sachem talk about these.
27:50 And what it does is it brings
27:53 RAC aware allocation to Azure's newest GPU systems.
27:56 So with the hardware like GB 2 hundreds and GB three hundreds,
27:59 the real unit of compute is an entire rack.
28:02 It's not just a single VM.
28:04 And now ICG ensures that your workloads land on the right
28:07 rack to unlock that full NV link and InfiniBand performance.
28:13 And if you have rack reservation, you can get visibility into your rack,
28:16 into the racks that your VMS are placed
28:18 on and you get health aware empty node queries.
28:22 And finally today we are excited to announce
28:25 the GA of Cyclecloud 8.8 and CCWS 1.2.
28:29 So in Cyclecloud 8.8, some of the big things we've added is Ubuntu
28:33 24 dot O four and enterprise Linux 9 support.
28:36 So we've basically expanded our OX flexibility across Red Hat,
28:40 OMA, Linux and Rocky.
28:42 And now next is one of the most important additions,
28:44 which is the node health agent.
28:46 Now this is important because in large AI and HPC clusters,
28:50 jobs can run for days and once they're
28:52 running you absolutely don't want to interrupt them.
28:55 And so the node health agent subscribes to these health
28:58 events and allows you to monitor them in real time.
29:02 It also brings in two different classes of checks
29:04 designed specifically for these large long running GPU workloads.
29:08 1 is these non invasive checks which run while your jobs are running.
29:12 They're very lightweight.
29:13 What they do is they monitor things like the GPU presence,
29:16 and if a problem is detected,
29:18 they'll automatically drain it without killing the active job.
29:21 The other one we have is invasive checks, right?
29:24 And so these invasive checks are deeper stress style
29:27 tests that only run when the node is idle.
29:30 They never run during a user's job.
29:33 What they do is that they validate the GPUs and interconnects,
29:36 make sure that they're fully healthy before the next job lands.
29:40 And so together these two give a customer
29:42 much more of a safer environment altogether,
29:45 especially for those long running AI jobs,
29:47 because it catches these hardware failures early and then
29:50 also protects the jobs that are already in flight.
29:54 We also have ARM 64 and HPV Vive support.
29:57 So for ARM 64 that includes the GB 200 and GB 300 architectures.
30:02 And we also have support for the HPV Vive
30:04 series for your next generation of AI and HPC systems.
30:08 And then finally, for the cycle cloud side, we have topology aware scheduling.
30:12 So the cycle cloud integration with Slurm ensures
30:15 that these distributed AI workloads align with the GPU cluster
30:18 and that includes the cluster back end topology so
30:20 that you can get the best training performance that you want.
30:26 Moving on to the right side.
30:27 So CCWS 1.21 of the biggest features
30:29 we have is the managed monitoring integration,
30:32 which visualises everything in an Azure managed Graphon
30:36 instance for a unified GPU and InfiniBand telemetry.
30:40 So it does this by installing the NVIDIA DCGM exporter and Slurm exporter,
30:45 which pushes end to end metrics into a self hosted Prometheus database.
30:50 And what this says is it visualises it in an Azure managed Grafondo.
30:53 So you get that job to GPU level, drill down that you want.
30:57 We also have availability zone support.
30:59 So what this does is it allows Co location of your GPU compute and storage
31:03 within the same availability zone so
31:05 that you can minimise your storage access latency.
31:09 Next, we have on Tri de authentication.
31:11 So that's supported right now with Cycle Cloud UI and Open on Demand.
31:16 And finally, we have the Open on Demand integration itself.
31:19 So for those of you who may not know,
31:22 Open on Demand gives you a simple browser based access to access shells,
31:27 files, interactive apps, VS Code like we saw on the demo.
31:32 And finally, we have the Linux VDI capability.
31:35 So we've done this and through a partnership with Sandy Withinlink,
31:37 they've built on top of an open source software.
31:40 And what that does is it enables interactive GPU
31:43 accelerated desktop workflows directly in the cluster environment itself.
31:51 And that is the end of the session.
31:53 If you enjoyed it and you want to learn
31:55 more about AI Foundry and Azure Kubernetes Service,
31:57 please feel free to check out these shessions on the screen.
32:01 So thank you all for coming and thank you BFL for being a valued customer.