Azure Cosmos DB Conf 2026 | Live Stream

Azure Cosmos DB Conf 2026 | Live Stream

Microsoft Developer

0:03 Hey everybody.

0:04 Welcome to azure Cosmos DB conf.

0:06 Hello.

0:07 Hello.

0:08 Hi, welcome.

0:09 We've got so much great stuff to show you.

0:11 Hi.

0:12 Hi.

0:12 Hello everyone.

0:13 Hey.

0:14 Hello, hello, hello.

0:16 Welcome.

0:16 Welcome to azure Cosmos DB conf 2026.

0:30 All right.

0:30 Hello everybod.

0:32 Welcome to Azure Cosmos DB Conf 2026,

0:34 a virtual event for developers building modern apps with Azure Cosmos DB.

0:39 I'm Patty Chow, product manager for DocumentDB

0:42 and I'm joined by my amazing co host,

0:45 Jay Gordon, Senior Program Manager for Azure Cosmos DB.

0:48 You know what today is all about?

0:50 Practical guidance that you can use right away.

0:53 So we're talking about, architecture patterns,

0:56 best practices and real stories from teams building

1:00 production systems with Azure Cosmos DB, DocumentDB and AI.

1:05 Yeah, that's right.

1:06 You'll see a mix of deep technical sessions,

1:08 quick interviews and demos and this will

1:11 cover everything from getting started to performance,

1:14 scaling and building AI powered experiences on Cosmos DB and DocumentDB.

1:20 And you know what, whether you're joining us live,

1:23 and I, wanna say thank you to our live audience.

1:25 I'm gonna give you all thank you.

1:29 Thank you.

1:29 Whether you're joining us live or watch Demand later, I know I will be.

1:34 We're excited to have you as part of the Azure Cosmos DB community.

1:38 And without our community, we can't have an event like this.

1:43 But before we get started, let's get the big hand emoji.

1:47 We want to say a big thank you to amd.

1:49 Yeah.

1:50 AMD is sponsoring Azure Cosmos DB Conf.

1:53 This year and their partnership has helped us significantly expand the show.

1:57 You know what that means, Patty?

1:58 What does that mean?

1:59 Well, you know what that means?

2:01 More sessions, deeper technical content,

2:03 which I know is important to all of you.

2:06 And, we'll have a broader set of speakers

2:08 and topics that we'll be covering today.

2:10 Yeah.

2:11 And this whole expanded programming really

2:13 wouldn't be possible without their support.

2:15 So thank you, amd.

2:17 Yes, thank you so much.

2:18 We really appreciate it.

2:19 Your support has just really made this into something huge.

2:24 I'm loving just seeing the LED wall.

2:26 I love seeing all these people here.

2:28 Thank you so much, amd.

2:29 Yeah.

2:34 Now what you're seeing live today is just part of that experience,

2:38 you know, that's right.

2:39 You know what, Patty?

2:40 There is a full library of sessions available live today and on demand.

2:44 You'll see them all go live around 1:30pm, Pacific.

2:49 And they're going to be covering everything

2:51 from AI architectures to design patterns, performance tuning.

2:55 We all need to do that.

2:56 And real world cases.

2:58 Yes.

2:59 So let's say that we're presenting here and if we can't catch everything today,

3:03 or you can't catch everything today,

3:05 you can go deeper and explore it all at your own place, at your own pace.

3:09 And you know what else we're doing, Patty?

3:11 And this is a big one.

3:12 We're also running the Azure Cosmos DB Conf.

3:15 Cloud Skills Challenge.

3:17 This is something.

3:18 You know what, Patty, I know you're excited about.

3:20 Yeah, I'm really excited about this, Jay,

3:22 because it is designed to help you build with hands

3:25 on skills using Azure Cosmos DB and it runs through May 8th.

3:29 Sounds put in your calendar.

3:30 Yeah, yeah, absolutely.

3:32 And I just hope that everybody who is watching you take advantage

3:36 of it because the first 500 people to complete the modules will get.

3:40 And here we go.

3:41 Patty, this is my favorite part.

3:42 What's your favorite part?

3:43 It is a, 100% free voucher to take the DP-420 exam.

3:50 So you know what?

3:51 Get certified.

3:51 Get certified for free.

3:54 Get learning.

3:55 Yeah, you know, I love free too.

3:57 And if you want to get a head start, visit AK Mscosmos CB CompChallenge.

4:03 That's right.

4:03 And like I said, our community,

4:05 the huge part of what we do to make sure that you're all involved,

4:09 you're all included and you know, you can take part in this event.

4:13 Yeah.

4:13 I mean, yeah, truly.

4:15 I think one of the things that makes Cosmos DB and Cosmos DB Conf.

4:19 So special is the people behind it.

4:22 And that's what makes it so special.

4:23 Like you all in our audience, thank you for being part.

4:29 And so we want to thank our incredible speakers,

4:32 our customers and partners who are sharing real world experiences today.

4:37 And that includes teams from OpenAI, Vercel, Walmart, AMD,

4:42 of course, the ODP Group, Southworks, Next, and many, many more.

4:48 And you may have gone on one of these sites today.

4:50 You might have bought something.

4:52 They're part of everyday life.

4:54 And these teams that we'll be,

4:56 sharing today that are building production systems at scale.

4:59 Yeah.

5:00 Today they will be sharing what's been working for them.

5:04 Sure.

5:04 And if you want to explore everything today that's happening,

5:07 you can head over to aka Ms.

5:11 Azurecosmos dbconf.

5:13 Yeah.

5:13 Just a quick note.

5:15 The YouTube description has all of today's links in one place.

5:18 There are over 30 sessions total,

5:21 with every session going live before the end of today's show.

5:24 Plus, if you're in the YouTube chat,

5:25 please tell us where you're watching from and ask questions as we go.

5:29 Sorry, I was working at my own pace.

5:30 No, it's okay.

5:31 We want to hear from you.

5:32 We want to hear from you too.

5:34 Well, you know what?

5:34 I think we've set you all up and it's now to get you to our, opening keynote.

5:39 Yeah, that's right.

5:40 Here's Kirill Gavrylyuk, Vice President of Azure Cosmos DB at Microsoft.

5:45 Developers.

5:45 Developers.

5:46 Developers.

5:48 Welcome to Azure Cosmos DB Conference 2026.

5:52 We are really grateful to have you here.

5:54 And let's take a look back at the year since

5:57 the last conference and what a world year has it been.

6:02 Of course we're all building AI apps at amazing pace, right?

6:07 And for Cosmos DB perspective, we invested heavily in semantic search.

6:11 That is core part of any AI app you have now.

6:14 Full text search, hybrid search, vector search,

6:17 semantic data, ranker, plethora of capabilities.

6:20 With confidence we can say if you're using Cosmos DB,

6:22 you do not need a separate search system,

6:24 just use built in capabilities and take advantage of cost efficiency,

6:29 performance, scale reliability and precision.

6:33 And of course more than half of our customers

6:36 are using coding agents today to build apps on Cosmos.

6:40 And for that we offered great agent kits

6:44 with skills to make your coding agent world class expert.

6:48 We added MCP servers and other capabilities to make this possible.

6:53 Now, AI does not remove the need for reliability,

6:57 security, performance and that is our priority number one.

7:01 This is where we invest the most of our energy.

7:04 We've provided throughput management,

7:06 fleet management for customers with large fleets of Cosmos DB.

7:10 We've invested with simplified partitioning with global security indexes,

7:14 hierarchical partition keys.

7:16 We've invested in per partition auto failover.

7:19 And now Cosmos DB is the only database that offers you 5,

7:22 9 reliability for any consistency mode.

7:25 And of course security is number one priority

7:28 and a number of improvements in security space.

7:32 Last year we've added second NoSQL database to Azure DocumentDB.

7:38 Unlike Cosmos, DocumentDB is fully open source.

7:41 It's based on an open source project at Linux foundation where Microsoft,

7:45 Amazon, Google and 12 other vendors are collaborating.

7:49 DocumentDB has a number of advantages over other document databases

7:54 and notably more than 40% cheaper than any other document database,

7:59 including MongoDB Atlas on AWS.

8:01 So please come over to Azure, use DocumentDB,

8:04 take advantage of performance and low cost.

8:08 Now AI transformation is on top of everyone's mind, right?

8:12 And for databases really means, in my opinion three things.

8:16 One is AI increase the importance of flexible data.

8:22 The data AI processes and emits is all semi structured, right?

8:26 It's prompts, it's memory, context, that's what AI is focused on.

8:30 And Cosmos DB provides the first class support for semi structured data.

8:37 It proudly presents flexible data model.

8:41 Now, AI accelerated the pace at which we develop applications.

8:46 Flexible data is paramount for you to enable the space.

8:50 Developers cannot be locked down by strict schemas,

8:53 they need schema flexibility which is again number one capability of Cosmos DB.

8:58 That's why we built a database is

8:59 to provide you schema less flexibility in data modeling.

9:03 Now AI brought semantic search as a first class operator

9:09 into queries and it's true for any database that matters.

9:13 And of course in Cosmos DB we invested far and beyond of that.

9:17 We added semantic search,

9:19 full text search vectors, hybrids, semantic grid anchor.

9:23 We invested in GraphRAG and we're about to announce

9:26 agentic retrieval which takes a graph rack to the next

9:29 level where the graph knowledge graph can be inferred

9:32 by LLM for you without you having to manually construct it.

9:36 Now the third important thing about AI transformation is

9:39 we all now are accompanied by our coding agents.

9:44 That's our favorite friends now, right?

9:46 Coding agents speed up development and for coding agents

9:50 to be efficient world class experts in databases, they need skills.

9:54 And that's why we invested in Cosmos DB Agent Kit.

9:57 We'll talk about it later.

9:59 That provides skills to your coding agent no matter which coding agent you use.

10:04 Along with MCP servers plugins for any favorite coding agents,

10:09 that you may choose.

10:11 Now let me introduce our guest speaker, Guillermo Rauch,

10:18 founder and CEO of Vercel, a popular AI app development cloud.

10:23 And Guillermo Rauch knows a thing or two about what AI apps need from databases.

10:30 In order, to have with us CEO and founder of Vercel, Guillermo Rauch.

10:37 Guillermo, you have built libraries like Socket.IO and Next.js

10:42 that form the modern web as we know it today.

10:45 On top of it, you've built a highly successful company, Vercel,

10:49 that has seen a really good growth of developers, that provides developer cloud.

10:57 What do you credit Vercel?

10:59 Success.

11:01 I think if I look back on the principles that we

11:04 created Vercel on, it was really about the ease of use.

11:08 At the time it was ease of use for humans.

11:10 Now lately I've been calling it agent ergonomics or ease of use for agents.

11:15 It's about, you know,

11:16 you have an idea and you need to bring it online and you actually

11:20 want it to scale and work really well and be a really high quality product.

11:25 So if, if the, if it was, if it could be summarized as something simple,

11:28 it would be the focus of ease of use,

11:31 but in the service of quality, not just, you know, getting slop as we call it,

11:36 this taste online that crashes and pages you at midnight makes sense.

11:41 As you transition to agents as your customers.

11:45 What, surprised you is what's happening in the industry.

11:49 It's really the scale.

11:53 As an entrepreneur, as a founder,

11:55 you start doing market sizing estimates when you're preparing

11:59 your early decks for your investors and would be investors.

12:04 The common wisdom was that even being generous,

12:08 Maybe there were 30 million people that could deploy applications because

12:12 we targeted the easiest programming language that we could think of, JavaScript,

12:17 TypeScript, trying to make the cloud super accessible.

12:20 But it was still tens of millions of people.

12:23 And Next.js also lowered a barrier to entry to react into this universe.

12:29 But when you think about agents,

12:31 they're empowering everyone on the planet to create software.

12:35 They might not even be realizing that they're creating software.

12:38 They might just ask for a solution to a problem.

12:41 And in the process, the agent writes an application and deploys it on our cloud.

12:47 So a lot of the growth that we're seeing these days is literally,

12:50 by the way, happened yesterday.

12:51 A high school friend of mine came to visit me.

12:53 He'd never written a line of code in his life and he's like, oh, by the way,

12:57 I know what Vercel does now because Claude code deployed

13:01 a Vercel and told me the application was now online.

13:05 So I could have never imagined that we would

13:08 see an increase in, let's call it a hundred times,

13:12 the addressable market of creators of software.

13:16 Maybe even higher.

13:18 I think we'll be soon in the future where, software is like water,

13:23 like it's just a normal thing to create, consume, sell, and it's everywhere.

13:30 That's amazing.

13:32 And as the platform that is, so frequently used by the agents,

13:36 what do you expect of the modern database?

13:40 How can we help you?

13:42 Well, I'm very thankful.

13:43 I always tell the team, because part of growing up as a company is

13:46 telling the stories and scars of the early days.

13:50 We started out using a database that was very managed by me in the early team.

13:56 And it was a horrifying experience.

13:58 And I reached out to you and I'm very thankful,

14:02 Carol, through an introduction from Nat Friedman,

14:04 because I was looking for a solution that would help

14:07 me focus on growing Vercel and not being a dba.

14:11 We had very few people on staff.

14:13 This is literally the garage story and one of the convictions that I had

14:18 that helped me choose Cosmos and be savvy about this because I'll tell you,

14:23 some of the less familiar developers look at Query languages.

14:27 And they might have a little bit of culture shock sometimes,

14:30 but one of the things that convinced me is that compute was becoming serverless.

14:35 The value of Vercel was you deploy it scales to zero,

14:40 it scales to billions of people.

14:42 I literally am supporting applications these days that were vibe

14:45 coded and go viral all over the Internet and have traffic

14:49 that even teams of engineers would have never dreamed of and can

14:53 now be sustained by someone that barely knows how to prompt.

14:57 And the magic there is that it was that bet on serverless.

15:00 Because I also tell you the reality is

15:02 that a lot of this software is quite ephemeral,

15:05 that's being written by the agents.

15:07 Maybe it's useful for a day, maybe it's useful for a presentation.

15:10 We're seeing a lot of sales engineers, including people at Microsoft,

15:13 use v0 to vibe code an app for a sales pitch,

15:18 for a demo for pre and post sales to demonstrate an integration of a product.

15:23 And so if software is ephemeral, serverless,

15:26 scale to zero and flexibility there is extremely useful.

15:29 So I'm very thankful that we made the right bet.

15:32 The other thing that I wanted to be super mindful

15:35 of that serverless compute helps a lot with is fault isolation.

15:41 So I didn't want to get a machine hot that could create noisy neighbor problems.

15:46 Imagine if the misbehavior of one tenant could impact the reliability,

15:51 uptime, availability of other of other customers.

15:54 But what I just described is the everyday database.

15:57 It's kind of nuts that we still.

15:59 There's so many people that kind of live in the shadows,

16:01 I think because they put systems online whose behavior they cannot predict,

16:07 they might become a wave of queries or users or whatnot that just ruins

16:12 the day for everybody and people start citing slow query logs and all that junk.

16:16 And so I wanted a system that gave me an economical

16:22 thinking where the developer writes a query and they understand its cost.

16:26 We can project out its cost.

16:29 A lot of what I think has made Cosmos successful is

16:31 that it kind of makes sense in the token world, right?

16:35 Like when I use a coding agent,

16:38 I know that certain traces of reasoning cost more compute.

16:42 When I make a cross partition query in Cosmos,

16:44 I get back an RU count of like how much I spent on that query.

16:48 And so it allowed me to scale Vercel from a few

16:52 developers that didn't want to have a DBA on staff

16:55 and to actually hundreds of engineers that are supporting

17:00 a growth in usage that would have never imagined with deployments,

17:06 which is one of the key collections that we Host on Cosmos Having grown,

17:10 our weekly Deployment count is 3x'd in a few months.

17:14 And I think there's probably no database system

17:17 in the world that I could have homegrown,

17:20 for example, that could have absorbed that kind

17:23 of growth that agents have gotten us.

17:27 Thank you.

17:27 Well, we're very grateful for you to be our customer.

17:31 We have many developers online right now watching this that are early in career.

17:37 What would be your advice to them?

17:41 Well, one of the things is that the agent is the new computer and this is

17:46 advice that I give myself every day is that I really have to update my priors.

17:51 Every day I give a presentation to the whole company saying look,

17:56 these are the things that I used to believe AI couldn't do.

17:59 And I was wrong.

18:00 I'm coming out and saying it like I

18:03 thought AI couldn't contribute to large code bases.

18:07 I was wrong.

18:08 I thought AI couldn't have, couldn't do good designs.

18:11 I was wrong.

18:12 And so it's very important to realize that the things

18:17 that you've learned you have to hold on to.

18:20 Of course you have to appreciate the skills that you've developed and be proud.

18:24 But also just be very light on your feet is my advice.

18:29 I like to think of the engineer of the future.

18:33 I'm someone that is very full stack in their understanding of the business.

18:38 When I hire an engineer I don't expect them

18:40 to just be typing code or even prompts for that matter.

18:43 I wanted to understand the full problem.

18:47 Again it's a little bit like I mentioned that economy system that Cosmos

18:51 has given us where I was able to shift the burden of understanding,

18:55 you know, what our database performance and cost

18:58 is going to be to the everyday developer.

19:01 They didn't toss the problem out to.

19:03 Oh, some other team will worry about scaling my database, right?

19:06 No, you are participating in that process.

19:08 I think the original ethos of DevOps was very much in that direction.

19:12 I think the difference now is that everyone is super full stack.

19:16 I do design, I generate videos, I generate SVGs,

19:22 I create front end code bases, I create backend code bases, I write clis.

19:27 During the holiday break I vibe coded my own Swift menu bar app.

19:32 And so I think the antidote here or like the way to really

19:38 stay relevant is to be super open minded and super full stack.

19:45 Thank you Guillermo.

19:46 Really grateful for us to have you here.

19:49 Appreciate everything what you're doing, love your posts,

19:53 love everything that vercel companies likewise.

19:55 It's been a great Journey and excited to continue collaborating

19:58 with you all and do more with Vercel in the future.

20:12 Thank you Guillermo.

20:13 Now Guillermo just covered few of the reasons why they chose Azure Cosmos DB.

20:17 But as these are the same reasons why thousands of other

20:20 companies choose Cosmos as a database for their agentic apps.

20:24 Its schema free flexibility enables your team to run fast iterate fast.

20:29 It's search built in into your query engine so that you

20:33 don't have to deploy different systems for search for database et cetera.

20:38 It's reliability, low latency and serverless

20:42 elasticity to enable your software to spike,

20:45 to enable your app to grow with demand without you touching anything.

20:50 Of course it's 5, 9 reliability that makes your sleep sound

20:56 while your software runs and serves millions and billions of people.

21:01 Now when it comes to AI use cases we have many

21:05 but the common ones with Cosmos DB of course include vector database.

21:09 Right.

21:09 That's why we invested in vectors, multimodal vectors,

21:12 the speed and cost effectiveness of vector search, semantic retrieval,

21:18 billing on top of vectors offers you

21:20 ability to do full text search, hybrid search,

21:23 precise increased precision of your retrieval,

21:26 taking advantage of the knowledge graph of your system.

21:31 But a recent trend is new use case that tripled

21:35 in size over the last six months is agentic memories.

21:39 We have more and more systems choosing Cosmos DB as agentic

21:43 memory for their apps and there is a reason for it.

21:46 Memory is semi structured, memory is dynamic,

21:50 memory needs to be reconciled over time and all

21:54 of these capabilities are built into Cosmos so you

21:57 can focus on what you app needs to do

21:59 and less about the infrastructure that it runs on.

22:05 Now let's get a couple words about the search in Cosmos DB.

22:10 The vector search is based on DiskANN algorithm built by Microsoft Research.

22:15 And the two key things about it is one, it's extremely cost efficient compared

22:19 to the traditional separate standalone search systems.

22:23 It's up to 50 times cheaper as you can see on the charts.

22:28 It is also low latency and scales no matter the scale of your app.

22:32 It provides you the same millisecond latency

22:35 for your vector search and full text search queries.

22:39 It scales well from few vectors to billions of vectors in your data set.

22:45 Now when you think about an AI app built on Cosmos,

22:48 of course OpenAI ChatGPT comes to mind.

22:51 That is well known.

22:53 It's a tremendous one of a kind application.

22:56 It executes more than 1.4 trillion transactions daily against Cosmos DB.

23:02 It stores more than 45 petabytes in Cosmos by now and at this scale,

23:07 it's the fastest growing app on the planet.

23:08 It still grows tenfold every year.

23:12 And to dive deeper into what enables OpenAI

23:15 to scale that fast and maintains a trap, it's a pace of innovation.

23:23 Let me invite John Lee, staff engineer from OpenAI,

23:26 to share a few thoughts with us.

23:30 Well, hi John, it's really great to have you here.

23:33 Would you mind introducing yourself?

23:36 Yeah, of course, it's good to be here.

23:38 My name's John, I'm on the online storage team at OpenAI.

23:42 Awesome.

23:43 Well, OpenAI is a unique company in many ways.

23:47 Right.

23:47 But one of the things that everyone knows about it is

23:50 that, and it's really visible is that it moves really fast.

23:54 You guys are doing many things concurrently and very fast.

23:59 What does OpenAI look for from databases

24:02 to enable users fast pace of innovation?

24:06 Yeah, that's a great question.

24:08 I think there are lots of standard database things that we care about, at scale.

24:13 So predictability, being able to scale to petabyte sizes,

24:16 being very flexible with how our use cases can be built,

24:21 and then of course observability is very important.

24:24 But I think from the scale perspective and being able to move really fast,

24:29 the most important thing here is being able

24:34 to scale from zero to millions of QPS,

24:37 being able to scale from zero bytes to petabytes,

24:39 features are launched on the regular, and they go immediately from no usage

24:45 to being used by hundreds of millions of people, on a daily basis kind of thing.

24:49 Right.

24:49 You know, so the way that you might design a normal database,

24:54 and then scale it out as your use case grows,

24:56 you know, that just doesn't work at this kind of velocity.

25:01 And I think we have thousands of developers that are actively building products.

25:05 So with Codex now things are iterating even faster.

25:09 It's really important to make it easy to onboard to databases really fast.

25:13 So we have thousands of tables, that we have to back by Cosmos DB.

25:20 And we have a system in front of Cosmos that helps

25:24 you be able to do sort of schema less design,

25:28 and allow these product developers to onboard very quickly.

25:33 Amazing.

25:35 Could you share any of the interesting patterns

25:37 that you guys have done on top of Cosmos DB?

25:40 Yeah, I think predictability is really key for how we use Cosmos DB.

25:46 We have a very well defined API that we have and expose

25:52 and all of these clients they build on top of this.

25:56 And that allows us to keep the system stable.

26:01 We don't want Every client to be writing their own Cosmos DB queries,

26:04 creating their own Cosmos DB accounts, creating their own tables.

26:08 So we try to build a multi tenant system on top of Cosmos, interestingly,

26:12 sort of at this global scale where we have

26:15 a lot of Cosmos accounts in a lot of regions, region failure is a big problem.

26:20 Right?

26:21 It doesn't happen often, but with dozens of regions it will happen.

26:26 And so we definitely use Cosmos DB multi region write.

26:31 We replicate our accounts, so that we can always shift reads

26:34 and writes to other regions in case of outages.

26:38 But generally also for good performance,

26:40 you need to be able to allocate your data

26:43 in Cosmos DB accounts that are close by.

26:46 One particularly interesting feature that we use for Cosmos db,

26:49 is that we have tiers of Cosmos DB accounts where some

26:53 of these accounts that are being used to store globally accessed metadata.

26:57 So like things like account, information and user settings,

27:01 we actually replicate that to dozens of regions,

27:05 as opposed to just a single region.

27:09 That's fascinating.

27:11 You've done a lot of innovations on top of Cosmos DB.

27:13 What's your favorite?

27:18 I would say that, being able to move data across Cosmos DB.

27:24 So we have a lot of products that will start,

27:30 on a single Cosmos DB account, and in a single container.

27:35 But as they grow or as their needs change,

27:38 iteration is very important at OpenAI.

27:41 As needs develop, maybe we realize

27:44 that they have much higher performance, constraints.

27:48 And so being able to transparently migrate their data from one Cosmos

27:52 DB account to a different Cosmos DB account that is globally replicated,

27:55 that is something that we've spent a lot of time trying to build around.

28:00 And the Cosmos DB team has given us a lot

28:03 of the building blocks that we need to be able to do that.

28:06 That's awesome.

28:07 We should at some point join the publishers.

28:10 A lot of viewers would love to see us as well.

28:14 Well, we have many app developers that are

28:17 early in career or potentially are, still in college.

28:21 What would be your advice to them?

28:24 Yeah, you know, I think this is evident also

28:26 in the way that our system has been built, to accelerate product teams.

28:31 There's a big culture at OpenAI about shipping things,

28:35 and shipping things quickly.

28:37 I think it's important to have like a vision of what you're working towards,

28:40 but don't over build solutions to problems.

28:43 Try and get something working into the hands of your users,

28:46 and then you can evolve and iterate on it.

28:49 Almost Never.

28:51 You will get the first iteration correct, but you'll learn from that experience.

28:57 Thank you so much John.

28:58 And thank you so much for being our customer

29:00 and thank you for joining us at this conference.

29:05 Thanks for the invite me.

29:10 Thank you, John.

29:10 Now what you noticed John mentioned Codex use in OpenAI,

29:15 everyone uses coding agents, right?

29:17 As I said, based on telemetry,

29:19 more than half of our customers are now using coding agents.

29:24 And for you to get started, let's say you don't know about Cosmos DB yet, right?

29:31 The best thing is to install our Agent Kit.

29:34 This is a set of the skills that gets you agent

29:37 up and running and becoming the world class expert in Cosmos DB.

29:40 It has 80 plus skills in growing encompassing best

29:44 practices with Cosmos DB across the entire app lifecycle.

29:48 From data modeling to partitioning to queries,

29:51 indexing, throughput management, high availability and monitoring.

29:56 Now to see it in action, let me invite our fearless product leader,

30:00 Andrew Liu to show Agent Kit in a demo.

30:05 Hi Kirill.

30:08 So to set the stage, I am building an app here.

30:12 It's a travel planner.

30:13 It's a multi agent Cosmos DB application and it's going to help me plan a trip.

30:19 So here I'm planning a five day family trip

30:22 to LA that includes my 30 year old daughter.

30:24 Now what's important about this is that it's not just helping me plan the trip.

30:28 It remembers where I left off.

30:32 This is done through by storing memory in Cosmos DB.

30:35 This includes short term memory, long term memory, and vector.

30:38 Now here's the thing.

30:39 It's not about the application, but rather, I'm new to building AI apps.

30:43 And the thing I constantly question myself is did I build it right?

30:47 Because at scale, getting it not right gets expensive.

30:52 What if I pick the wrong partition key?

30:54 What if I have the wrong data model, the wrong index?

30:57 I mean that's burning RUs in production.

30:59 That's the kind of thing that you don't want to surprise you when you launch.

31:03 And scale.

31:05 So let's jump over into what this looks like in vs, code.

31:10 First, thing I want to show you is

31:11 actually our new VS code extension or relatively new.

31:14 Many of you guys have seen the portal.

31:17 We get it.

31:18 We've also gotten a lot of the feedback on.

31:19 We really want to go and improve the desktop tooling.

31:21 So we've been putting a lot of effort into the VS code extension.

31:25 What you'll find here is you'll get a lot of the same operations but you

31:28 can do it now in line with your code without having to context switch.

31:32 So here I have my Cosmos DB account, and it works just the same way.

31:36 I can go and query the database right behind this application and what

31:41 you'll see is I'm storing a bunch of data in Cosmos DB.

31:44 I have my short term memory in the form of messages.

31:46 I have my long term memory, in the form in this memories container where I

31:50 have things like my declarative memory which is durable facts,

31:55 procedural memory where I have behavioral preferences and episodic memory,

32:00 which is going to be things that are a bit more trip specific.

32:03 Now let's go into the magic on Agent Kit.

32:08 Did I even get this right?

32:10 So installing this is actually very easy.

32:13 Just1 CLI command away.

32:15 You can use npx skills add and go and add the Cosmos DB Agent Kit.

32:20 It's going to go through a few different questions,

32:23 like setting it up for the project scope or doing it globally on your computer.

32:31 But when it sets it up, what it's going to do is it's going to go

32:34 and create this on the project scope a agents container,

32:38 or rather directory and it'll load all of my skills in here.

32:44 Now that I have my agent skills set up,

32:47 let's go and go check out the quick application.

32:54 So I am going to go take a look here and I'm going

32:56 to have it just go review my Cosmos DB data model for performance issues.

33:02 Now this isn't just a generic LLM looking at the code.

33:05 The Agent Skills is a package set of specialized expertise modules.

33:09 Right?

33:09 We're taking a lot of context about Cosmos

33:12 DB and bringing that into modules, just for this.

33:17 Now there's no services, no accounts.

33:19 It's a repo of skills so the agent gets smarter about Cosmos.

33:22 It's teaching the agent how Cosmos DB actually works.

33:25 It'll include data modeling,

33:27 best practices, partitioning, indexing, RU economics,

33:30 and make sure that that actually matches up with my data access patterns.

33:34 These are things that it took us decades

33:36 to learn over building real world Cosmos DB applications.

33:41 Now you can get that directly in the form of agent skills.

33:47 Now this is what's really cool about this.

33:50 It's not just doing some generic pattern

33:52 matching and giving me generic NoSQL advice.

33:55 It's reading my actual code.

33:57 It's taking a look at My code, my infrastructure as code, Bicep templates,

34:01 correlating my partition key, my index policy,

34:05 my document shapes, to the different data access patterns.

34:09 It knows what good looks like.

34:12 Specifically for Cosmos DB,

34:14 this is expertise that used to live in one senior architect's head on my team.

34:20 But it was hard to get a hold of those individuals.

34:23 I had to book them a week out.

34:25 Now it's just in my editor here on demand.

34:30 So let's take a look at the summary here.

34:32 It was able to go and find an order of priority and highest roi.

34:38 Oh shoot, I missed some composite indexes for my order by query

34:42 that can shave off a bunch of RUs for each of my different queries.

34:46 And when I'm running at scale with high QPS, that's really going to add up.

34:50 It's also going to go and prioritize additional things like hey,

34:54 I should go and fix my memory's partition key once again.

34:58 Hey, if I get the wrong partition key,

35:01 as this application scales to many partitions,

35:03 if I have a lot of cross partition queries that's going to go add up.

35:07 So what I'm really happy about this is

35:09 it's catching this before I hit live production.

35:16 So there's two key takeaways I'd like

35:19 to for everyone here in the audience number one, if you're a new developer,

35:24 Cosmos DB doesn't have to feel foreign or scary like

35:29 all the cognitive load of partitioning data modeling right here.

35:33 I can go and interact here with my Cosmos DB Agent Kit and go and get a live

35:39 code review of my system and make sure

35:41 that what I'm doing is well architected by default.

35:45 Also if you're an existing Cosmos DB developer,

35:47 well this also lets you go and scan your code,

35:51 scan your past projects and go look for opportunities for optimizing

35:55 the performance as well as cost characteristics of that application.

35:59 Now back to you Kirill.

36:07 Wow, thank you Andrew.

36:08 I think the last point resonated with me so much.

36:11 This is not just allows you to increase your productivity and build apps faster.

36:16 It's actually world class expert in Cosmos DB as your code reviewer.

36:21 You can plug it in into your CI CD processes in your code,

36:25 review flows and it's right there catching any bugs that you

36:29 may introduce or helping you tune your code to be more efficient.

36:34 Now we have great coding agents skilled on Cosmos DB.

36:38 We have fantastic AI apps,

36:41 but with all of that we need systems to continue to be reliable,

36:45 secure and performant.

36:46 And Cosmos DB is built for that.

36:49 That is the number one priority for Cosmos DB as a database.

36:53 Cosmos DB is a database that can take your applications anytime,

36:56 from gigabytes to petabytes,

36:58 from hundred transactions per day to trillions of transactions per day.

37:02 You don't need to worry about it if you hit on overnight success at some point.

37:08 Cosmos DB is the only database in the cloud

37:10 that offers you guarantees not only five nines down,

37:14 uptime across any consistency modes, including the strong consistency.

37:20 It also provides you money back guarantees for zero data loss.

37:23 Thanks to the consistency SLA's financially backed.

37:29 Finally, it's the only database that gives you full autonomous resiliency,

37:33 zero touch for your system.

37:36 Now this scale and reliability is made

37:38 available to you not only just by software, but also by an awesome hardware.

37:43 And to discuss how we run the fleet of Cosmos DB,

37:47 I would like to invite Steve Berg,

37:50 Corporate Vice President and General Manager of the Service

37:53 CPU Cloud Business Group at AMD on stage.

37:56 Welcome Steve.

37:57 Thanks for having me.

37:59 Glad to be here.

38:00 Great to have you.

38:03 Now, Steve, from your perspective,

38:04 how does AMD and Microsoft partnership impact how customers experience Azure?

38:09 That's a great question.

38:10 Before I answer the partnership piece, let me give you a little bit

38:13 of history about our working together with Microsoft.

38:16 Back in 2017 we launched our first EPYC processor, Naples,

38:20 and Microsoft was actually the first cloud provider

38:23 to offer it to make it available in the cloud.

38:26 That created a really unique bond between us and Microsoft.

38:30 Over the past decade,

38:31 we've worked together to launch five more generations of of EPYC instances

38:35 that are available in about 60 different compute families across the globe,

38:39 with our most recent one being Turing, which is an excellent product.

38:43 So we're constantly working to optimize our hardware

38:46 along with your software to be more performant,

38:49 more efficient and provide new capabilities with each generation.

38:53 We do this with a deep engineer

38:55 to engineer collaboration on making roadmap changes

38:57 that improve our products and help them

39:00 get tuned better for Microsoft Azure needs.

39:03 So some of these examples include things like confidential computing,

39:06 3D V cache and high performance bandwidth that's used on some of your VMs.

39:13 This partnership has led to more Azure first parties adopting our hardware.

39:20 And that is what the partnership provides is Cosmos DB.

39:24 And users on Cosmos DB get all the benefits of decades of work

39:28 that we do together and all the hard work that happens within our collaboration.

39:32 I could agree more.

39:34 We actually use AMD quite heavily in our Cosmos DB fleet and that's what

39:38 enables us to scale with demands like

39:40 OpenAI and other super fast growing applications.

39:43 Yeah, exactly.

39:44 OpenAI and all those just get all the benefits

39:46 of us working together with everything under the hood.

39:49 So it's really, that's what the partnership really brings.

39:53 That's awesome.

39:53 Now, in today's world with global situation, energy efficiency is top of mind,

39:59 how should builders think about balancing

40:01 the performance needs and sustainability and energy efficiency?

40:05 Yeah, we absolutely focus on that.

40:07 Our product teams and our engineers are

40:09 always looking at what's energy efficiency and sustainability.

40:13 And fortunately, builders who choose to run Azure get all of this benefit

40:17 without even really having to think much about it at all.

40:20 Because we do all of the for them.

40:22 We work to deliver generation after generation EPYC CPUs with more performance,

40:27 better efficiency per watt and improved TCO,

40:30 which directly translates into users and builders being able to get

40:35 benefits of cost and maximizing their performance per watt per dollar.

40:40 That's the critical metric.

40:42 So when you consider the global reach

40:43 of azure with over 70 regions and counting,

40:46 it's absolutely imperative that we maximize performance

40:49 for those users in a sustainable way.

40:51 So focusing on that performance per dollar is where we concentrate on.

40:57 And when Cosmos DB team rolls out new improvements to your SaaS offerings,

41:01 builders implicitly get all those benefits

41:04 just by using the modern infrastructure.

41:06 And when end users roll their own infrastructure as a service,

41:11 if they gravitate towards amd, epyc, vms, they also receive those benefits.

41:16 That's amazing.

41:17 That's exactly what we all need to do, right?

41:19 We don't want to think about problems that we can't solve.

41:22 We want you to solve them for us.

41:23 We want to make it easy for you.

41:24 Yes.

41:25 And semiconductors has always fascinated me

41:27 as a really tough industry because you

41:29 have to plan ahead five years ahead for what CPUs customers might need.

41:35 How do you do that?

41:36 Yeah, it's interesting you asked that because I don't think anyone could

41:41 have predicted what was happening with AI in the past year or so.

41:45 And it's really fascinating to see what you're doing with agentic AI today.

41:49 I think the rapid pace of AI took everyone many by surprise and no

41:54 one would have guessed the importance that CPU provides the role in this market.

41:58 Right.

41:58 CPUs are becoming what I like to refer to as the new bacon.

42:02 The inference, workloads, orchestration, pre and post processing, agentic AI,

42:07 general purpose, all continue to use CPUs in a very greater detail.

42:14 And so cloud infrastructure will continue to move from general

42:18 purpose Commodity type machines

42:19 to increasingly customize differentiated platforms.

42:23 Now we're going to focus on what's important,

42:25 that performance per dollar per watt that is so critical in our infrastructure.

42:30 So we already do a lot of this custom work today.

42:32 I provided some of the with Microsoft,

42:34 I presented some of those examples earlier and we'll

42:37 continue to expand in this space in the coming years.

42:40 So if I leave you with a prediction

42:42 of what's going to happen in the five year span,

42:44 I can tell you that we're going to launch

42:47 more and more VMs with Azure and Cosmos DB

42:51 and user developers in the planet will be able

42:55 to get benefits from AMD and our collaboration together.

42:59 That's exciting.

43:00 So Cosmos DB customers can look forward to more transactions per

43:03 dollar thanks to the work that we at a better cost efficiency.

43:08 Absolutely awesome.

43:09 Well, thank you so much Steve.

43:10 Thank you for our partnerships.

43:12 Thank you for sponsoring this conference

43:14 and made this wonderful sessions available.

43:16 Appreciate you having me and thank you all for the time.

43:20 Thank you.

43:27 That was amazing.

43:29 Now Cosmos DB is not the only NoSQL database we have on Azure.

43:34 Last year we launched our second NoSQL database, Azure DocumentDB.

43:38 And unlike Cosmos, DocumentDB started from an open source project.

43:42 This is a DocumentDB project at Linux foundation that Microsoft,

43:46 Amazon, Google, Yugabyte and 12 other partners are collaborating on.

43:53 Both Amazon and us are offering fully managed services.

43:57 Amazon DocumentDB and Azure DocumentDB on top of this open source project,

44:01 this is a fully managed document database that has

44:06 a number of advantages over other document databases.

44:09 It's faster, it has faster AI search for example.

44:12 You can deploy it anywhere in any cloud or on premises.

44:17 And it's more than 40% cheaper than any other document database on the cloud,

44:23 including AWS MongoDB Atlas, for example.

44:27 So let's take a look at an example of this.

44:30 Let's say I want to send 1 million

44:32 writes per second to my database on Azure DocumentDB.

44:37 A single instance, single shard 80 allows me to do that no sweat.

44:45 An equivalent scenario,

44:47 equivalent setup on AWS with MongoDB Atlas with the same SKU M80

44:53 would only give me 800k transactions per second according to AWS documentation.

44:59 We have the codes if you want to run it and see for yourself what it does.

45:02 But according to docs, it's only 800k and it cost me 60% more.

45:09 So all in all, I can get twice as many

45:14 transactions per dollar with Azure DocumentDB for every SKU.

45:19 It's at least 40% cheaper on skewer SKU basis overall,

45:23 as you scale your application load more data,

45:26 you get more than 85 reduction in TCO for your application.

45:31 So DocumentDB is really cost effective.

45:35 Take it for a spin and get the cost efficiency.

45:39 Now when should I use which?

45:43 If you are having a lot of simple queries,

45:46 if you want to have a mission critical application,

45:51 if you want to scale limitlessly, if you like serverless,

45:55 form factor, you like global distribution, Cosmos DB is for you.

45:59 This is the only database in the cloud that can do all that.

46:02 But if you prefer open source, if you like MongoDB API compatibility,

46:06 if you want to potentially port your application to other

46:10 clouds or on prem and you like the core architecture,

46:14 then DocumentDB is a good choice for you.

46:19 Multi cloud, open Source, Complex Queries,

46:24 DocumentDB, Lots of Scale, Serverless,

46:27 5, 9 Reliability, Cosmos DB now what if I want both?

46:32 What If I want 5, 9 resilience and I want cross cloud portability?

46:39 As of today you can do it both.

46:42 We created an open source project, we launched it today called MultiCloudDB.

46:48 It's fully open source,

46:49 MIT licensed and it provides you cross cloud portability across Azure AWS

46:56 and GCP by providing one data access layer on top of Cosmos DB AWS,

47:03 DynamoDB and Google Spanner, we provide you cross cloud disaster recovery.

47:14 We will provide you cross cloud replication with sync

47:16 as part of this project you don't have to choose.

47:20 You can have five ninth availability and cross cloud portability

47:24 at the same time using this project and with Cosmos,

47:28 Dynamo and Spanner from each of the hyperscalers.

47:34 Now these are just some of the topics we will cover in the next few hours.

47:38 And while we have your attention, join us for Skill Challenge to win free Cosmos

47:44 DB certification and stay tuned for more exciting sessions today.

47:50 Install Cosmos DB Agent Kit, look for more sessions,

47:54 build an app and maybe you can talk

47:56 about it at Our next Cosmos DB conference 2027.

48:00 Enjoy the rest of the event.

48:06 Thank you so much Kirill.

48:08 Hey, wasn't that a great keynote?

48:12 Thank you so much Kiril.

48:15 Thank you Andrew.

48:16 Yeah, thank you so much.

48:17 There is so much to take in, from AI to real world applications.

48:23 Even with the interview with AMD that was so interesting.

48:27 There's so much that you can do with Azure Cosmos DB.

48:30 Yeah, and you know what, the demo really brought it to life.

48:33 I love AgentKit.

48:34 I want to say hi to Theo and Saji.

48:37 They're doing such amazing work.

48:39 Yeah.

48:39 And if you're enjoying the event so far, please let us know in the YouTube chat.

48:43 We love to hear what you think.

48:46 Up next, we've got a quick story from one of our partners.

48:49 Let's take a look.

48:53 Hi, my name is Mick Feller,

48:54 I'm a distinguished software engineer for the ODP Group.

48:57 But you might better know us as Office Depot.

49:05 We are currently using Cosmos DB mainly, in our homegrown AI,

49:10 application that we built for our enterprise called the Personal Assistant.

49:13 We use it for time series to do analytics,

49:16 but more importantly we use it for, for HR profiles.

49:19 So we store each employee's profile so we can give more context

49:23 to the AI models to be able to answer questions better in contextual form.

49:32 The main benefits we've gained from using Cosmos DB is the hands off management.

49:38 It scales automatically.

49:39 We have full control over how much we use, how much we don't use.

49:43 And it's been just amazing around the whole management of the actual

49:46 tool because the scalability we can control cost to fine grain level.

49:51 And that's why we chose Cosmos DB.

49:58 We've seen almost 100% year over year increase

50:01 in usage and big driver of that was Cosmos DB.

50:05 Because it scaled so flawlessly with our needs,

50:08 we were able to seamlessly scale to that level of usage for the customers.

50:13 That's our success story with Azure Cosmos DB.

50:16 And I can't wait to hear about yours.

50:21 Yeah, welcome back and if you're just joining us,

50:25 yes, you missed a fantastic keynote.

50:29 I can't believe they missed it.

50:30 I know, but don't worry, we've got you.

50:33 Everything from today will be on demand,

50:35 including exclusive sessions so you can catch up at any time.

50:40 So grab the full playlist here at aka Ms.

50:44 Cosmosconf26 playlist.

50:47 Let's give some shout outs.

50:51 Do you want to just say to anybody who's watching,

50:53 I saw some great people from Sweden.

50:56 Yeah, I saw people from the uk, Mexico, all around.

51:01 So tell us where you're watching from.

51:03 We want to give you shout outs if you got

51:06 some interesting stories also about what you're doing with Cosmos DB,

51:09 we may just throw here.

51:10 Oh, we've got a few of our.

51:12 So, well, we've got Mike, from Norwich, thank you so much.

51:18 And then one of our great influencers, Ravi Jain.

51:21 Thank you so much for all the support.

51:23 Yeah, no, thank you so much for being

51:25 here and please keep it coming in the chat.

51:28 So if you can tell us what you're building with Cosmos DB,

51:31 we'd love to shout you out.

51:34 We promise we Won't ask for your partition key unless you bring it up, which.

51:39 Okay.

51:39 Hey, that's why we have that Agent Kit, right?

51:41 Yeah, exactly.

51:42 Andrew showed us.

51:43 If you want to keep connected, aside from the YouTube chat,

51:47 after the event, you can join us on the NoSQL channel on Discord.

51:52 You can go to aka Ms.

51:55 Discord Cosmos Conf.

51:57 Yeah.

51:57 And also the Cloud Skills challenge.

52:00 Yeah, absolutely.

52:01 We want you to get, that ability to win a, 500.

52:07 Excuse me.

52:07 The first 500 of our entrants will be able to enter

52:11 to win one of 500 free certificate vouchers for the DP-420 exam.

52:16 So don't miss out on that opportunity.

52:18 Yeah.

52:18 And if you want to get started, visit AKA Ms.

52:22 aka.ms/cosmosdbconfchallenge.

52:24 And you know what, quick feedback helps a ton.

52:27 So make sure that you fill out our evaluation survey.

52:30 We want to know what you thought, so go to aka Ms.

52:33 Cosmosconf2026survey.

52:36 So mouthful of all those URLs I know.

52:40 All right, well, it's back to keynote programming and it's the last part of it.

52:44 And then we're going to get into all the other sessions.

52:46 But we wanted to do a really great featured session.

52:50 I used to watch the Jetsons as a kid

52:52 and this one kind of near and dear to my heart.

52:55 Jade, stop this crazy thing.

52:56 Cosmos DB Best Practices A lesson from Spacely Sprockets.

53:01 Oh, wow.

53:02 Wait, that sounds amazing.

53:04 Here's Sid Anand from Walmart and welcome to my talk.

53:09 My name is Sid Anand.

53:11 I will be talking today about

53:13 Cosmos DB Best practices and architectural patterns

53:16 in the context of building a fictional E

53:19 commerce site for a company called Spacely Sprockets.

53:23 First, a little bit about me.

53:26 My name is Sid Anand.

53:27 As I mentioned earlier, I'm a technical fellow in the data platform at Walmart.

53:32 I've spent about 25 years in the SF Bay Area working on end

53:37 to end data infra at scale at a variety of companies big and small.

53:44 Let's dive into a, product overview of the site we're trying to build.

53:49 First of all, since it's an E commerce site,

53:51 we expect to see some core shopping experience features such

53:55 as a homepage with some sort of integrated search landing page.

54:00 On this page you'd expect to see some product carousels and a search box.

54:05 A user may search for items or click on one of the personalize

54:09 recommended items and that will take him or her to a product details page.

54:13 If the user likes what he sees,

54:16 he can click on add to cart to add that item to his or her cart.

54:21 And later on he or she might decide to check out and that would

54:26 kick off order processing flows that would result in delivery of this item.

54:32 Last but not least, after the user has bought multiple items from the site,

54:38 he or she may want to go back and look at current and past orders.

54:43 And also, like other modern e commerce sites,

54:46 we'll also have other features like profile and preference,

54:51 payment and billing, authentication and account management.

54:54 All of these things are important to have,

54:56 but they're outside of the scope of this talk.

54:59 We're going to focus primarily on the core shopping experience.

55:04 So let's start with a hundred foot view of what we need in engineering.

55:08 We'll start with a classic three tier

55:11 architecture made up of a presentation tier,

55:14 an application tier and a persistence tier.

55:17 There are many options and choices

55:19 for stacks for presentation tiers and application tiers,

55:23 so we'll just go with what we know best.

55:25 We'll pick React Vue and Angular for the presentation tier,

55:29 and we'll pick a spring boot based, platform for the application tier.

55:36 The next big question we need to ask ourselves is what will be

55:40 the wire protocol and format for data

55:42 that is transacted between our different tiers?

55:46 When making a decision like this, we often have to consider multiple things.

55:51 For example semantic richness.

55:53 Does the wire format, have a rich type system that is natively supported?

56:01 And usually if it does have a rich type system,

56:04 it comes with schema enforcement and schema evolution.

56:10 But sometimes schema enforcement and evolution go against developer velocity.

56:15 So we have to decide which of those is more important for us.

56:18 And because we'll have incidents in production,

56:21 we'll need a way to debug the data flying back and forth between microservices.

56:26 So we'll want to also ensure that we have

56:28 a rich tool and ecosystem to support whatever we pick.

56:32 Most often teams are given two choices or they're faced with two choices.

56:38 Either they go with gRPC and protobuf, which is much more efficient,

56:42 then the other option which is REST over HTTP with JSON.

56:47 However, the rest case with JSON is more widely used

56:51 and it is what we'll pick for our talk today.

56:55 So now we've picked our top two tiers and the interchange format between them.

57:01 We have to make a decision about what

57:03 to pick for our database in the presentation tier.

57:07 While making this decision, we can consider either a tabular,

57:13 data model or a JSON based document model.

57:17 If we pick The JSON based document model.

57:19 It simplifies our mental model and our cognitive load.

57:23 It reduces our cognitive load because what we're transacting

57:26 between microservices is also what we're storing in the database.

57:30 We don't have to concern ourselves

57:32 with flattening from JSON into a tabular format.

57:38 And we also want to keep the SQL

57:40 as our query language because it's very expressive.

57:44 Therefore, we need a database that provides both

57:47 the flexibility of JSON and the expressiveness of SQL.

57:52 Luckily, Cosmos provides both as part of its NoSQL API.

57:57 So now we have a full stack.

58:00 What are some of the other benefits of Cosmos?

58:03 Well, we get scalability thanks to its support for horizontal partitioning.

58:09 It's globally distributed, it provides a range of consistency options.

58:15 So we get tunable consistency and through this we get support for transactions,

58:20 which is very important.

58:22 Finally, it gives us a flexible cost model.

58:25 Now being an enterprise database, it also supports change,

58:28 data capture, snapshot and point in time recovery and other features.

58:36 But this is all very good and what we need in a database.

58:40 Now that we have picked our stack, let's talk about the product.

58:45 While we have five different things to pick from, we're

58:48 going to focus on the three most challenging ones.

58:50 The shopping cart, order processing and also viewing a customer's order history.

58:56 Let's start with the shopping cart.

58:58 What are our engineering requirements for the shopping cart?

59:02 First of all, we need it to be globally distributed.

59:05 If a region has an outage, we don't want an outage in our shopping cart.

59:10 We want people to be able to add to their cart

59:13 and view cart no matter what is happening in a given cloud region.

59:18 Now that we have picked a database that works in multiple cloud regions,

59:22 we also care about the data being consistent.

59:26 If a user makes a change to a cart in one region,

59:29 he should see that exact same change reflected in other regions.

59:34 And last but not least,

59:35 we need all of these interactions to be low latency because any

59:39 type of latency friction will cause a drop off or conversion problem.

59:46 Luckily, Cosmos DB provides all of these.

59:49 So we can create a cart container

59:51 that operates in multiple regions and uses session consistency.

59:56 This provides lower latency and better

59:58 availability for writes than globally strong consistency.

1:00:03 Let's look at this in detail.

1:00:05 Let's say the user wants to add an item

1:00:08 to his cart after he's called that add to cart.

1:00:12 The presentation tier forwards the request to the application

1:00:16 tier which will call upsert on the Cosmos DB instance.

1:00:22 Once it's successful, the database will return,

1:00:26 200 or 201 error code, response code, along with a session token.

1:00:32 And that session token will flow all the way upstream back

1:00:35 to the user where it will be stored in a cookie.

1:00:38 Now let's say after this set of steps is finished, the region goes down.

1:00:44 Like we have a total region, one outage.

1:00:46 But now the user wants to view his cart

1:00:49 because we stored the session token in the cookie.

1:00:53 The user will present that session token when he calls View cart.

1:00:58 That will go to the presentation tier and get forwarded to the application tier.

1:01:03 The application tier will then call query

1:01:05 against the cart collection presenting that session token.

1:01:09 If the data has been replicated from the first region

1:01:13 to the second region by the time this call is made,

1:01:17 the item will be found and returned to the user

1:01:19 so the user can view a consistent cart.

1:01:23 But if there are any delays in replication,

1:01:26 rather than showing the user an inconsistent cart,

1:01:29 what will happen is the application tier will continue to retry until it,

1:01:34 syncs with the initial region, at which point it will show the final cart.

1:01:41 The design provides a globally consistent low latency shopping cart.

1:01:45 And it's exactly what we need for our application.

1:01:49 Now that we talked about shopping cart, let's talk about order processing.

1:01:53 Order processing starts when we're, checking out.

1:01:57 We convert the cart that we check out with into an order.

1:02:01 An order is made up of one or more order line

1:02:04 items where each line item encapsulates a product and its quantity.

1:02:09 Let's look at an example.

1:02:11 Let's say the user had a cart with ID 123 and checked out.

1:02:16 As you can see, in this cart there are three items.

1:02:19 The first one is for the foo sprocket, the second one is for the bar sprocket,

1:02:24 and the third one is for the monkey sprocket.

1:02:26 And each of them come with different quantities.

1:02:29 After checking out, the cart will be converted into an order with ID

1:02:34 A and we can see the order line items are carried over.

1:02:40 So what are the engineering requirements for order processing?

1:02:44 When an order is inserted into the database,

1:02:47 we want the order and all of its line items to be inserted, atomically.

1:02:54 It should also be possible to upsert line items

1:02:58 independent of each other and of the order itself.

1:03:01 On the read side, we should make it really easy for an application

1:03:06 to read and order and all of its line items in a single query.

1:03:10 But they may also want to read either

1:03:12 the order or align item independent of one another.

1:03:16 To achieve this, in Cosmos, we have two options for data models.

1:03:21 Option one gives us a Denormalized model,

1:03:24 also known as embedded or a single document model.

1:03:28 Option two gives us a normalized AKA reference, AKA multidoc model.

1:03:33 Let's have a look at these.

1:03:36 Let's say I pick option 1.

1:03:38 In option 1 I have an order JSON.

1:03:41 I see that the partition key is the order ID,

1:03:44 which happens to also be the ID of the document.

1:03:50 There are other, properties in this document that are relevant and important.

1:03:53 For example the status of the order.

1:03:56 Also, the shipping address and payment address are included,

1:03:58 as well as the customer's id.

1:04:00 And if you notice there is an array type called orderline items.

1:04:04 In this array we have each of the orderline items that are part of the order.

1:04:11 In option two, what we do is we create a parent order document and its

1:04:17 children order line item documents and we

1:04:20 store them together in the same container.

1:04:23 Because they use the same partition key,

1:04:25 they are also part of the same partition.

1:04:29 The order line items have their own line item IDs,

1:04:32 but they are all stored with its parent order in the same partition Now,

1:04:39 how do you choose which model to use?

1:04:42 Well, if you have orders that don't have too many line items,

1:04:47 if you know reliably that there will be less

1:04:49 than 50 of those, option one is a good choice.

1:04:52 It's also a good choice if your dominant access pattern

1:04:55 is always on the order with all of its line items,

1:04:58 such as reading the entire order.

1:05:01 And also if you want simplicity and lowest read ru cost,

1:05:05 option one is a good one for you.

1:05:08 But if the orders can be unbounded,

1:05:11 have unbounded number of line items, you must choose option two.

1:05:16 And option two gives you some other benefits as well.

1:05:18 It allows you to update individual line items.

1:05:22 Because you may want to update the fulfillment status,

1:05:24 for example it's been delivered, and also the quantity.

1:05:27 If you amend your order, you also may want to independently query the line

1:05:33 items from each other from the parent order.

1:05:35 And so the multi doc approach gives you this.

1:05:39 And how does it work?

1:05:40 Well, if you use option one whenever you want to read an item,

1:05:44 in this case read an order, you just call readitem for the order and it'll

1:05:49 return the order and all of its nested line items.

1:05:53 If you want to, modify or delete the order and its line items,

1:05:59 you can also issue point operations.

1:06:01 For those with option two, you get more flexibility.

1:06:05 For example, if I want to, read a parent and all of its children together,

1:06:11 I would use the query shown here.

1:06:13 I would select star from order where the partition

1:06:16 key is in this Case the order id.

1:06:18 But I'm also given the ability to do point reads and read

1:06:22 a single document like a single line item and update its quantity, for example.

1:06:27 And last but not least,

1:06:29 whenever I'm doing any changes across the order and its line items,

1:06:33 I can use transactional batch to either update them

1:06:36 or delete the whole order and the line items.

1:06:41 Now that we've solved two of the hard problems, let's get to the third one.

1:06:44 In this case, a customer wants to see its order history,

1:06:48 which means all of the orders that it

1:06:52 has completed and their associated line items.

1:06:55 And we want to do this with low latency and low cost.

1:06:58 Now, as the order container is partitioned by order id,

1:07:02 querying by the customer ID results in a cross partition query.

1:07:08 Cross partition queries incur higher cost and latency than point queries.

1:07:13 The greater the number of partitions,

1:07:15 the greater the latency and cost of cross partition queries.

1:07:20 Now, does Cosmos provide a way to efficiently query by customer id?

1:07:25 Yes it does.

1:07:26 And it's called Global Secondary Indices.

1:07:30 Let's look at how this works.

1:07:32 Let's start with our normal container, the order container.

1:07:36 As we mentioned earlier, it is Partitioned by Order ID,

1:07:39 but it contains other JSON properties like Customer ID.

1:07:43 Now, I define a new index container called CustomerOrders,

1:07:48 which is defined by a covering query shown here.

1:07:51 Select star from order and a partition key.

1:07:55 In this case, the partition key will be customer id.

1:07:59 Once I set this up, the customer ID will get

1:08:02 live updates from the order via the full fidelity change feed,

1:08:07 filtered by, the covering queries seen below.

1:08:13 Now, the customer orders container is called an index container.

1:08:17 It's a read only container.

1:08:19 Apps cannot write to it,

1:08:21 but writes can only come to it through the order container.

1:08:25 And the beauty of this is a user can query the order container

1:08:29 by order ID or against the customer

1:08:31 orders container by customer ID at both times.

1:08:35 They are point queries, so they're cheap and fast.

1:08:37 We avoid expensive cross partition queries.

1:08:42 Now that we've covered three of the main core pieces, of functionality,

1:08:46 let's talk about some optimizations in terms of latency and cost.

1:08:51 There are three types of optimizations worth considering.

1:08:54 We've already talked about the first one which is to use a gsi.

1:08:58 The second one is about partition scoping.

1:09:02 First of all, what is partition scoping?

1:09:05 Let's look at an example.

1:09:08 Let's pretend that we have a user collection.

1:09:11 That user collection stores documents as shown here.

1:09:14 User JSON and it's backed by a POJO class shown here on the Right user job.

1:09:22 If I want to select like George Jetson from the database,

1:09:27 I would provide a point query as shown here.

1:09:30 Select star from C where C ID equals a bind param.

1:09:34 There's one problem here.

1:09:36 Every time I issue this query, there's an extra round trip that is used

1:09:41 to parse the query and generate a query plan.

1:09:47 If, I execute this query every time,

1:09:51 if I use this code every time I execute the query,

1:09:55 the query will incur an extra round trip.

1:09:58 But if I add the line shown here where I

1:10:01 set the partition key on the Cosmos query request options,

1:10:05 then the first time will be the only time that the query will be,

1:10:11 will generate a query a plan,

1:10:12 and then it will be cached in whichever pod this application

1:10:16 is running in so that subsequent queries avoid an unnecessary round trip.

1:10:21 And this can cut your point query latency by up to 50%.

1:10:27 Last but not least, let's talk about payload size reductions.

1:10:32 What I'm showing here is how RU costs vary as a function

1:10:37 of of the payload size and also based on the type of operation you're doing.

1:10:44 As you see, as I go to larger and larger documents in the payload,

1:10:51 the cost in terms of RUs increase,

1:10:53 but they increase more dramatically for point reads.

1:10:57 Point reads, double in cost, as a power of 2 of the document size.

1:11:05 However, you'll notice that under 32 KB it's cheaper than point queries.

1:11:10 But because point queries have a very flat,

1:11:14 slope over time, they end up costing less.

1:11:19 So let's look at the example from before, where we had an order I.D.

1:11:24 in model one with its array of line items.

1:11:27 In this case, let's pretend I had 100 items in my cart when I checked out.

1:11:34 If, I want to save space, what I can do is compress the order

1:11:38 line items using an algorithm like zstandard and base64,

1:11:43 encode it so that I can store it as a string field, a string property.

1:11:48 I've also shown a, property called compression,

1:11:50 which has a bunch of metadata, but that is not needed.

1:11:53 Assuming we kept that metadata, what will happen?

1:11:57 If we use Zstandard, with its default compression level of 3,

1:12:03 we can save 75% of space,

1:12:06 and this will result in almost, up to 75% in RU reduction.

1:12:13 At this point, I'm, at the end of my talk and I

1:12:15 want to thank the following people for their guidance and support.

1:12:18 Thank you.

1:12:21 Hello everyone.

1:12:22 My name is Chander Dahl and I'm the CEO of Kasden,

1:12:26 an elite AI consulting firm delivering high impact Enterprise grade solutions,

1:12:32 including custom built AI chatbots and advanced

1:12:36 automation systems tailored to complex business needs.

1:12:41 Today we're going to talk about one of the most important concepts,

1:12:45 which is agent memory.

1:12:48 And by the way, it's the same for your chatbots.

1:12:51 And why is agent memory so important and how to get it done right.

1:12:58 So today I'm going to show you four strategies,

1:13:01 starting from a strategy in which we talk directly to LLMs.

1:13:06 And let's not forget LLMs are stateless

1:13:08 and they do not have any memory whatsoever.

1:13:12 And then we'll use the other three strategies,

1:13:15 sliding window, hierarchical and entity graph.

1:13:19 And one of them will lead to the highest precision

1:13:23 and that will be the high precision or recall winner.

1:13:27 However, that doesn't always mean that it's going to be hundred percent recall.

1:13:32 With that particular strategy,

1:13:34 we'll show you different use cases and from that you

1:13:38 can decide what's the best for your particular use case.

1:13:43 We will start with 60 seed messages,

1:13:46 and the reason we chose 60 seed messages is we saw

1:13:50 that in a lot of companies when they do AI testing,

1:13:54 especially on their chatbots,

1:13:56 a lot of times they go up to 10 to 15 conversations only.

1:14:01 And guess what?

1:14:03 All the three strategies provide 100% recall up to 30 turns in most cases.

1:14:10 And that's why by the time you end up in production,

1:14:13 you may not even know that you have less than 100% recall.

1:14:19 And that's a, big no, no.

1:14:22 We also will show you 10 recall questions.

1:14:26 And the first five questions are really easy

1:14:29 and all the three strategies get them right.

1:14:33 However, the next six questions are more nuanced

1:14:36 and that shows how you need to change your testing strategies,

1:14:41 especially if you're leaning more towards

1:14:43 the easier questions during your AI POCs.

1:14:48 So let's have some fun and define the problem.

1:14:52 If you started working like me before ChatGPT was launched on AI applications,

1:15:00 you may have noticed the context limits really small.

1:15:05 The token window limit would be the most important thing.

1:15:10 Before you even started creating these applications,

1:15:13 you would notice right up front that these LLMs actually are stateless.

1:15:19 So after ChatGPT was launched,

1:15:22 a lot of people assumed that LLMs are actually stateful.

1:15:26 And in some cases they even thought that it has a memory.

1:15:30 A lot of that has to be done by engineers like you and I.

1:15:34 You might have heard of the famous quote, I told you my name,

1:15:37 my budget, my team, my deadlines, and you forgot everything.

1:15:42 And that is why memory happens to be

1:15:44 the most important conversation in enterprise applications worldwide.

1:15:50 So now that we know that LLMs have this problem

1:15:53 that they are stateless and every API call starts from scratch.

1:15:58 How do we solve this problem?

1:16:01 Well, very easy.

1:16:02 We can send the full conversation history every single time.

1:16:07 Yes, except we know all LLMs have a limit.

1:16:11 All right, then maybe you can just summarize it.

1:16:15 Let's try it out in case you send the entire

1:16:19 conversation and let's assume that's within the limit an LLM expects.

1:16:26 In that case you're still having costs that are going to explode.

1:16:32 You're also going to slow down the LLM.

1:16:35 What if you start summarizing parts of it?

1:16:39 In that case you still have that issue

1:16:41 that you may be losing some context, maybe some fact,

1:16:47 anything real world that you really care about

1:16:50 that got compressed because all summarizations have a compression problem.

1:16:56 And then finally users will have to repeat themselves quite a bit.

1:17:02 So why does memory matter?

1:17:04 One of the most important things to remember is that more

1:17:07 than 70% of enterprise projects will require multi time agents.

1:17:13 And the number one gap cited by developers happens to be long term context.

1:17:19 And by the way, if you go with the LLM alone,

1:17:21 which really has no memory, and you go with the best memory strategy,

1:17:26 we're going to compare today,

1:17:27 you may be looking at a 20x token cost difference between these strategies.

1:17:33 And that's why memory isn't just a feature checkbox.

1:17:36 It's an architectural decision that determines cost,

1:17:39 recall quality and user experience.

1:17:42 So the first strategy happens to be sliding window which

1:17:45 keeps your recent messages and then summarizes the messages before it.

1:17:50 An analogy would be a security camera that keeps the last eight hours

1:17:54 of footage and it writes a one page summary of everything before that.

1:17:59 That's really easy to implement and that's a strength.

1:18:03 And it's also going to have a low token cost, in this case roughly 1100 tokens.

1:18:09 And it's great for short conversations, especially less than 30 turns.

1:18:14 And this is something to keep in mind.

1:18:16 Just because it's great for that particular use case does not mean it's going

1:18:22 to be great for another use case where we actually require a real world fact,

1:18:27 even the one that was hundred turns ago or a thousand turns ago.

1:18:33 In that particular case, recall is going to be lower.

1:18:36 And in today's demo it's got about a 60% recall.

1:18:40 Let's not forget you will have 100% recall if it is less than 30 turns.

1:18:46 So remember, it isn't about the strategy, it's more about the use case.

1:18:50 And based on that use case, we need to pick the strategy.

1:18:54 Strategy 2 hierarchical memory think of this as three different tiers,

1:18:59 the first year being hot.

1:19:02 For example, a company's knowledge management system will

1:19:05 have your team's messages as your first tier

1:19:08 and that's part of a conversation and you

1:19:10 want to send them as it is uncompressed.

1:19:14 Number two is your weekly meeting notes,

1:19:17 let's call that tier two and kind of compressed

1:19:21 summaries because tier two not as important as tier one.

1:19:26 And then finally the least used out of the three which could be

1:19:30 your company wiki and those are extracted facts in your long term storage.

1:19:36 So better recall than sliding window.

1:19:39 And of course it can do better than it, but not as good as something that will

1:19:45 always present the right fact every single time.

1:19:48 Because tier three, as you can notice it

1:19:51 will miss on certain facts here or there.

1:19:54 Finally we got entity graph and I love this analogy,

1:19:57 it's like a detective's case board where you have all

1:20:00 these different flags and then you got linkages between them.

1:20:03 Right as every single card with connections between them.

1:20:07 And that's entity graph.

1:20:09 So what does it store?

1:20:10 Stores all the entities, Anything real world,

1:20:13 A name, organization, anything real world.

1:20:17 Any entity that is real world gets stored facts.

1:20:21 These are key value pairs for that particular

1:20:24 entity and you can have any amount of these.

1:20:26 And finally embeddings you may have heard

1:20:29 of vectors and vector databases got created simply because

1:20:34 of these LLMs wanting to search basically do semantic

1:20:38 search on your data and get you that particular retrieval.

1:20:44 Now in this case what I love about Cosmos DB

1:20:46 is that you have the native vector search inside Azure,

1:20:50 Cosmos DB and then finally relationships which are connections between entities.

1:20:56 Now why do you get 100% recall in this particular case?

1:20:59 I mean the reason is simple because we don't really have any compression.

1:21:03 We're doing the vector search,

1:21:05 we're getting those entities and those entities have

1:21:08 the facts and we can get all of that back.

1:21:11 So in this particular case, for this particular data set,

1:21:14 we are able to get 100% recall except we're also using average 1660 tokens.

1:21:21 That's pretty much 50% more than the first strategy.

1:21:24 That's sliding window and remember that's cost.

1:21:27 So that's another consideration.

1:21:30 So even though I would say there are four approaches,

1:21:32 you might have noticed I start with zero.

1:21:35 Why is that?

1:21:36 That's because I'm an engineer and as engineers

1:21:39 we use arrays and we start with zero.

1:21:41 That's not the case.

1:21:43 I really don't think it's an apples to apples

1:21:45 comparison because LLMs really have no memory, no context.

1:21:49 What you're trying to do unless you provide it.

1:21:52 So for that reason I don't think it's a good comparison.

1:21:55 But I still want you to see that you may be able to get 92 tokens,

1:21:59 but you got a 0% recall.

1:22:02 However, the other three strategies, very important, as you can see,

1:22:06 we go from 1100 tokens all the way to 1660.

1:22:09 And in this particular scenario, you go from 60% recall to 100% recall

1:22:14 at a, one and a half times the token limit.

1:22:17 But here's what I love about Azure Cosmos DB.

1:22:20 Same database, same SDK, same partition key, different recall guarantees.

1:22:25 And I can use all these three strategies at the same exact time.

1:22:31 So that leads me to why Azure Cosmos DB For agent memory.

1:22:35 Well, that's because we are in Azure Cosmos DB conference.

1:22:39 That's not the case.

1:22:40 It's really because at the end of the day all my data is in Azure Cosmos DB.

1:22:46 And by the way, for someone like me who's been working on DocumentDB, that was,

1:22:50 I think it was 2016 onwards I started working

1:22:52 on it and then it became Azure Cosmos DB.

1:22:55 And I know we now have another DocumentDB which is

1:22:57 a little different than that DocumentDB that was there long back.

1:23:03 My data and my client's data lives on Azure Cosmos DB.

1:23:06 It used to be a big problem to now have another vector storage,

1:23:11 a completely new database just for the vectors.

1:23:15 Well, not anymore we can actually have that data.

1:23:18 Your embeddings, your vectors live right alongside your data.

1:23:21 I don't actually need to have another native vector search database.

1:23:26 The schema is flexible and we talked about that.

1:23:29 I, love the session isolation.

1:23:31 My partition key in this case happens to be session id,

1:23:34 so I have zero cross partition queries.

1:23:37 And then finally, for this demo, I'm not worried about global scale at all

1:23:42 because all we have is about 60 different messages.

1:23:44 But once I go into production,

1:23:47 I don't have to worry about global scale because now I

1:23:51 can scale as much as Azure Cosmos DB would allow me,

1:23:55 which is literally all over the world.

1:23:58 And that's why Azure Cosmos DB.

1:24:01 So let's talk about the benchmark where we have about 10 recall questions.

1:24:05 You got the core questions, basic recall.

1:24:08 Why do we have that?

1:24:09 A lot of times what we noticed was all these POCs

1:24:12 that people were doing had a bunch of the same kind of questions.

1:24:16 And a lot of times those questions were created by AI itself.

1:24:19 Well, but it was also 100% recall based on questions

1:24:24 that probably don't understand how these strategies actually work.

1:24:27 So what you'll notice is 100% recall

1:24:31 for all three strategies for the core questions.

1:24:35 We also notice that every single data set is different.

1:24:38 We work in healthcare, you know, manufacturing, finance,

1:24:42 fintech, tech, pretty much name any major domain.

1:24:46 And all these bigger corporations worldwide

1:24:48 in those domains have very nuanced data.

1:24:52 That's why I wanted to show you how much difference gets made in just five

1:24:56 questions the moment you start asking nuanced questions

1:25:00 on that particular data and how the Mrs.

1:25:04 Happen.

1:25:05 So for example, in this case,

1:25:06 sliding window only gets one out of those five correct.

1:25:10 So your recall happens to be only 20%.

1:25:13 But if you have the right use case,

1:25:15 sliding window could be at 100% and that's why it gets confusing.

1:25:20 Before we do the demo, I would like for you to ask yourself which strategy is

1:25:25 going to remember a, URL mentioned once 40 plus turns ago?

1:25:32 And I hope you get the answer right.

1:25:34 Let's dive into the demo.

1:25:36 As you can see here, I've got two providers.

1:25:38 Here's Azure OpenAI and you can also click OpenAI.

1:25:40 It'll work with both the keys.

1:25:42 So if you have Azure OpenAI,

1:25:44 assuming you have the same exact keys for whatever model you're trying to use,

1:25:49 here's two different demos.

1:25:50 One is Kasden and another one is default.

1:25:53 And remember, they're both fake data, so please feel free to use your own data.

1:25:58 I'm not going to ask anything from the Direct LLM because

1:26:01 at the end of the day we're not going to get any answers.

1:26:04 Now we're going to compare the other three strategies which is sliding window,

1:26:08 hierarchical and entity graph on the right hand side.

1:26:12 It's just a way to see what the results are and what to expect.

1:26:17 So let's ask the first question and see

1:26:20 what it does while it's answering the question.

1:26:23 I just want to show you that all three of these are going to pass.

1:26:27 So this is actually a really easy question.

1:26:30 No problem whatsoever.

1:26:31 All of these strategies will get you the same exact answer,

1:26:33 which is cast and offerings are consulting, training and recruiting.

1:26:39 Same exact answer pretty quick.

1:26:42 And you'll notice hierarchical was extremely fast,

1:26:45 but entity graph is actually slower still gets you the same three answers.

1:26:51 So now I'm going to skip the next four questions and go to number six.

1:26:56 The reason is because the next four questions we're

1:26:58 going to get the same result from all the strategies.

1:27:03 So here, here's a result which says about us in contact us.

1:27:08 All right, so let's go to round number six

1:27:11 and open this and you notice it's missing Consulting.

1:27:15 Let's go to Hierarchical.

1:27:18 Run this round.

1:27:20 It's got about us and it's got consulting, which is true.

1:27:23 And that's exactly what we were expecting.

1:27:25 On the right hand side, here's your Entity graph.

1:27:32 It's got about us and it's got, consulting.

1:27:34 It's got both.

1:27:36 The answer is correct.

1:27:38 All right, so now let's go back to Sliding

1:27:40 window and we're going to go to number seventh.

1:27:44 And here we're going to open the round and see what we were expecting.

1:27:47 All right, it's got a about, it's got webpage, it's missing products,

1:27:54 and it's also missing from our stored graph context, which is true.

1:28:00 Let's go to Hierarchical.

1:28:05 And as, you can see, it's got pretty much everything,

1:28:08 except it's still missing products and also

1:28:11 missing from our stored graph context.

1:28:15 Let's go to Entity Graph.

1:28:21 And you may have noticed Entity Graph is actually slower.

1:28:24 So that's another thing to keep in mind is what your, use case is.

1:28:30 So you've got, you've got actually everything.

1:28:32 You got products, you've got the About Chander Dahl, you've got the webp.

1:28:36 It's actually perfect.

1:28:38 And it also says from our Stored Graph context.

1:28:41 That's pretty amazing.

1:28:43 Let's go back to the eighth round.

1:28:45 I'm going to pass because they all get the training URLs.

1:28:48 It's the same exact answer for all three of them.

1:28:51 And now let's just show one more.

1:28:53 And let me go back to Sliding Window.

1:28:56 Let's go to number nine, very quick.

1:28:59 And it's got something.

1:29:01 Let's see what it is.

1:29:02 Okay, so it's actually missing presentations, which is true.

1:29:07 Let's go to Hierarchical.

1:29:09 It's got presentations, so hierarchicals got all of those.

1:29:13 And same with Entity Graph.

1:29:20 And then the 10th round, it's a very nuanced round.

1:29:24 All that is getting wrong is not the response.

1:29:28 They have all the responses correct, both of them,

1:29:30 except they're missing from our stored graph context.

1:29:37 All right, so how many of you got that right?

1:29:40 I hope it's all of you.

1:29:42 So what you just saw was directllm has no prior contacts.

1:29:46 We didn't do that part of the demo.

1:29:48 But the idea is you're not gonna get any response if you didn't.

1:29:51 If the LLM does not have that data, unless you fine tune it to have that data.

1:29:56 Whereas, Sliding Window, the first five questions correct.

1:29:59 The next six, it only got one right, which is the eighth one,

1:30:03 which we didn't run but it was really easy and it was going to run.

1:30:07 Now you've got hierarchical, which was three out of ten,

1:30:11 which means in this case about eight out of ten.

1:30:14 So you notice what's happening, even though it was three out of five.

1:30:17 Sorry, not three out of ten.

1:30:19 We had the first five really easy questions and you may think you have an 80%

1:30:25 recall where all you had was a 60%

1:30:28 recall because three out of five were correct.

1:30:31 An entity graph in all the cases was 10 out of 10.

1:30:34 And, one of the things I want you to remember is that if

1:30:36 you had a sliding window which never requires you to go beyond 30 turns,

1:30:42 you would have 100% recall and sliding window.

1:30:45 So it's really not the recall as much as it is the combination of your strategy,

1:30:49 your use case and your data done right?

1:30:52 So here's your scorecard.

1:30:55 If you remove the blue, which is the first five questions,

1:30:57 sliding window is only one question.

1:31:00 That's 20% recall.

1:31:01 Hierarchical, 3 out of 5, that's 60% recall in entity graph, 10 out of 10.

1:31:07 So.

1:31:07 And 5 out of 5 and 5 out of 5.

1:31:09 So that's 100% recall cost versus recall.

1:31:13 The trade off is right here.

1:31:14 And it's important because if your use case is literally less than 30 turns,

1:31:19 you're better off going with sliding window because the cost is

1:31:22 one and a half times more in terms of entity graph.

1:31:25 You're using the least amount of tokens.

1:31:27 You're hoping you get the response also faster.

1:31:30 So that's another cost to keep in mind is performance cost.

1:31:33 Otherwise, if you really care about higher recall, well,

1:31:37 in that particular scenario,

1:31:38 entity graph would be a much better deal altogether.

1:31:42 So an easy way to look at this would be sliding

1:31:44 window for less than 30 turns supports the chat recent contacts hierarchical,

1:31:50 about 30 to 100 turns.

1:31:52 And again, you know, this is more,

1:31:54 it'll depend on your data, but let's just say ballpark,

1:31:57 that's what you're looking for, especially if you're using it

1:32:01 for planning or consulting and you have to have some key facts.

1:32:05 And then entity graph, for example,

1:32:07 you have your CRM bot or something where every fact matters,

1:32:11 like you don't want a response

1:32:13 with a financial number that's actually not correct,

1:32:16 you know, and it can also scale to way more than 100 terms

1:32:20 because at the end of the day you're really getting the fact correctly.

1:32:25 So again, you can start with one, you can also upgrade later.

1:32:30 You can also do a hybrid, which is also not a bad idea.

1:32:34 And in some cases you can have a combination

1:32:36 of, let's say a sliding window as well as Entity Graph.

1:32:39 You know, for 90% of your cases you may have sliding window,

1:32:43 especially for less than 30 turns.

1:32:45 And then you can go to Entity Graph

1:32:48 whenever there's a premium use case behind it.

1:32:51 So again, memory is a spectrum and you need to choose by recall, not by height.

1:32:57 And you gotta remember what's your use case and what your data is like.

1:33:03 What we covered was the problem,

1:33:05 which is LLMs are stateless, they have no memory.

1:33:08 Then we really covered the baseline plus

1:33:11 the next three strategies and then the architecture,

1:33:14 which was really easy by the way.

1:33:16 You must be wondering why five Cosmos containers to remember for Entity Graph.

1:33:21 You will have different containers for Entity Graph.

1:33:24 And that's not at all a bad idea in this particular case.

1:33:28 That's why we did that.

1:33:29 You can, you can do it multiple different ways,

1:33:31 but for that reason we had more than just one container.

1:33:36 And then you've got the live demo which had 10 recall questions.

1:33:40 Then you had the evidence which was 100% recall in terms of Entity Graph.

1:33:45 And then the decision,

1:33:46 which isn't really about the recall numbers that we shared here,

1:33:49 but a lot more than that and we discussed every single of those.

1:33:54 Now just remember this is not production data.

1:33:58 And this is literally fake data.

1:34:00 But this is also not production level code.

1:34:02 However, it's a good POC if you want to run.

1:34:05 And that's your barcode.

1:34:06 If you click that barcode,

1:34:08 it will take you to this blog post that explains the nuances,

1:34:12 how the decisions were made,

1:34:13 and it also explains to you the different parts of the code.

1:34:17 I just want to show you this real quick for your code to remember.

1:34:23 Here's your direct LLM.

1:34:25 And if you're a dot net developer you can look at the right

1:34:27 hand side of the code and if you're a Python developer,

1:34:29 you can look at the left hand side of the code.

1:34:32 Directllm.

1:34:33 It's literally sending the system prompt and then

1:34:36 creating the message chain and then making a call.

1:34:40 Very simple.

1:34:41 That's how you've been making all your calls so far.

1:34:44 Next, the sliding window.

1:34:45 And as you can see here we have

1:34:47 the summary for anything beyond 30 messages and then

1:34:51 just 30 messages where everything is uncompressed and it's

1:34:55 part of the context we're sending to the LLM.

1:34:59 And as you can see here, the window size is 30 and this is your one hour,

1:35:03 which is we just chose that time to live.

1:35:06 And that could change for whatever you're trying to do.

1:35:10 And then you've got the hierarchical memory

1:35:11 and here is where all your tiers are.

1:35:14 And as you can see, here's your tier one size,

1:35:17 your tier two block and the max tier.

1:35:20 Same for both Python as well as Net.

1:35:23 And then you've got the entity graph.

1:35:24 This is actually really easy to produce because

1:35:27 your entity entities and your embeddings are living together.

1:35:32 So one thing to remember is that this is exactly how we anyways code

1:35:37 when we are doing programming because we have the same exact objects now we just

1:35:42 have a way to retrieve those objects and then make it part of of your LLM

1:35:48 answer by giving it the context which it needs to give you the response.

1:35:53 And finally here's your read path on entity graph.

1:35:57 And this is just a small blog post that I prepared only for the demo.

1:36:02 However, a lot more detailed blog post is right here.

1:36:06 If you want to do a more involved code walkthrough,

1:36:09 you can click that link and it goes through a lot more than what I just shared.

1:36:16 So let's not forget,

1:36:17 here's the link to that and I hope you take advantage of it.

1:36:21 Let me know if you have any feedback for me.

1:36:24 I had a blast and hope you have

1:36:26 a blast coding and using this in your applications.

1:36:29 Thank you for having me.

1:36:31 Have a great day and a great rest of the conference.

1:36:37 Welcome back.

1:36:38 Thank you so much to our speakers.

1:36:40 I believe that was Sid or Ann Chander.

1:36:43 They were both great, right?

1:36:45 Yeah.

1:36:45 Honestly that was a very fun and super practical look at Cosmos

1:36:49 DB's best practices through the story of get excited, spacey sprockets.

1:36:54 Jane, get me off this crazy thing.

1:36:57 Right?

1:36:58 Well, at least if you don't have a well architected partition key.

1:37:01 Right.

1:37:03 Always on availability,

1:37:05 low latency across the galaxy and keeping costs under control.

1:37:09 Those are all trade offs we all face when we build real systems.

1:37:14 But you know what, we're going to talk way more about

1:37:17 that and I want to just make sure that we're acknowledging you, the audience.

1:37:21 So let's take a quick look.

1:37:23 We've got some really, really nice comments, but I wanted to just show this one

1:37:29 the thought processes people go through when designing solutions.

1:37:34 You know, there's a lot you have to make.

1:37:36 So many considerations when you're designing a production application.

1:37:40 And you know what?

1:37:41 You all are here to learn about that.

1:37:42 Patty, what do you got?

1:37:43 Yeah, honestly, for me,

1:37:44 I've got nothing but love in this group chat or in the comments section.

1:37:49 So please, if you are around.

1:37:53 We'd love to hear from you.

1:37:54 Feel free to drop a comment, share what you're building or, just show some love.

1:38:00 We're back into the program.

1:38:02 Thank you for the love.

1:38:03 Oh, thank you.

1:38:04 Thank you for the love.

1:38:05 We're shifting into one of the most important

1:38:07 patterns in Azure Cosmos DB in our next session.

1:38:11 One of my favorite teammates.

1:38:12 I know yours too, Justine Cocchi takes us

1:38:15 deep into the Azure Cosmos DB Change feed.

1:38:17 How it works under the hood,

1:38:19 the patterns that hold up at scale and the live demo of debugging and lag,

1:38:24 recovering processing and production is going to be great.

1:38:27 Yeah, you know Jay, Justine is amazing.

1:38:30 And if you love Justine.

1:38:31 Yeah.

1:38:31 And if you haven't used Change Feed, every write becomes a durable ordered event

1:38:36 stream that you can fan out to services, search and analytics,

1:38:39 event driven architecture without staying on a message bus.

1:38:43 So let's hear it all from Justine.

1:38:46 Here's mastering the Azure Cosmos DB Change Feed patterns,

1:38:49 scaling and real world architectures.

1:38:51 We'll see you soon.

1:38:56 Hi, I'm Justine Cocchi and I'm a program manager on the Azure Cosmos DB team.

1:39:00 I'm really excited to talk to you about Change Feed.

1:39:03 Today we'll cover what it is, how to read it,

1:39:05 and some tips for debugging for real world applications.

1:39:10 So first, what is the Change feed?

1:39:12 It's a persistent record of all changes to items

1:39:15 in your container that's ordered by modification time.

1:39:18 As your client app is making writes to your items,

1:39:21 you can then read them in the Change feed with your consumer app.

1:39:25 Your Change Feed consumers are processing changes per partition.

1:39:29 So let's take a look at what that looks like under the hood.

1:39:32 Because Cosmos DB is a distributed database,

1:39:35 your data is distributed across multiple physical partitions in a container.

1:39:40 In this example, I've got three physical partitions.

1:39:43 Each physical partition can have its own independent change feed

1:39:46 reader that's reading the changes in order for that partition.

1:39:50 Order is guaranteed within a logical

1:39:53 partition key in your physical partition range.

1:39:57 Your continuation token tracks progress of that change

1:40:01 feed processing for a given container.

1:40:04 And this means that each of your partitions

1:40:06 can have independent continuation tokens and independent processing.

1:40:09 This is great for efficiency because it really allows you to scale your chain

1:40:13 tree processing even for extremely large

1:40:16 containers with a high volume up writes.

1:40:20 There's a couple of different Change feed modes that you can choose from.

1:40:23 The default Change Feed mode and what's enabled

1:40:26 on all of your accounts is latest version mode.

1:40:29 In this mode you get the latest create or replace

1:40:32 for every item items that are deleted from your container.

1:40:36 No longer appear in the feed, and this means you won't get a notification

1:40:40 for deleted items in the change feed itself.

1:40:43 Any item that still exists in the container,

1:40:45 you can read the latest version of it.

1:40:47 There's infinite retention of all of these changes

1:40:49 in your container and you can go back and reading

1:40:52 even from the very beginning of your container

1:40:54 to get all of the items and their updates.

1:40:58 This is really great for a lot of change

1:41:00 feed scenarios and powers most change feed apps today.

1:41:03 There's also all versions in deletes mode.

1:41:06 With all versions and deletes, you get every create,

1:41:09 update and delete including TTL expirations.

1:41:13 Items for TTL expirations will appear in the feed in order of their purge time,

1:41:18 which may be slightly later than their actual expire time.

1:41:22 Retention of changes in this feed is based

1:41:25 on the continuous backup window for your account.

1:41:27 So either seven days or 30 days, depending on what you've configured,

1:41:31 you can read all of these changes within that window.

1:41:34 This mode is really great for audit logs or real

1:41:37 time processing where you need to react to every change.

1:41:40 Consider an application that is replace heavy where you have a high

1:41:44 volume of updates to a single item within a short period of time.

1:41:48 For all versions in deletes mode, you would get every single update in the feed,

1:41:52 whereas latest version you would only get

1:41:54 the latest version per change feed pull.

1:41:57 Depending on the needs of your application, each mode may be a better fit.

1:42:01 So it's important to consider what your app is trying

1:42:04 to do and choose the mode appropriate for your app.

1:42:09 There's a couple of different ways to actually consume the change feed,

1:42:12 and we're going to cover three of these ways.

1:42:15 The first is the change feed processor.

1:42:17 This works on a monitored container

1:42:20 or the container that you're reading your changes from.

1:42:23 When you create a change feed processor,

1:42:24 you can create multiple instances to handle scaling of all the partitions.

1:42:30 And this is the compute environment that you're

1:42:34 actually deploying your change feed processor on.

1:42:37 So you can imagine this like AKS,

1:42:39 where you'd have multiple pods that represent multiple instances.

1:42:43 The code that's actually being executed when a batch

1:42:46 of changes is processed is your handler delegate.

1:42:49 This is like the business logic or the brains of your change be processor.

1:42:53 Here is where you handle what you're reacting to for the change.

1:42:58 Maybe it's writing to something downstream or maybe you're doing some analytics.

1:43:03 This is really where you write your change feed

1:43:07 logic and the Change feed processor is available in both.

1:43:10 Net and Java.

1:43:11 Now, as your changes are being processed,

1:43:14 the lease container helps to Maintain state and checkpoints for keeping track

1:43:20 of where you are in processing for all of your various partitions.

1:43:24 Let's take a look at some lease

1:43:25 management and really understand how our physical partitions

1:43:29 map to leases which map to the instances

1:43:32 we've already established as a distributed database.

1:43:35 Cosmos DB can scale out to multiple physical partitions

1:43:38 and each physical partition owns a range of partition key values.

1:43:43 That top row is our physical partitions.

1:43:47 In this example we have six physical partitions

1:43:49 and each of them have their own lease.

1:43:52 Leases are always one to one with partitions

1:43:54 and this is managing the state of where we are processing.

1:43:57 In terms of the modifications to that individual partition.

1:44:02 You'll notice they each have their own

1:44:04 continuation token which is what maintains the state.

1:44:08 Now your Change Feed processor is deployed across multiple

1:44:11 instances and each instance owns a number of leases.

1:44:15 These are evenly distributed across instances to ensure

1:44:19 that as you're processing you can keep up with these changes.

1:44:23 Assume you only had one instance that's handling all of your leases.

1:44:27 If that instance goes down, it might impact the resiliency of your application.

1:44:31 However, if you have more instances than the number of physical partitions,

1:44:34 you would have idle instances that aren't able to do any work.

1:44:38 The Change Feed processor automatically rebalances leases across

1:44:42 these instances to ensure the maximum efficiency of your processor.

1:44:46 Let's say one of your instances goes down.

1:44:49 The leases that are owned by that instance will

1:44:51 automatically be rebalanced across all of the other healthy options.

1:44:56 The next way to read change feed is through the Cosmos DB trigger.

1:45:00 This is an Azure functions trigger that is

1:45:04 a really simple way to read change feed.

1:45:06 It's built on Change Feed processor under the hood.

1:45:08 So all of the lease and distribution of partitions

1:45:10 that we just reviewed still applies in the Azure functions trigger.

1:45:15 However, the management is done for you.

1:45:18 Because it's built on Azure functions, it's serverless.

1:45:21 It has built in auto scaling to dynamically react to the number of instances

1:45:26 you really should be having without over

1:45:28 provisioning instances that are not actively being used.

1:45:32 This is a really easy way to build the change

1:45:35 feed and the logic of your Azure function is like

1:45:38 the delegate in the Change feed processor where you're just

1:45:41 responsible for the business logic of processing the changes themselves.

1:45:46 The last option is the Change Feed pull model.

1:45:49 This model gives you the most control over how to read the change feed,

1:45:53 but also means you need to have all the safeguards put in place yourself.

1:45:58 So there's no lease container for maintaining state.

1:46:01 You instead use continuation tokens where you're

1:46:04 responsible for storing and maintaining those continuation tokens.

1:46:08 You're responsible for any error handling.

1:46:10 There's no built in retries.

1:46:13 You need to handle all of that in your client code.

1:46:16 While it does give you full control, it's also more code to manage.

1:46:21 One additional benefit of the pull model say you

1:46:24 wanted to only process changes for a specific partition range.

1:46:27 This is possible in the pull model because again,

1:46:30 you have full control over what is actually being pulled.

1:46:34 Now that we've got our Change Feed app,

1:46:36 let's talk a little bit about monitoring and how we can debug

1:46:39 some common failure modes to ensure that our Change Feed app is resilient.

1:46:45 Some of the key metrics to monitor for Change Feed are the estimated lag.

1:46:50 This measures the gap between the latest change that actually occurred

1:46:54 in your container and the last processed item in your change feed.

1:46:58 This is going to be your most

1:46:59 important metric for monitoring your change feed health.

1:47:02 While this estimated lag gives you a number of outstanding changes,

1:47:06 it is intended to be an estimate and it's not

1:47:08 always the exact number of changes that are actually pending.

1:47:12 The key thing to measure here is not exactly what the number is,

1:47:15 but the slope of how the number changes over time.

1:47:18 Is your lag steadily increasing?

1:47:20 Is it decreasing?

1:47:21 Are you staying constant?

1:47:22 This will really help you understand if your Change Feed app is healthy.

1:47:27 Some other things you can take a look

1:47:28 at is the throughput of your change feed processing.

1:47:31 You can look at the number of items that are processed

1:47:33 per second and the RU consumption of your change feed processor.

1:47:38 It's important to look at our use not only of your source

1:47:41 container that you're monitoring the change

1:47:42 feed of, but also the lease container.

1:47:45 If you have a lot of lease rebalancing

1:47:47 and redistribution that can put strain on your lease container,

1:47:51 which typically is provisioned with a relatively low number of RUs.

1:47:54 Monitoring these things will ensure

1:47:56 that your change feed application is healthy.

1:47:59 Digging deeper into the leases, you can look at the owner distribution.

1:48:03 Are your leases evenly distributed across instances?

1:48:06 Is the last checkpoint increasing for each of them,

1:48:09 or is one of your leases stalled with processing?

1:48:13 For setting up the monitoring itself

1:48:15 you can use application insights or OpenTelemetry,

1:48:17 which is built directly into the SDK to make it really

1:48:21 easy for you to export this telemetry for your change feed handlers.

1:48:26 For alerting.

1:48:26 Some things that you might consider creating

1:48:28 alerts for is sustained lag growth or if

1:48:31 there's any partitions that have stalled

1:48:32 and are no longer making progress in processing.

1:48:36 We can look at a couple key failure

1:48:38 modes and Some resilience patterns to help combat these.

1:48:42 The first is as Changes arrive, they're processed in batches.

1:48:46 When everything is going smoothly in the happy path,

1:48:50 your process handler succeeds and your lease checkpoints are all updated.

1:48:55 However, because these changes are processed in a batch,

1:48:58 what if there's an error?

1:49:00 Unhandled exceptions can create issues in your processing.

1:49:05 While the change feed processor does have retry,

1:49:07 you don't want to infinitely retry these poison messages

1:49:10 that are not able to be handled after multiple attempts,

1:49:14 and instead consider writing them to something like a dead letter queue.

1:49:17 This will allow you to write the item

1:49:20 that is failing to take a look at it later,

1:49:23 instead of blocking the entire processing for that container for that partition,

1:49:27 because your partition would not be able to continue

1:49:29 making progress if you're getting a consistent error.

1:49:34 There may be other causes of silent stalls for a given partition,

1:49:39 and one of the common causes is outgoing calls.

1:49:42 So if you're making calls to an outgoing service and they're failing,

1:49:45 this can also cause errors in your processor that will stall your change feed.

1:49:51 You can implement the circuit breaker pattern to ensure

1:49:54 that you don't flood this downstream service and inhibit,

1:49:57 your ability to actually recover from these errors.

1:50:01 A lot of the key resilience patterns

1:50:03 for any application also apply to change feed,

1:50:05 and it's important to keep these in mind

1:50:07 for the most resilient application and processing.

1:50:11 Now, we also talked a little bit about RU throttling.

1:50:15 For consistent RU throttling,

1:50:16 you can consider implementing priority based execution

1:50:20 to ensure that your mainline transactions on your container,

1:50:25 maybe your writes and your reads,

1:50:27 perhaps they have a higher priority than your change feed processor.

1:50:31 With priority based execution, you can specify that in the client.

1:50:35 To ensure that the change feed is not

1:50:37 accidentally consuming extra RUs that you'd maybe prefer,

1:50:41 go to the rest of your application.

1:50:44 There's also throughput buckets that will

1:50:46 help you indicate the percentage of RUs that you want to distribute across

1:50:50 your various applications and including your change feed.

1:50:55 Now, while you may see ru, throttles across your entire container,

1:50:59 it's also important to consider hot partitions because the change feed

1:51:04 works at the partition level and your progress is per partition.

1:51:08 If you have a surge of writes that is generating a hot partition,

1:51:12 it may be difficult for your change feed to keep up.

1:51:15 Ensure that your leases are appropriately balanced across partitions,

1:51:19 across instances, so that if there is a hot partition,

1:51:23 it doesn't overload one of your instances

1:51:25 and result in higher lag across the board.

1:51:31 All right, let's take a look at a demo to see how the Change

1:51:35 Feed processor is actually configured and see

1:51:37 it in a real application application,

1:51:38 I'm going to pull up VS code with a simple app here.

1:51:42 I've got a social media app simulator which is really just two console apps.

1:51:48 My first console app is simulating users creating,

1:51:51 editing and deleting some of their posts.

1:51:54 And the program that we've got on screen here is

1:51:56 a second console app that has our Change Feed processor.

1:52:00 Now, because we know that we've got updates and deletes

1:52:02 and I want to react to all of them,

1:52:05 I'm going to use the Change Feed processor with all versions and deletes mode.

1:52:09 When I configure my Change Feed processor, I can choose a processor name.

1:52:14 This tells me the unique processor instance and will allow me

1:52:19 to coordinate multiple compute instances all to the same processor application.

1:52:23 This is really critical for least rebalancing

1:52:25 because if you're using different processor names,

1:52:28 the change feed will assume that it's four actual different,

1:52:33 change feed instances and different purposes.

1:52:37 So we've got the processor name which sort of ties all

1:52:39 of our instances together and we also have our instance name.

1:52:43 We'll take a look in the console app how the instance name allows

1:52:46 us to tie multiple different processes

1:52:48 of these together into that same processor.

1:52:52 I'm using the same lease container for all

1:52:54 of these, which is storing all those checkpoints.

1:52:57 And I'm also setting up a couple of notifications.

1:53:00 These lifecycle notifications are really important

1:53:03 for debugging the lease ownership in your application.

1:53:06 And you can get notifications for lease acquired and lease released,

1:53:11 as well as printing out some, error messages that may occur.

1:53:15 So this is the Change Feed processor itself.

1:53:18 But let's also take a look at the delegate.

1:53:20 Once we get a batch of changes,

1:53:23 this is the code that will execute and actually process those changes.

1:53:27 So first we'll simulate some extra processing with a short delay.

1:53:33 And of course, in your real application

1:53:34 this is whatever business logic you may have.

1:53:37 Then we can process our changes

1:53:40 slightly differently depending on the operation type.

1:53:42 The first thing we want to check for is any deletes.

1:53:45 The ID and partition key of deleted items

1:53:47 will come through in the metadata so we

1:53:50 can pull out those key pieces of information

1:53:52 to be used later for creates and replace operations.

1:53:56 We can directly get the ID and the user id,

1:53:59 which is our partition key from, the current aspect of this change.

1:54:06 Scrolling down a little bit further,

1:54:08 this change feed is writing to a notifications container

1:54:13 which then can power notifications for our social media app.

1:54:17 So now that we Looked at the code, let's see it running in action.

1:54:21 And I've got three separate terminals pulled up here.

1:54:24 The first thing I'm going to do is start running my social simulator.

1:54:27 This console app is going to be creating

1:54:30 items and posts from users in our application.

1:54:33 Then I'm going to run my notification processor.

1:54:36 Now notice I'm not submitting any arguments here,

1:54:39 it's just that basic, code that we talked about.

1:54:43 And it will have a default instance name which we'll see spin

1:54:47 up here in just a second as, the detailed stats come online.

1:54:52 Now in the meantime,

1:54:53 I'm going to start a stream of posts in my application and let's make sure

1:54:58 our processor is printing in verbose mode so

1:55:00 we can see all those creates coming in.

1:55:02 But we're using all versions and delete.

1:55:04 So let's simulate some delete events and some edit events and you

1:55:08 can see all of those are actually flowing through into our processor instance.

1:55:14 All right, let's turn it back into quiet mode so we don't have every single,

1:55:18 change, but just the batches of changes that are coming in.

1:55:21 And let's simulate a couple of viral bursts.

1:55:25 There's a lot of buzz on our social app and users are making

1:55:29 more posts than ever as we see these viral bursts start to flow through.

1:55:34 We see the lag is steadily increasing on our change feed processor instance.

1:55:40 So we can create a second instance to help us,

1:55:43 manage all of the changes that are coming through for all

1:55:46 of these partitions and balance it a little bit better across both of these.

1:55:51 So I'm going to print out the detailed stats for our second instance.

1:55:57 And you notice that this one is named Instance 2 as it starts up.

1:56:02 We start by acquiring Elise and we get an error on our initial prefacer.

1:56:07 Now this is really important because this is not actually a bad error.

1:56:12 This error is just telling us, hey, I lost a lease and it went to someone else.

1:56:16 This is the behavior that we actually

1:56:18 want change Feed processor is automatically rebalancing

1:56:21 leases for us so that we can

1:56:23 ensure they're evenly distributed across our processors.

1:56:26 If we print out the leases that are actually owned by each of these, we can see,

1:56:31 our first instance owns three of our leases and our second instance owns two.

1:56:36 This is great because this means that our processors

1:56:39 are working properly and our load is evenly distributed.

1:56:43 Now I'm going to go ahead and shut back down our second

1:56:46 instance and we should see these leases

1:56:49 redistribute back onto our initial processor.

1:56:52 We see those lease acquired messages now and when we print the leases

1:56:56 owned this time we see all five leases are back onto our main processor.

1:57:01 Let's go ahead and shut down this application

1:57:04 and we see how easily Change Feed allows us

1:57:08 to spin up new instances and dynamically react

1:57:11 to the volume of processes and changes coming into our container.

1:57:18 Flipping back into slides now that we've shown a basic Change Feed app,

1:57:22 let's take a look at some advanced patterns.

1:57:25 We can combine the event sourcing CQRs and materialized views

1:57:29 pattern and to create a really powerful application in Cosmos DB.

1:57:33 While our demo showed creates updates and deletes

1:57:37 inline to items with the event sourcing pattern,

1:57:41 every write, every create, edit,

1:57:43 delete actually becomes its own independent create event.

1:57:46 So it's a depend only event store.

1:57:49 We can write these commands into our event

1:57:52 store still stored in Cosmos DB as container and we can use Change Feed to read

1:57:57 this ordered log and populate some downstream systems.

1:58:01 These are also known as materialized views and it helps us craft a view

1:58:06 of our data that is better suited for the read patterns of my application.

1:58:13 In my app I've got a couple of read patterns that I want to support,

1:58:16 like show me my notifications.

1:58:18 This is that notifications container that we

1:58:20 were just populating in our example.

1:58:22 But let's say we also want to get posts by topic.

1:58:26 Or maybe we want to create a search index container.

1:58:29 We can use built in full text search with Cosmos DB

1:58:33 to create a second container that has full text policies on our content.

1:58:39 This allows our users to really easily search across those data patterns too.

1:58:45 With Materialized Views pattern.

1:58:48 Setting up all of these separate containers based

1:58:50 on the read pattern that we're trying to optimize really

1:58:53 gives us the best configuration for Cosmos to give

1:58:57 us these efficient reads without impacting our write path.

1:59:01 Our write path is still optimized for the writes.

1:59:05 Now the Materialized Views pattern is really common in many databases,

1:59:09 but I do want to highlight a Cosmos

1:59:11 DB feature that makes this pattern really easy.

1:59:14 We have a feature called Global Secondary Indexes which

1:59:17 are effectively managed materialized views directly into Cosmos DB.

1:59:22 So while I just showed you how you can

1:59:24 kind of build this yourself with the Change feed.

1:59:27 Reading the change feed of social events,

1:59:29 populating a new container with a new partition key,

1:59:32 you can add a global secondary index.

1:59:35 This will do the exact same thing for you,

1:59:37 but it will automatically maintain the secondary container without you

1:59:41 having to write the code to build and manage changepie yourself.

1:59:45 This is a really powerful container,

1:59:48 this is a really powerful feature that allows you

1:59:50 to build containers specialized to the read patterns that you have.

1:59:54 My container, I had five physical partitions in my source.

1:59:57 If I were to use my source container to serve this query by topic,

2:00:01 it would be really inefficient,

2:00:03 it would be slow because I'm checking every physical

2:00:05 partition and it also would cost a lot of RUs.

2:00:09 With GSI I'm able to automatically create this and I

2:00:13 don't have to manage the change feed myself.

2:00:16 So let's take a look at the Azure portal

2:00:18 and we can see how this is actually implemented.

2:00:22 I have my Change Feed application,

2:00:25 my Cosmos DB here that was powering that Change Feed app we

2:00:28 just looked at and you can see my source container is social events.

2:00:34 Now my source container is partitioned on user ID and we

2:00:37 see all of the posts that my users are making here.

2:00:41 But I added a GSI already.

2:00:43 Let's take a look at our GSI and we can see this has

2:00:47 a persistent copy of all of the items in our source container,

2:00:50 but now it's partitioned by topic

2:00:52 and this is automatically maintained by the Change feed.

2:00:56 We ensure any writes only go to social events.

2:00:59 We don't need to handle any dual write or error handling.

2:01:01 That's all given to us by the platform

2:01:03 and we're able to have this read optimized view.

2:01:07 Now let's take a look at a query.

2:01:10 I want to get the top 100 items where the topic is Cosmos Conf,

2:01:15 and this first query is against my social events container.

2:01:18 If I look at my query stats I can see this costs almost 27 RUs.

2:01:23 So that's quite a bit of RUs to get these topics.

2:01:27 But if I issue that exact same query against my global secondary index,

2:01:32 flip over to query stats.

2:01:34 I see this costs just about five and a half RUs.

2:01:36 So this gives us an roughly 80% savings.

2:01:40 All because we're using the GSI which

2:01:42 is partitioned and configured for our read pattern.

2:01:50 This demo really shows us how you can use Change Feed

2:01:55 and manage features to give you

2:01:57 the best configuration for your specific application.

2:02:01 All right, we covered a lot today, so the key takeaways are that Change Feed

2:02:06 really powers a lot of event driven applications.

2:02:09 There's many architecture patterns that help

2:02:11 you build these apps for your business.

2:02:13 Use case.

2:02:15 When you're using the Change Feed, consider the Change Feed mode that is best

2:02:19 suited for your application and also the consumption model.

2:02:22 Whether that's Change Feed, processor Azure functions or the pull model,

2:02:26 it's important to invest in monitoring and you can use

2:02:30 the estimated lag to ensure that your processors are Keeping up.

2:02:34 Keep resiliency best practices in mind in your handlers.

2:02:38 Ensure that they're idempotent as retries are happening and their resiliency

2:02:43 is all the best practices with circuit breaker retry timeouts, et cetera.

2:02:49 Advanced features like Global secondary Index really help

2:02:52 you simplify these patterns and let the platform

2:02:54 do the heavy lifting instead of you

2:02:56 needing to manage that Change Feed application yourself.

2:02:59 Now you can get started and learn even more about Change Feed at this link,

2:03:03 aka MsAzureCosmos DB ChangeBead.

2:03:07 Thank you so much and I hope you enjoy the rest of the conference.

2:03:13 Okay.

2:03:14 Hello everybody.

2:03:15 So, this is MultiCloudDB.

2:03:17 Write once, run anywhere.

2:03:19 I'm Theo van Kraay, I'm a PM in the Cosmos DB team.

2:03:23 I work on SDKs, connectors,

2:03:26 developer experience and lots of fun stuff like that.

2:03:29 But I'm going to be doing something very different this time.

2:03:32 Probably the first time we've done anything like this, on Cosmos DB Conf.

2:03:37 I'm going to be talking about something that we're calling MultiCloudDB.

2:03:41 This is a unified data access layer for best of breed cloud databases.

2:03:47 So for the first time ever we're talking about databases other than Cosmos DB.

2:03:52 But before I go into that, before I talk about what that does,

2:03:55 I want to set some context here.

2:03:58 So why teams still want cloud native databases?

2:04:03 Of course, databases like Azure, Cosmos DB, Amazon,

2:04:07 DynamoDB, Google Spanner, are still very highly sought after.

2:04:11 They offer high availability and elasticity built

2:04:14 from the ground up in those platforms.

2:04:16 These platforms are built to exploit cloud core,

2:04:22 properties in those environments.

2:04:24 They have strong operational guarantees and SLAs and manage scaling.

2:04:28 It's simply not possible to get this same kind of performance

2:04:32 if you're installing a traditional database software and putting that onto

2:04:37 VMs for obvious reasons and of course you have a fast

2:04:39 path to production without owning the database or the control planes.

2:04:44 Everything is very easy, it's very elastic,

2:04:46 it's very scalable, it's very highly available.

2:04:50 The benefits there are pretty obvious.

2:04:52 But teams tend to want managed database upside,

2:04:55 without the irreversible coupling that you tend to get.

2:04:59 There's a tension here.

2:05:01 The deeper that you adopt native SDKs of these types,

2:05:04 of databases and query models, the harder it becomes to move anywhere else.

2:05:13 So this becomes something that is concerning for customers,

2:05:17 if they want to exit the architecture or move into another cloud

2:05:22 and Portability concerns show up pretty much before any line of code is written.

2:05:27 And so sadly for us working on Cosmos

2:05:29 DB&M for our counterparts in Amazon and Google,

2:05:33 often these very good database services get excluded

2:05:37 right off the bat because there's no portability,

2:05:40 they only run in Azure or Amazon or Google etc.

2:05:45 And so the cloud native benefits are obviously significant,

2:05:48 everybody knows that, but so is architecture level lock in and risk.

2:05:52 And this is very real and we are obviously recognizing this.

2:05:56 So the customer reality is that customers do want

2:06:00 the benefits of cloud native databases as we've said,

2:06:02 but they also need a story for portability,

2:06:05 procurement and regional strategy and so on.

2:06:08 This is the gap that we're trying to fill with MultiCloudDB,

2:06:11 and it's designed to close.

2:06:13 So portability is no longer optional for teams

2:06:17 who are building these types of applications,

2:06:20 operating across regions, customers and different clouds and so on.

2:06:23 We see two different types of multi cloud

2:06:27 portability pressure if we can put it that way.

2:06:29 The one is cross cloud deployment.

2:06:31 So this is where you have the same solution delivered into multiple clouds,

2:06:35 typically white labeled products,

2:06:37 sovereign deployments or maybe partner managed environments,

2:06:40 where the application needs to behave consistently even when

2:06:44 the backing database changes by the customer or region, et cetera.

2:06:49 And then the other one maybe is more familiar is where you

2:06:53 want to deploy into one cloud but you want a credible exit strategy.

2:06:57 In other words you want to avoid cloud vendor locking at least.

2:07:02 And portability obviously matters there as well.

2:07:05 Even if the team never actually switches cloud providers.

2:07:11 In both cases the requirement is the same.

2:07:13 Application code doesn't want to have to be

2:07:16 rewritten if you're moving into a different cloud.

2:07:19 And again sadly for us working Cosmos DB this means a great database product

2:07:23 is often excluded right out of the bat because it only runs in Azure.

2:07:32 So there are some abstractions out there in the Java world you have things

2:07:37 like Spring Data and Hibernate of course

2:07:39 and the Net world you have Entity Framework,

2:07:41 Core and Python you have things like django etc.

2:07:46 That offer something close to this.

2:07:47 So you'll get something like a programming to an interface model.

2:07:51 In Spring Data you have repository abstractions and object mapping and so

2:07:55 on and you get something like a productive lowest common denominator.

2:08:00 The problem with this is that they're

2:08:04 not based strictly on portability contracts.

2:08:06 This is a side effect of developing an abstraction that is Meant to simplify

2:08:14 development on many different databases rather

2:08:17 than it being the core design goal.

2:08:20 So what you get over time is divergence still appearing.

2:08:24 You still get query features drifting by backend annotations and mappings code,

2:08:29 level changes that you need to apply even though

2:08:32 more or less things tend to be very similar.

2:08:35 So portability is sort of possible in part but it's not guaranteed by design.

2:08:39 And in practice this shows up when you

2:08:42 want to migrate from one cloud to another.

2:08:43 You still have in some cases almost as many changes if not quite as many.

2:08:49 And divergent still accumulates because as I

2:08:51 said this is not the core design goal.

2:08:55 It's a side effect of some principles to create an abstraction that you

2:08:59 get some level of portability but it's

2:09:01 not really completely portable in that sense.

2:09:06 So why MultiCloudDB and what we've built

2:09:08 here and what we are building is different.

2:09:11 We have a strict portability design goal from the ground up.

2:09:15 This is not a side effect capabilities explicitly

2:09:18 surface what is and what is not portable.

2:09:21 The promise that we're making is write once,

2:09:24 run anywhere at semantics MultiCloudDB

2:09:28 treats portability as a product guarantee,

2:09:31 not a best effort convenience that you get

2:09:34 as a side effect of some goals of abstraction and whatnot.

2:09:39 So again design goal strict portability same code

2:09:43 and query code should run unchanged across Cosmos DB,

2:09:47 DynamoDB and Google Spanner should be portable by default.

2:09:50 Transparent limits.

2:09:52 We will be providing some native escape hatches

2:09:54 but we don't expect customers to be using this.

2:09:58 The whole point of this abstraction is that you shouldn't be using those because

2:10:03 that will obviously break portability and that's

2:10:05 what you want if you're in this space.

2:10:08 So simple architecture view here you have a Java application

2:10:11 that's going to be connecting to multi CloudDB client and then

2:10:13 you have a service loaded discovery that's going to load

2:10:16 the appropriate module that is being driven by config files only.

2:10:21 And this is set very similar to what you get out of things like Spring,

2:10:24 and Hibernate and so on.

2:10:26 But again as we've said this is portability from the ground up and strict

2:10:30 guarantees around not only the feature set

2:10:34 but the behavior of those features as well.

2:10:37 And it follows from that we need a portable query dsl.

2:10:40 So you can see on the left there this is MultiCloudDB query dsl the same

2:10:45 regardless of the platform that's running on it

2:10:47 would be different on the right hand side,

2:10:49 syntax differences between those different databases and so on.

2:10:53 It also follows that we have a fairly conservative feature roadmap.

2:10:58 Bottom line is example a feature will not

2:11:03 be supported if all providers cannot support it.

2:11:06 So if it can't be supported in every

2:11:08 available provider then it's simply not shipped.

2:11:13 So you might be wondering at this point why

2:11:15 would we be crazy enough to do something like this?

2:11:20 And the answer is of course we've done this before.

2:11:22 Those of you who are Familiar with Cosmos DB,

2:11:24 Azure Cosmos DB was the first database platform

2:11:27 to build broad compatibility APIs at this level.

2:11:31 Those of you who've used MongoDB API, Cassandra API,

2:11:34 Table API, Gremlin API and so on, these were compatibility APIs,

2:11:40 APIs sitting on top of Cosmos DB,

2:11:42 providing the surface area of a different database along on the same platform.

2:11:47 This gives us years of real world experiences handling compatibility drift,

2:11:52 behavior gaps, edge case semantics, et cetera.

2:11:55 We've learned that how portability fails in practice

2:11:58 where feature gaps create friction and so on.

2:12:01 How to set clear contracts and manage the relevant trade offs et cetera.

2:12:07 Why this matters.

2:12:09 Well this is not our first rodeo as my American

2:12:12 colleagues would say we're applying hardened compatibility discipline

2:12:16 to multi cloud DB from day one and we

2:12:20 actually think this is the easier problem of the two.

2:12:24 Compatibility APIs are hard.

2:12:26 Fitting a database engine to a programmability surface

2:12:30 area that we don't control is well difficult.

2:12:34 Let's say surface area keeps expanding,

2:12:37 behavior contracts keep shifting, goalposts move continuously.

2:12:41 It's very difficult.

2:12:43 Now I'm not going to say that multi clouddb is easy

2:12:45 but certainly from our perspective given

2:12:47 our experience it's a more tractable problem.

2:12:51 We can constrain the programmability layer that we control across

2:12:56 a focused set of databases and therefore we define the contract,

2:13:02 the portability of the contract.

2:13:03 We scope that to known database targets and we

2:13:06 can enforce limits explicitly so we have more control.

2:13:09 Even though there are challenges, will be challenges around supportability.

2:13:14 It's certainly a more tractable problem than the one that we

2:13:17 have already solved in the past with Cosmos DB nonetheless.

2:13:23 So of course where this fits best,

2:13:27 of course if your most important things that you want to solve

2:13:31 that you want the benefits of cloud native databases, the elasticity,

2:13:35 the high availability guarantees,

2:13:38 the SLAs and so on the managed service aspect but you also

2:13:42 want Portability then this is going to be a great fit for you.

2:13:46 Of course there are trade offs to accept as there are everywhere.

2:13:49 Strict portability means some trade offs.

2:13:52 Of course you are going to get access to more granular features if you're

2:13:56 using the native SDKs for Cosmos DB and DynamoDB and Google Spanner and so on.

2:14:03 So bottom line you want to be choosing

2:14:06 this when portability and the benefits of a cloud

2:14:09 native database are really the higher order

2:14:11 bit for your architecture and for your strategy.

2:14:15 So then again ultimately the key takeaways,

2:14:18 managed cloud databases are still very,

2:14:20 very attractive to many customers for some very good high availability,

2:14:25 elasticity, performance maintenance, operational guarantees.

2:14:28 The list goes on.

2:14:29 The blocker isn't really the database capability, far from it.

2:14:34 It's of course the application level coupling to one provider's SDK.

2:14:38 It's lock in.

2:14:39 We don't like to say the words lock

2:14:41 in but that's really what we're talking about here.

2:14:43 And MultiCloudDB is designed to blow that open by having

2:14:50 a strict portability design from the ground up so teams can

2:14:54 keep an exit strategy and deploying cross cloud scenarios without giving

2:14:59 up the massive benefits of modern managed cloud databases like Azure,

2:15:04 Cosmos DB and our counterparts in Amazon and Google and so on.

2:15:10 So before I do a demo I just want

2:15:11 to mention we do have some documentation and you'll see there,

2:15:14 it says at the bottom there in small print.

2:15:16 The SDK is currently available as a public preview.

2:15:19 It is not yet fully ready for production use.

2:15:22 Expect breaking changes, incomplete features, limited support during this phase.

2:15:28 But we definitely will encourage you to try this out,

2:15:31 raise issues in GitHub, ask questions, make feature requests, et cetera.

2:15:36 We want to know how you want to use this, what this is solving for you.

2:15:39 We're definitely committed to it.

2:15:42 So yeah, check out the documentation.

2:15:43 There's a nice getting started guide, a developer guide as well.

2:15:47 It's fully comprehensive with everything that we have so far

2:15:50 and obviously we're going to be

2:15:51 developing on this, information about compatibility,

2:15:53 a couple of samples that you can get

2:15:55 your teeth into, an API reference and so on.

2:16:00 All right, so without further to do, let me set up my demo here.

2:16:05 Okay, so here we are.

2:16:06 In VS code I have an application called Risk Platform App.

2:16:11 It's a multi tenant app that is giving a dashboard

2:16:14 of different applications that are providing risk management services and so on.

2:16:21 And this is built on Multi Cloud DB and you

2:16:24 can See the multi cloud client is being used here,

2:16:28 and that is going to do a bunch of things and one

2:16:31 of the things it's going to do is load a config file.

2:16:34 And again I mentioned earlier,

2:16:35 the configs are the only thing that are different per provider.

2:16:39 So you've got one for Cosmos DB here

2:16:41 and one for DynamoDB and everything else is identical,

2:16:45 the code is the same regardless of which provider you're pointing to.

2:16:48 You don't need to change it.

2:16:49 You can treat it as basically the same database

2:16:52 and we guarantee that the behaviors and the functionality

2:16:55 that is available to you should be the same

2:16:58 and performance differences within some

2:17:00 acceptable limits of a statistical difference.

2:17:04 So let me just start up my Cosmos DB,

2:17:09 version here with the Cosmos DB properties that I have.

2:17:12 So let's go ahead and start this one up.

2:17:17 Okay, so that's started up.

2:17:18 So let me go into this app here and we can see I've just put

2:17:23 a little moniker there just based on the config

2:17:25 file to denote which database it's pointed to.

2:17:27 But everything else should be identical if we were going to run this in Amazon.

2:17:33 And so I can see here I've got different portfolios,

2:17:36 I click on these, I'm going to see

2:17:39 a different list of positions on those portfolios.

2:17:41 So just to prove it let me go back here and let me cancel

2:17:45 this app and let me start it up

2:17:49 again but use the Dynamodb properties files instead.

2:17:57 Okay, so that started up, now you can see there provider, DynamoDB.

2:18:01 Instead of just clicking on the link I'm just

2:18:03 gonna go straight back to my original link here

2:18:05 and I should able to refresh or hit that URL

2:18:08 and you see it's just now pointing to Dynamodb.

2:18:11 But everything else is exactly the same.

2:18:14 The data is the same and all the behaviors should be identical.

2:18:20 And if we were to look in the portal for Cosmos DB,

2:18:23 now of course this is where things may look different.

2:18:25 So you'll have databases and containers in Cosmos

2:18:28 DB for those of you familiar with Cosmos DB.

2:18:31 In Dynamo of course you have tables but you don't have databases.

2:18:35 So we prefix this with the database

2:18:37 name in order to replicate the resource hierarchy.

2:18:40 And so that's from the SDK perspective that's what it's abstracting.

2:18:43 But of course the data is ultimately identical

2:18:45 and most importantly of course how the SDK interacts

2:18:48 with the data and the results that come

2:18:50 back and the behaviors should be identical as well.

2:18:53 And ultimately that's what you will care about

2:18:55 if you want to maintain a single playbook,

2:18:57 a single set of application code for your solutions

2:19:00 that are being deployed into different cloud environments,

2:19:02 or if you want to make sure that you

2:19:05 have an easy migration and exit strategy from one,

2:19:08 cloud environment to the other.

2:19:12 So that's it.

2:19:13 That's everything for MultiCloudDB again, check out the documentation,

2:19:20 and the links that we gave here earlier on.

2:19:23 Thanks for watching.

2:19:30 Hey everybody.

2:19:30 My name is Andrew Ruffin.

2:19:31 I'm a senior Product Manager here at AMD and I

2:19:34 focus on launching and growing new products and services on Azure.

2:19:36 With AMD under the hood, you're here at Cosmos DB conference most likely because

2:19:41 you care about building systems that actually work at scale.

2:19:43 They're globally distributed, they're highly available,

2:19:46 and they're predictable under real workloads.

2:19:48 Azure Cosmos DB gives you all of that as a managed service,

2:19:51 and that abstraction is incredibly powerful, especially as a managed service.

2:19:55 This all comes by design.

2:19:57 They come from real engineering decisions, joint enablement,

2:20:00 and the deployment of actual infrastructure across

2:20:02 the global scale that only Azure can provide.

2:20:05 The hardware choice isn't your concern,

2:20:06 but the physics absolutely stick around and they matter.

2:20:10 And that's where AMD comes into the picture.

2:20:12 Today, Azure Cosmos DB runs on modern AMD EPYC processors across the globe.

2:20:17 We launched AMD EPYC in the cloud on Azure nearly 10 years ago,

2:20:21 and we've continuously executed and generated new products and services

2:20:25 that give the features that you need generation over generation.

2:20:29 You'll hear more details on how Azure

2:20:31 Cosmos DB takes advantage of AMD infrastructure,

2:20:33 including the latest 5th gen processors later on.

2:20:37 While you, as an Azure Cosmos DB user, don't need to spend any cycles

2:20:40 worrying about the hardware or the infrastructure.

2:20:42 You get the benefits in terms of cost

2:20:44 savings enabled by efficient and sustainable performance across Azure,

2:20:48 you'll find benefits when you're running on top of modern EPYC processors.

2:20:52 The newest V7 generation, for example, provides up to 35% more performance

2:20:56 in performance per dollar over the prior generation.

2:20:59 With these newest generations integrated into Azure Cosmos DB,

2:21:02 you get more performance per request unit

2:21:04 and a better experience to the many dynamic scaling

2:21:07 and elasticity features that the team continues to roll

2:21:09 out and you hear more about later today.

2:21:12 This is all enabled by the close partnership

2:21:13 we at AMD have with Microsoft as a whole,

2:21:15 and it creates a feedback loop between hardware and software development.

2:21:19 This results in optimized performance per dollar

2:21:21 per watt for the most advanced workloads,

2:21:23 like the ones you run every day on Azure Cosmos DB.

2:21:26 So thank you to the Azure Cosmos DB

2:21:28 team and all the organizers of the conference.

2:21:30 We appreciate the opportunity to help put

2:21:32 this all together and make this possible.

2:21:33 Enjoy the show.

2:21:36 Welcome back and a quick thank you to Andrew Ruffin and our partners

2:21:40 at AMD for being part of the Azure Cosmos DB conf.

2:21:43 Yeah, absolutely.

2:21:45 They were a key part of how we were able to extend this show to five hours,

2:21:50 more sessions, more speakers and more depth.

2:21:53 Quick reminder, you can explore the full agenda of all those talks at AKA Ms.

2:21:59 Azurecosmos dbconf.

2:22:00 All of the talks.

2:22:02 All of them.

2:22:03 All of them.

2:22:03 They're going to be live.

2:22:04 Tell them about it.

2:22:05 Yeah, there's going to be $0.21 sessions total and every

2:22:08 session will go live before the end of today's show.

2:22:11 You can find all of today's links up in the YouTube description, a lot of links.

2:22:16 So you're going to be looking past the comments.

2:22:19 So, we're going to talk a little bit about agents AI,

2:22:22 but before we do that we've got some comments

2:22:24 from you that have to do with agents in AI.

2:22:27 So let's take a look.

2:22:28 So, one of our viewers, a very, very engaged viewer, Chenola,

2:22:34 has said I asked the chat and I'd love you to answer those questions too,

2:22:39 what your AI agent development, environment may look like.

2:22:44 They said we are using GitHub Copilot at the enterprise level.

2:22:48 We've developed our agents to run widely with Copilot CLI,

2:22:53 using it for code test DevOps.

2:22:56 That's huge.

2:22:57 It's an end to end kind of experience.

2:22:59 And then Sandra says, I am a DB admin who usually uses the Azure

2:23:05 Portal for DB management and Bicep templates for deployment.

2:23:09 So they do have a copilot with PowerShell, for doing that automation.

2:23:13 So.

2:23:13 Hey, Patty, I hear you also got

2:23:15 another question about something related to AI coding.

2:23:18 Yeah, also from St.

2:23:20 Andrew.

2:23:21 First of all, thank you so much for the great comment on the keynotes.

2:23:24 Now the question is, about the agent skills.

2:23:27 Does it only work with the extension?

2:23:29 No, you can use NPX to install it.

2:23:34 Installation and usage instructions should be in the documentation.

2:23:40 Yeah, but what do we have, Jay?

2:23:42 Absolutely.

2:23:43 And up next, when you're building with AI agents,

2:23:46 token usage isn't just a billing detail, it's a design construction.

2:23:50 And so every extra turn and chunk of context adds up.

2:23:55 So you need an approach that stays efficient as your agent scales.

2:23:59 And I know I may get access to lots and lots of tokens,

2:24:03 but not everyone else does.

2:24:05 I know efficiency really is key and that's where memory becomes critical.

2:24:11 Farah Abdou will show you how an agent memory

2:24:13 layer on Azure Cosmos DB helps you reuse Cosmos context,

2:24:16 reduce tokens and keep response quality high while improving

2:24:20 latency and keeping costs as predictable as you scale.

2:24:23 Absolutely.

2:24:24 And I love this session.

2:24:26 Farah is so great.

2:24:27 This is cutting AI agent costs with Azure Cosmos DB, the agent memory fabric.

2:24:33 Here's Farah.

2:24:36 Hello everyone, my name is Farah Abdou.

2:24:39 I'm a lead machine learning engineer and AI researcher.

2:24:43 And today I will walk you through how

2:24:45 to cut AI agent costs with Azure Cosmos DB.

2:24:50 This is a war story about building an AI system

2:24:54 that costs way too much money and how I fixed it.

2:25:00 So for the agenda today you are going to talk about the problem.

2:25:04 Why is the multi agent AI based model the solution?

2:25:08 We have one database, three features,

2:25:11 then a live demo and the real production numbers.

2:25:17 So global AI spending in 2026 will be $2.5 trillion.

2:25:25 The agent market grows from 7.84 billion to 52.62 per year.

2:25:34 This is between 2025, 2030, a compound of annual growth,

2:25:40 rate of 46.3% and 75% of large enterprises will be running AI agents by 2026.

2:25:51 And inference demand will grow about a, thousand times by 2027.

2:25:57 And this will keep that host law is not an optional thing.

2:26:04 It's how we can stay in the business.

2:26:12 So this is the architecture that everyone is building right now.

2:26:16 Store data systems, a cache, a relational database,

2:26:19 a vector database and event pass.

2:26:22 We have four terms for this, four filler points,

2:26:25 we have dual coordination between them.

2:26:28 And this is exactly what I built first.

2:26:30 It worked, but it didn't work after this.

2:26:39 So what about the production?

2:26:43 So in the demo, first of all we have three agents,

2:26:48 we have 100 requests, it's about six donors.

2:26:53 Then we come to production.

2:26:54 In production we have now three agents,

2:26:57 still our agents, but now we have 10,000 requests.

2:27:03 This will cost us $18,000 a month.

2:27:06 Nothing changed about the design, we just got real traffic.

2:27:15 So we have five filler modes.

2:27:17 I have seen with my own eyes, number one is about the cost exclusion.

2:27:23 We have from 5 to $15 in a demo becomes 18,000 to $90,000 a month in production.

2:27:33 Number two is the reliability collapse.

2:27:36 So 95% per agent multiplied across three agents equals for example like 85.7%.

2:27:45 This is 14:30 every day at 10,000, huge number.

2:27:55 So number three is about the latency cascade.

2:27:59 For example agent one let's say take three seconds.

2:28:02 Agent B take about four seconds.

2:28:05 Agent C take about five seconds.

2:28:09 There are 12 seconds total plus 200 to 500 millisecond coordination between

2:28:15 the handoff and for the state corruption

2:28:20 like parel rights with no concurrency control.

2:28:24 It means silent data and 5 is the context degradation.

2:28:32 Agency gets only 60% of agents a original entity after 47 steps.

2:28:45 So what is the reliability mass that kills you?

2:28:49 One agent at 95%, 500 filler a day at 10,000 rubles.

2:28:59 Three agents at 85.7%.

2:29:03 1,430 filler per day.

2:29:08 We have five agents at 77% 2,300 Philip and they're gonna

2:29:16 say 40% or more of agent project will fail by 2027.

2:29:22 So look at this math.

2:29:23 I think huge numbers the little Pill now

2:29:30 that nobody plants are Winterbeat analyzed 100,000 production queries.

2:29:38 18% were exact duplicates.

2:29:42 47% were semantically similar.

2:29:45 Same meaning, different words.

2:29:48 Only 35% were genuine new questions.

2:29:51 When I run the same analysis on my own system,

2:29:55 about 65% of my LLM Spent was answering question I already answered.

2:30:05 So the breaking point is $47,000 a month.

2:30:11 40% of the sprint time debugging agent fillers.

2:30:15 Multi agent debugging takes three to five

2:30:19 times longer than the single agent debugging.

2:30:22 And only 11% of the organizations have AI agents in production at all.

2:30:29 We were in that 11% but it didn't feel like success.

2:30:34 Something now had to change and this is why we are here today.

2:30:41 We have one database, three capabilities.

2:30:44 The agent memory subgroup which is built entirely

2:30:49 ON Azure Cosmos DB for NoSQL this replaces four systems.

2:30:54 Cut the cost 73%, cut the latency 65% and let

2:31:00 me show you now how So why Azure Cosmos?

2:31:07 Because we have four properties that make

2:31:11 it uniquely suited for the multi agent footprint.

2:31:14 First, the native vector search using DiskANN sub 20 milliseconds

2:31:20 at 10 million vectors under 100 milliseconds at one period.

2:31:26 The second is a change feed like

2:31:29 the built in event streaming no message queue needs.

2:31:33 And the third is the ETag optimistic concurrency at the document

2:31:39 level no distributed op manager is here and the global

2:31:44 distribution so multi region 99.999% SLA single digit multi second point

2:31:54 triads all in one SDK one string connection all in one.

2:32:04 Now we know that a traditional cache matches exact strings 18% hit rate.

2:32:15 A semantic cache matches meaning using the vector similarity 65% hit age.

2:32:25 The extra 47% is the money you are leaving with a table evidence the query

2:32:32 on the screen does it in one SQL call vector distance with a threshold of 0.80.

2:32:43 If the incoming question scores at or above that returns a cache answer.

2:32:48 No LLM.

2:32:53 Now for the embedding pipeline we have four components.

2:32:56 An embedding model which will be a text

2:33:00 embedding.3 small by the Azure OpenAI 1536 dimensions.

2:33:07 Then we have a Cosmos DB container with a distance in vector index DTL

2:33:14 item from 32nd to 7 days partition key which will fill the query hash.

2:33:20 And we have the vector index configuration which where we set the indexing

2:33:26 search list size between 10 and um.500 to tune the accuracy versus build speed.

2:33:33 And then the lookup query which will be the select top

2:33:37 one with the vector distance greater than or equal to 0.80.

2:33:42 Below that call the LLM And store the new entry.

2:33:50 So now let's have a look how this works.

2:33:54 Like.

2:33:59 So we have here, this is the entire cache lookup one SQL query.

2:34:13 We have the vector distance built into the Cosmos DB.

2:34:19 No extra library, no separate vector database.

2:34:24 Then I have some agents to try on.

2:34:28 This is the first one with three questions and then we have a next

2:34:32 run with another three question to see the cache hit and the cache.

2:34:39 So if I run the demo, we have here three pilots on the screen.

2:34:47 Semantic cache change feed, optimistic concurrency, all in one database.

2:34:52 The first agent as you can see here, like for example it have a support agent

2:35:00 with what is the refund policy for electronics?

2:35:03 It have a cache miss.

2:35:05 It doesn't find this before.

2:35:07 So now we have the answer.

2:35:09 The second agent, the return agent.

2:35:12 How can I return an electronic product?

2:35:14 The third agent called the product itself or anti on laptops.

2:35:20 So now we have missed right now what if I just Change the question again?

2:35:29 So I will just comment this uncomment the second

2:35:34 one or the second run it has similar questions,

2:35:38 similar meanings but not exact same writing.

2:35:43 So here for example what is the electronics return undefined policy

2:35:47 Then join item I put Then the warranty comes with that.

2:35:52 So let's run this one more time.

2:36:03 So as you can see we have here a cache

2:36:06 hit similar to 0.89 which is more than that 0.80.

2:36:13 We have got it in the slide.

2:36:16 The second one is also a cache hit similar to 0.91.

2:36:21 The third is also cache hit with 0.92 so

2:36:25 it finds all of this info on the database.

2:36:28 It doesn't need to call the LLM again and cost me money.

2:36:33 So how much it saved you?

2:36:35 It saved me.

2:36:36 Now just 327.

2:36:39 Okay, let's go back to our slides.

2:36:46 So what you just saw was three agents, three different questions, zero impulse,

2:36:52 all answers compact in about one second I

2:36:58 say what is electronics return and refund policy?

2:37:01 And I found it matches the refund policy for electronics.

2:37:04 For example at 0.89 similarity a traditional

2:37:08 cache would have made a semantic cache.

2:37:12 Catch it every time.

2:37:16 So the production results and this is the result

2:37:20 from a real production deployment which is constant with published benchmarks.

2:37:28 Before $47,000 a month,

2:37:32 803.50millisecond average latency with 10% cash hit rate.

2:37:39 After that I have $12,700 a month.

2:37:45 These 100millisecond average three months,

2:37:49 300millisecond average latency and 65% cache hit rate and 73% cost reduction,

2:38:01 65% latency improvements, one architecture.

2:38:06 Change it all.

2:38:10 So what about the disk?

2:38:11 And at scale, let's see.

2:38:14 So under 20 milliseconds at 10 million vector was greater than 90% recall.

2:38:24 We have under 100 millisecond at 1 billion vectors.

2:38:30 This is still greater than 90% recall.

2:38:34 70iu per query let down a complex SQL query and when

2:38:40 the index grow 100 times the latency increases less than 2 times.

2:38:50 And so the change feed we have two ways,

2:38:53 the traditional way and the agent memory separate way.

2:38:56 The traditional is where the agent A will

2:38:59 be sent some message queue and Q delivers To Agent B 200 to 500 millisecond end

2:39:08 times handoff plus the cost of the queue service.

2:39:12 With the agent memory fabric that we have

2:39:15 just had the agent A will write Cosmos DB

2:39:19 change feed triggers agent pay in real time

2:39:23 and no queue to deploy monitor or to pay for.

2:39:29 So the database is the message us Here.

2:39:37 The three patterns I use in production

2:39:40 are first of all the fan out coordination.

2:39:44 So one agent will write many react and partition by the task ID.

2:39:55 The button 2 is that state machine translation where the status

2:40:02 moves pending to planning to execution to validating to complete.

2:40:09 So we have all the status like in front of us.

2:40:14 Change feed fires on every transition.

2:40:18 You get a full audit tree for free.

2:40:21 Button three is competing the consumers.

2:40:25 Multiple agent instances should work from one partition the least procedure

2:40:31 will rebalances automatically when you scale from 1 to 100 instances.

2:40:42 The problem was that agent A and C post read the same decode,

2:40:48 post, write back, last write, win and the first write is silently lost.

2:40:55 That's why I have here the mistake concurrency so we can prevent the corruption.

2:41:02 And the solution is the ETag where every document have a version stamp.

2:41:08 If a document changes between you read and you write,

2:41:13 Cosmos DB will reject the write.

2:41:15 No distributed lock manager needed, just a streamline retry rule.

2:41:26 What about the conflict resolution?

2:41:28 So a 409 conflict is not an error, it's an information.

2:41:34 And it has three patterns.

2:41:37 First one, the last writer wins.

2:41:40 Discard your changes, reread, retry, use when refresh data.

2:41:46 Is here is always more correct.

2:41:50 Next is the merge and retice.

2:41:52 So fetch is a fresh state merge.

2:41:55 Both changes retry with the new ETag.

2:41:58 Use when the agents update different fields of the same documents.

2:42:03 The pattern three is about partition and avoid.

2:42:08 So design the partition keys so agents never write at the same document.

2:42:14 The best conflict resolution is no conflicts at all.

2:42:22 So here's the full architecture we have.

2:42:27 Before we had like four systems,

2:42:31 four pills, four filler points, zero coordination.

2:42:36 After that we have one database, one perl,

2:42:40 one SDK, one dashboard, sync capabilities and only one fell.

2:42:51 So the lessons that I made that I learned the hard way are listed here.

2:42:57 These are the four things I got wrong.

2:43:00 First one is a partition key.

2:43:02 The partition key design done last, not first.

2:43:05 It caused a painful live migration.

2:43:08 The right key was/, task ID,

2:43:13 even distributed plus the change feed locality design itself.

2:43:20 Number two is that I killed the vector index after loading the data.

2:43:25 That causes a 90 minute window of hundred percent cache messes.

2:43:31 Build the index before loading the data.

2:43:34 Number three is the compliance team asked for an audit trail change.

2:43:41 The feed was already captured every state transition for free.

2:43:45 It saves reproduction incidents through the replay.

2:43:49 And number four is that start with the partition and avoid,

2:43:55 not the merge and retry.

2:43:58 The best conflict resolution is no conflict.

2:44:06 So for you now what to call Monday morning.

2:44:09 Four steps.

2:44:11 Create a Cosmos DB serverless account five minutes.

2:44:16 Then you'll create a container with a disk

2:44:19 in and vector index on/embedding 2 minutes.

2:44:26 You need then to store the LLM Responses with embeddings.

2:44:31 That is my existing code with plus let's say 10 other lines.

2:44:38 And before every LLM call, run the vector distance query first.

2:44:46 You can have semantic caching in production by the end of the Monday.

2:44:55 But why this pattern lens is because we can see

2:44:59 that McKinsey Global Institute published a Report in November 2020.

2:45:04 Five says that AI powered agents

2:45:08 and Reports could generate $2.9 trillion in U.S.

2:45:13 economic value per year.

2:45:17 So the winners will not be the teams that build the most agents,

2:45:22 there will be the teams that run them efficiently.

2:45:26 This is what the agent memory fabric gives you.

2:45:35 Thank you all the code is open source.

2:45:38 You will find this at this repo.

2:45:40 I'll be around.

2:45:42 If you have any question please open an issue on the repo and let me know.

2:45:50 Thank you.

2:45:54 Your favorite apps light up your world.

2:45:57 Azure Cosmos DB powers the spark behind their shine

2:46:00 with massively scalable data and built in AI capabilities.

2:46:06 Intelligence on the move.

2:46:07 Audi uses Azure Cosmos DB to fuel

2:46:10 AI assistants that deliver answers at highway speed.

2:46:15 Creativity that connects with Azure Cosmos DB.

2:46:19 Adobe helps brands craft personalized experiences with real time relevance.

2:46:25 Color reimagined built on Azure Cosmos DB.

2:46:29 Pantone's palette generator turns inspiration into expression.

2:46:34 Goal.

2:46:36 Premier League uses Azure Cosmos DB to keep fans close

2:46:39 to the pitch with insights drawn from match day model moments, talking cars.

2:46:45 The future of in car infotainment starts

2:46:48 with TomTom and Azure Cosmos DB answers uninterrupted.

2:46:54 Together.

2:46:54 Azure Cosmos DB and ChatGPT, the world's number one AI app help 800

2:46:59 million users keep work moving with near zero downtime.

2:47:03 From connected transportation to creative exploration and so much more.

2:47:08 Azure Cosmos DB powers the AI apps that move the world forward.

2:47:18 Hey folks, welcome to a session about Azure DocumentDB.

2:47:22 We have a really exciting session for you where I'm going to show

2:47:24 you first what and how things work under the hood for Azure DocumentDB.

2:47:29 What are the key functionalities and how you can modernize

2:47:33 your MongoDB compatible workloads that are on prem or on the cloud,

2:47:37 with just a few clicks.

2:47:39 So before we get into the slides and the content,

2:47:42 let's go over a brief demo and I'm going

2:47:44 to show you how everything works behind the scenes.

2:47:46 So over here I have a website where you place orders.

2:47:52 And our website is called Contoso Retail.

2:47:55 Think about it like a website like Walmart

2:47:56 or seven eleven where they have an online database,

2:48:00 where they sync up everything with their local instance and every

2:48:04 local store would have their own on prem DocumentDB server.

2:48:09 High level, to brief it up we have

2:48:12 a centralized database which connects all the local

2:48:16 stores that are in person with all

2:48:19 the local database for DocumentDB for these stores,

2:48:22 that are available out there today.

2:48:26 Over here we have the global database.

2:48:28 Let's also look at other databases that we have,

2:48:31 we have one store in Seattle and then we have a second store

2:48:34 in Chicago and it's all connected to our global database as, I mentioned before.

2:48:40 Now let's go through our Seattle store

2:48:42 and see what orders have been placed so far.

2:48:46 The last order that we had was a customer demo for like $249.

2:48:50 And the status still is still pending.

2:48:53 You can even add your own orders right now.

2:48:55 Let me briefly see how it works.

2:48:58 We place an order and the last order was just populated on the screen.

2:49:03 Everything is going to be synced up to a global database,

2:49:06 which is the Azure DocumentDB database.

2:49:10 If I hit refresh, the same order has

2:49:12 been popped up over here with the same numbers, and the same amount.

2:49:17 But let's say for some reason your local databases are not working today,

2:49:22 your WI fi, is off, or there's some kind of disaster that's happening.

2:49:28 We're going to try to replicate that in this demo as well.

2:49:31 I'm going to click on simulate Outage.

2:49:34 Everything that's connected to your local database

2:49:36 to your online database has been cut off.

2:49:39 There is no sync between the two.

2:49:42 If I go back to my Seattle database and add another order.

2:49:46 Let's see, let's add four items to the list and, place an order.

2:49:52 As you can see, we have an order Created

2:49:55 with order ID as last four digits as 5471.

2:50:01 If I hit refresh, it's right here, 5471.

2:50:04 But let me go back to my global database and see it has been synced up or not.

2:50:13 Let's give it a few more seconds.

2:50:16 Now we are connected to a global database.

2:50:18 Let me hit refresh.

2:50:21 As you can see, the last order ID is 5470.

2:50:26 Now let's take a look at what goes on behind

2:50:29 our database to make sure that we actually track,

2:50:34 this order in our local database.

2:50:37 I have, three sources over here.

2:50:39 First is my Azure DocumentDB instance, which lives on Azure.

2:50:43 Then I have a local DocumentDB instance which is, I've named that Chicago.

2:50:47 And then we have Seattle.

2:50:48 If I go to Orders and Documents and search for my last order,

2:50:55 which was 5471, you can see it has popped up.

2:50:59 It has all the items that I just placed an order for.

2:51:03 They are titled by SKU names.

2:51:05 It's still here.

2:51:06 But if I do the same thing, with my Azure database collection and if I go

2:51:12 to JSON view and search for 5471, nothing shows up.

2:51:18 We don't have anything in our global database just yet.

2:51:23 I'm going to cancel the outage and click to restore.

2:51:28 The disaster mode is off and the connection has been restored.

2:51:32 It knows all the changes that are in the pipeline

2:51:35 are going to be flush to our global database.

2:51:38 As you see it just mentioned,

2:51:39 auto sync 1 new order from Seattle store to global.

2:51:44 It's going to take a few minutes for it to pop

2:51:46 up on the recent orders and we can take a look.

2:51:50 The reason this is really crucial for you is because Walmart as a store,

2:51:53 at any store that are out there you have

2:51:56 local stores that are running on a 247 basis.

2:52:00 You cannot just rely on the cloud at every second.

2:52:03 Everything like we know some things could go wrong.

2:52:06 Customers if they want to buy something, we cannot be like hey,

2:52:09 the Internet is down, we cannot process a transaction.

2:52:13 So that's why having a local database is really important for these stores.

2:52:17 So as you can see it just flushed one new change to the global database.

2:52:21 Let me hit refresh again and you can see the order

2:52:25 id 5471 has been popped up on our global database.

2:52:29 And the amount everything is like searched and filled up over here.

2:52:34 So this is really crucial.

2:52:37 Since we are in truly open source database,

2:52:39 you can run DocumentDB on your local laptop computer wherever or even.

2:52:44 You can run it on aws, GCP or Azure.

2:52:48 It makes it so flexible for you to just

2:52:50 manage your workload based on what your needs are.

2:52:53 You're not tied into one vendor or one service at a time.

2:52:58 Also I want to show you that there is a bidirectional sync

2:53:01 between global and from local to global as well as global to local.

2:53:06 Let's say in this scenario the Contoso retail store,

2:53:09 they want to sync up everything once a day,

2:53:12 everything that's happened in 24 hours

2:53:14 of a local store to their headquarter database.

2:53:17 So they can do analytics,

2:53:19 do reviews and create dashboards on how many orders were placed.

2:53:23 We allow that as well.

2:53:24 Let's say like global database makes an update to the product sku,

2:53:30 or a product pricing has been changed.

2:53:33 Let's say the wireless headphone charge price goes from 249 to like 200.

2:53:38 They're running a sale for Thanksgiving or something like that.

2:53:41 So we can pull our data from a global database and it's going

2:53:44 to sync everything up with the local database that we have out there.

2:53:50 Right now we only have three databases stored.

2:53:52 We have one Chicago location, Seattle location and a global database.

2:53:56 But you can have multiple stores and have

2:53:58 a bidirectional sync between all these stores.

2:54:01 If something goes wrong in one of the local stores,

2:54:04 don't worry, your changes are not going to be gone forever.

2:54:07 They're just going to be queued up and they're going to be stored using change

2:54:11 streams later once the connection has been

2:54:13 restored and the disaster recovery has been fully.

2:54:17 You can also track what's going on for all these stores in our Azure service,

2:54:20 at the same time as well.

2:54:23 Now let's take a look at how our data is being

2:54:26 laid out behind the scenes using our own VS code extension.

2:54:29 DocumentDB for VS code extension let's go and take a look at how

2:54:33 our data is laid out for our online

2:54:36 database which is the Azure DocumentDB database.

2:54:39 Our local instances of DocumentDB which is

2:54:41 the Chicago store and the Seattle store.

2:54:45 Let's go through our orders and you can see we

2:54:48 have all the information necessary needed for our online headquarter database.

2:54:55 So if I click on the view for a specific document or a specific data point,

2:55:00 you can see it has all the information necessary

2:55:02 that an audience that a headquarter person would need.

2:55:06 Like where was the data synced from?

2:55:08 It said it was synced from the Seattle store at time April 16 at 7

2:55:14 30pm it also has all the information of what the status of that information is,

2:55:18 what was the total amount that was collected and what

2:55:21 was the items that were added to their sku.

2:55:25 Now let's go and take a look at the performance for our online database

2:55:30 and we want to make sure that the performance is optimal because let's say we

2:55:35 have an online store like Walmart where people want to search up what products

2:55:39 they are looking for and if it takes forever for them to load the website,

2:55:43 load the product they're looking for, it's

2:55:45 just going to be a bad user experience.

2:55:47 So now let's take a look at the orders and see how

2:55:51 long it takes to find a certain item that you're looking for.

2:55:55 So you can see we have multiple fields over here.

2:55:57 One of the fields is status.

2:55:59 So let's try to filter out based on the status of an order.

2:56:06 So if I just do status and shift and hit

2:56:11 run it's going to give me all the data points.

2:56:14 All the documents that have status equals shipped.

2:56:17 Now I want to take a look at the query performance for this exact same query.

2:56:22 As you can see it took like 3.5 milliseconds to examine over 15,000

2:56:28 documents to return only 2866 document the ratio for examine to return ratio

2:56:34 is 5 to 1 which is not good because you don't want your database

2:56:38 to do a full collection when

2:56:39 retrieving information for what they're looking for.

2:56:44 In this case it's only 15,000 documents but a store because Walmart there are

2:56:48 going to be millions and millions of millions

2:56:49 of documents or orders in the database.

2:56:53 If we scale it up to those levels it's just going

2:56:55 to be forever for them to get the information they are looking for.

2:56:58 So that's why we built something which is very useful using AI,

2:57:02 which is an AI powered performance insights.

2:57:06 Well how it works is it already has

2:57:08 all the information about your cluster, your database,

2:57:10 how your indexes are stored,

2:57:12 how your data is laid out and the performance of your database.

2:57:16 So it takes all that information and we

2:57:18 have a bunch of skills that MD built right

2:57:20 into it which gives you the recommendation of what

2:57:23 we think is going to be optimal for you.

2:57:26 Before we even give you the recommendation it

2:57:27 tells you what exactly is wrong with your query.

2:57:30 Over here you can see the performance summary said

2:57:32 the query performance is poor because you basically did

2:57:35 a full collection scan and which is inefficient for filtering

2:57:40 on the status field that we are looking for.

2:57:43 You don't want to do a full

2:57:44 collection scan when you're working with your database.

2:57:47 So and then it's going to give you a one

2:57:50 click recommendation solution where all it says is hey,

2:57:53 just create an index on the status field.

2:57:56 And it gives us, we have the button for it.

2:57:58 Just click on Create index and it's just going to create the index for you.

2:58:02 So as you can see it took like two

2:58:04 seconds but the index has been created on status field.

2:58:09 We are looking for.

2:58:10 So let me click over here, go to indexes and just hit refresh.

2:58:15 We can see our status index has been created.

2:58:17 Now let's run the exact same query to see

2:58:20 the performance difference with and without an index.

2:58:26 It's 1.74 milliseconds so we actually reduced it by exactly half that time.

2:58:33 We only examined 2,866 documents.

2:58:36 Return 2860 is the document that we are looking for.

2:58:41 This is how an index was actually optimized your performance.

2:58:47 You don't have to read any documentation or find the best indexing policies.

2:58:51 That's why we brought everything together using AI to make sure

2:58:54 your database performance is the most optimal you are looking for.

2:58:59 Not just this only works on an Azure DocumentDB service.

2:59:02 It works on any local DocumentDB service.

2:59:05 Or any Mongo compatible database that's running out there.

2:59:10 So let's do the same thing over here and let's do,

2:59:12 I'm going to open up a Seattle store and filter out by status

2:59:16 equals shift and run a fine query and it's going to give us,

2:59:24 since we don't have an index on, since

2:59:28 we don't have an index on the status field,

2:59:30 it's just going to examine all the documents that we

2:59:34 have to give us the response that you're looking for.

2:59:39 So AI performance also works on your local

2:59:42 instance as well as an Azure DocumentDB instance.

2:59:45 Now since we look at the demo, everything worked well.

2:59:48 Now let me give you a brief overview of what Azure DocumentDB is

2:59:51 or what is DocumentDB and why we even built it in the first place.

2:59:55 So you must be already familiar with MongoDB databases.

3:00:00 It's the most popular non relational database that is out there.

3:00:04 People allow the specifically people it allows to specifically

3:00:09 doc to specify documents in a JSON format.

3:00:12 And JSON fits naturally into many web developers code bases with the popularity

3:00:17 of JavaScript and it is very popular

3:00:20 to build solutions like web and mobile apps,

3:00:23 AI and rag applications and product catalog and personalization.

3:00:27 Personalization.

3:00:28 So what is DocumentDB?

3:00:30 It is an open source MongoDB compatible

3:00:33 database built for flexibility, scale and AI.

3:00:36 It is now part of the Linux foundation

3:00:38 and it has the industry support from Azure,

3:00:41 aws, gcp, Snowflake Cockroach Labs and so many more.

3:00:47 It also gives us ability to run your DocumentDB

3:00:50 instance on AWS GCP on Azure the same

3:00:53 time where you can replicate information across all

3:00:56 the clouds so you're not vendor lock in.

3:00:59 And also you can run it on prem.

3:01:02 As I showed in the demo few seconds ago,

3:01:05 the mission of DocumentDB is it's built on principles of transparency,

3:01:10 freedom and standardization visibility.

3:01:13 We want to ensure developers have

3:01:15 full visibility in the underlying architecture.

3:01:18 That's why with the MIT license users have complete freedom

3:01:22 to use the project as they please with no restrictions.

3:01:25 It is an open standard.

3:01:26 The goal is to create an Open Standard

3:01:28 for DocumentDB databases for a universally accepted implementation standard.

3:01:33 All right, so based on DocumentDB's popularity we thought

3:01:37 about let's also build an Azure first service built

3:01:40 on those open source DocumentDB and we call it

3:01:44 Azure DocumentDB With mongodb compatibility it is built innovative.

3:01:48 You can build innovative apps with truly open

3:01:51 source 99.03 MongoDB compatible document database in any environment.

3:01:55 You can also deploy in Azure for the best

3:01:57 Enterprise experience with hybrid and multi cloud capabilities.

3:02:02 It has AI driven AI built right into it.

3:02:05 It is a built in no cost vector search powered by DiskANN.

3:02:10 So you can combine your vector search queries with and your full

3:02:13 text queries with hybrid search right into the database.

3:02:17 Think about one database for all kinds of your search workloads.

3:02:20 You don't need to ETL your database out

3:02:23 of operational database to vector database like Pinecone or Vivid,

3:02:28 something like that.

3:02:29 Everything lives under one database so you can run all your queries together.

3:02:33 It is enterprise ready.

3:02:35 We have authentication with Entra ID.

3:02:37 You can get.

3:02:37 You get all the perks of being an Azure service.

3:02:40 We also give you up to 99.99995 availability SLA across the full service stack.

3:02:47 It also is cost effective.

3:02:48 It reduces the total cost of ownership

3:02:51 by scaling compute and storage independently.

3:02:54 And also 35 days of free backup and 24,7 support included in the price.

3:02:59 So there is no extra cost for licensing or support or even backup.

3:03:04 The price you see on the website is the price you pay every month.

3:03:07 It is also only based on your compute

3:03:09 and the storage that is required for your workload.

3:03:13 There are a lot of customers that are actually using Azure DocumentDB today.

3:03:17 We have AB and UBS, KPMG, UnitedHealthcare,

3:03:21 a lot of customers for different workloads.

3:03:25 Some of them like EY and KPMG are using DocumentDB

3:03:28 for their vector search workloads for the AI native applications.

3:03:33 One of the key differentiator for Azure DocumentDB is its

3:03:36 hybrid and multi cloud freedom like the demo I showed before.

3:03:40 You can run DocumentDB on any service that is out there,

3:03:44 local, aws, GCP and Azure at the same time.

3:03:47 And there's about unidirectional replication between all these services because

3:03:51 it runs on the same engine that powers them all.

3:03:54 And Also with Azure DocumentDB you can lower your monthly cost by up to 45%.

3:04:01 Also we give reserve instances.

3:04:03 If you commit for a year we give you up to 40% discount and if you

3:04:06 come in for three years we give up

3:04:08 to 60% discount on your MongoDB compatible database.

3:04:13 You can save up to 50% by switching to Azure DocumentDB

3:04:16 and it gives you the best perks that are out there.

3:04:19 We Give up to 32 terabytes per shard free point

3:04:23 in time backup and restore and no additional support contracts or licensing.

3:04:27 I'm going to emphasize this again.

3:04:29 The price you see on the website is the price you pay at the end.

3:04:33 There Are no surprises, no additional contracts,

3:04:36 licensing or any other fees that are associated with Azure DocumentDB.

3:04:39 But wait, there's more, there's more discounts.

3:04:43 We also give up to additional 20% customer incentive,

3:04:47 discounts on top of 1% and 3% RI.

3:04:51 If you're interested, feel free to reach out to us@DocumentDB@microsoft.com

3:04:54 and today we are introducing Premium SSD V2.

3:05:00 With Premium SSD V2 it gives you the maximum iops

3:05:04 and throughput for every disk selected at no additional fees.

3:05:07 So you can do 80,000 IOPS and around 1200 milliseconds per disk selected so

3:05:14 you can upgrade your speed but this the same price that you paid today.

3:05:19 We also have the LLM based Index Advisor

3:05:21 for DocumentDB which I demoed a few seconds ago.

3:05:24 So make sure you get the best optimum performance from your database,

3:05:28 wherever you're running it from.

3:05:30 And say goodbye to manual troubleshooting with just one click.

3:05:34 AI Optimization Then we also bringing advanced

3:05:36 full text search capabilities in Azure DocumentDB.

3:05:39 Think about Elasticsearch workloads that people usually run with the database.

3:05:44 We are bringing all those functionalities

3:05:45 built right in directly into Azure DocumentDB.

3:05:48 So no more paying for third party services for your search workloads.

3:05:52 Everything is built right into your database.

3:05:54 Some of the functionalities we're bringing in the next

3:05:57 month or so is going to be fuzzy.

3:05:59 Search proximity search,

3:06:01 multi language support like Chinese and Thai BM25 Ranking and rank Fusion.

3:06:05 So with Rank Fusion you can run

3:06:07 your hybrid search workloads with just one query.

3:06:10 So you combine your semantic search

3:06:12 and your full text search and use Rank Fusion

3:06:15 to find the best optimal solution best optimal

3:06:18 documents in the database for your exact query.

3:06:23 And in the future we're going to bring on multi

3:06:25 field indexing analyzers and tokenizers and tokenizers filters and so on.

3:06:30 And did I mention we have the best in class vector search.

3:06:33 We have performed significantly better than

3:06:36 any other MongoDB compatible database out there.

3:06:38 There are so many benchmarks that have

3:06:40 been reported about the performance with latency, lower latency and higher rps.

3:06:45 We also have Microsoft AI built in vector indexing which is the disk scan

3:06:51 and technology which supports richer vector search

3:06:54 capabilities so you get high accuracy and speed.

3:06:56 You can scale up to millions of vectors

3:06:58 embedding with high accuracy and fast speed.

3:07:01 We support up to 16,000 dimensions for context.

3:07:05 The largest embedding model that is out there

3:07:09 is 4096 dimensions but we support up to 16,000.

3:07:12 So let's say in the future, OpenAI, Anthropic or Gemini comes up with their own

3:07:17 model which has higher embedding dimensions.

3:07:19 We still support that in Azure DocumentDB.

3:07:21 No more thinking about what would happen.

3:07:24 So we are compatible with all the models

3:07:26 that are being released in the world today.

3:07:31 One other thing about Azure DocumentDB is you can do vertical

3:07:33 scaling as well as horizontal scaling along with the storage scaling.

3:07:38 So vertical scaling involves increasing the resources like vcores,

3:07:41 virtual cores and RAM of your existing nodes or shards.

3:07:44 Easily horizontal scaling means you are

3:07:47 distributing your data across multiple nodes,

3:07:50 enabling support for larger datasets and higher throughput.

3:07:53 You can also instantly independently increase

3:07:56 your disk size which without altering the compute

3:07:59 tier which gives you providing more flexibility

3:08:04 if your data grows in the future.

3:08:06 Let's go into the scale up process in much more detail.

3:08:10 Since Azure DocumentDB is built on scaling up process,

3:08:13 we support scaling compute which is vcores and RAM and storage iops this size

3:08:18 independently you can increase or decrease your computer

3:08:21 storage without service downtime or application changes.

3:08:25 We have larger disk up to 32 terabytes and 64 terabytes

3:08:29 per shard and we allow up to 10 shards at a time.

3:08:33 The architecture supports smaller compute

3:08:35 larger disks depending on your workload.

3:08:38 We also have horizontal scale out

3:08:40 when scaling vertically is no longer sufficient.

3:08:42 You want to scale horizontally by increasing the fuzzy physical shard count.

3:08:48 Azure DocumentDB manages logical physical shard mapping and the logical

3:08:52 shards are mapped into physical shards and requests are routed automatically.

3:08:56 Rebalances is handled behind the scenes.

3:08:59 We do everything for you and we go up to 10 shards at a time.

3:09:03 We are the only MongoDB compatible database with instant autoscale.

3:09:08 How it works is ensuring there's a peak demand.

3:09:10 We are always sure that your database is running at the same time.

3:09:14 Let's say there's a spike,

3:09:16 during a certain time of the month or a certain time of the year

3:09:19 where you see a huge spike from a user from your user using your website.

3:09:24 So instead of manually scaling up like the month

3:09:27 before to make sure that you meet the user demands,

3:09:30 we instantly auto scale for you.

3:09:32 And if we see the demand has slowed down,

3:09:35 we can manually lower the computer based on your workloads so your application

3:09:41 is still running without that unexpected need

3:09:44 that is going to crash the website.

3:09:46 We ensure that all our users have

3:09:48 instant autoscale so customers will choose this right

3:09:52 here for their workloads and over under

3:09:54 provisioning and reduce the escalations and budgeting.

3:09:58 Since we are a first party Azure service

3:10:00 you get all the authentication and access management,

3:10:02 data protection and network security built right in for no additional cost.

3:10:08 We have database security in Azure DocumentDB with in transit security,

3:10:12 Data Encryption and transit and at rest security.

3:10:15 So with a first party Azure service you

3:10:18 get all the Azure requirements that you're looking for.

3:10:20 And we also made sure migration is

3:10:23 seamless for whatever workflow you're looking at.

3:10:25 We offer three types of migration.

3:10:26 First we have the offline migration.

3:10:29 So think about it like a snapshot bulk copy from the source to the target.

3:10:35 If new data has been edited after we start an offline migration,

3:10:40 it's not going to be copied over.

3:10:41 But for that exact reason we have online migration.

3:10:44 But apart from the bulk start copy activities done in offline migration,

3:10:48 a change stream monitors all the addition,

3:10:51 updates and deletion that has happened.

3:10:53 Once the bulk copy start has been worked on, it remembers all that information.

3:10:58 Once that information has been migrated to Azure Document db,

3:11:01 it updates and deletes whatever the updates

3:11:03 have been done when the migration has started.

3:11:05 You can also cut over at any time when performing

3:11:08 an online migration to ensure it remains active until manually finalized.

3:11:14 Using cutover to prevent data loss and only proceed with cutovers

3:11:18 when the replication gap has been eliminated across all your collection.

3:11:24 Choose the best way you want to migrate your data

3:11:27 from your local or any other MongoDB compatible database to Azure DocumentDB.

3:11:32 We have an excellent tool for this where we have a VS

3:11:35 code extension where you just have to provide the connection string

3:11:39 from your source to the destination which is going to be

3:11:42 an Azure DocumentDB and you could

3:11:43 just migrate your data without any interruption.

3:11:47 Since we are MongoDB compatible database up to 99.03%

3:11:51 we will tell you if certain changes are required,

3:11:53 which I don't think is going to happen, but we give you an assessment,

3:11:57 a pre migration assessment which tells you what are at risk,

3:12:01 what needs to be changed before you even start your migration.

3:12:04 So you can have all the information you need before migrating

3:12:08 your huge workloads from on PREM or any other provider to Azure DocumentDB.

3:12:15 So since you have, let's say you are running through migration and you

3:12:18 have an issue migrating your data to the cloud to Azure DocumentDB.

3:12:22 We provide free migration with Cloud Accelerate Factory.

3:12:25 You can also reach out to the product team directly by just emailing

3:12:29 us@cbd migrationsupport@microsoft.com and one of the people

3:12:33 working on Azure DocumentDB will definitely reach

3:12:35 out and make sure your migration is seamless and there are no other issues

3:12:39 or data loss while you're migrating your data from On Prem to the cloud.

3:12:43 So we give you all the tools that are necessary to migrate

3:12:45 from On Prem art from any other Mongo compatible database to Azure DocumentDB.

3:12:51 If you have any questions, feel free to reach out.

3:12:53 And I think you should get started with Azure DocumentDB with a free

3:12:56 tier where we give up to 64 gigabytes of storage for completely free.

3:13:02 No credit card required to sign up.

3:13:04 Just try out and see how you like it.

3:13:06 Or you can even try our local instance, run a local DocumentDB server,

3:13:10 using a Docker container or even checking out all the samples that we

3:13:14 have built for Azure DocumentDB or DocumentDB

3:13:16 by checking out the website DocumentDB.io/samples.

3:13:20 The demo that I showed over here is also going to be available over there.

3:13:24 So if you want to take a look at the code,

3:13:25 how it worked, all the information is provided over there.

3:13:28 But feel free to reach out to us if you have any questions or drop a question

3:13:32 in the chat on the YouTube comments and we

3:13:34 will make sure we're able to answer it.

3:13:36 Thank you.

3:13:40 My name is Johnny Jalife and I'm CTO with Tao Works.

3:13:48 We typically use Cosmos DB in the solutions we build for our customers.

3:13:52 As the operational backbone for cloud native applications.

3:13:55 It works really well for customer facing platforms

3:13:57 that need to handle real time data like user context,

3:14:01 configuration and workflow state.

3:14:03 We also use it a lot in distributed and event driven

3:14:06 systems where data is coming in from fast and changing constantly.

3:14:10 And that's where Cosmos DB really stands out.

3:14:18 The main benefits we've seen are speed,

3:14:20 flexibility and less operational overhead.

3:14:23 Cosmos gives you low latency performance,

3:14:25 a data model that adapts easily as requirements evolve

3:14:28 and the ability to scale without having to redesign the system.

3:14:31 And it's fully managed, teams can stay focused on billing and shipping

3:14:35 instead of spending time running the database.

3:14:43 The main outcomes have been faster delivery,

3:14:45 consistent performance and the ability to keep expanding without friction.

3:14:49 We are able to ship features quicker,

3:14:51 keep performance strong as usage grows and add

3:14:53 new use cases without having to rework the data.

3:14:57 At the end of the day it helps

3:14:58 our customers customers move faster and scale with confidence.

3:15:05 Welcome back and welcome back.

3:15:08 And all I can say right now about what we've seen so far.

3:15:11 Amaze, amaze, amaze.

3:15:13 That's right Jay.

3:15:14 We've seen how developers are building

3:15:16 across clouds and designing flexible systems.

3:15:19 Yeah.

3:15:20 And you know what?

3:15:20 We've heard back from some of our great, great members of our community,

3:15:25 who have left some things in the chat, Chenola said,

3:15:29 looks like there is a great future for DocumentDB.

3:15:32 You know there is Patty.

3:15:34 I know that there is.

3:15:35 We're doing a lot of great things

3:15:38 with DocumentDB and especially with open source DocumentDB.

3:15:42 So we're always available and if you want to learn more about it,

3:15:46 just visit DocumentDB.io.

3:15:48 Also we have a great comment from behind.

3:15:52 Building real time dashboards is a breeze

3:15:55 using change feeds and it really, really is.

3:15:58 Absolutely.

3:15:59 And we heard so much great stuff from Justine today about

3:16:02 the change feed and how it's really helping people with real time applications.

3:16:06 But now it's time to go into some real,

3:16:09 real deep content things that we know are so important to having

3:16:14 a well created efficient and cost

3:16:19 especially considered when we're talking about indexing.

3:16:21 Yeah, I couldn't have said it better Jay.

3:16:24 This is really where performance comes together in just an application.

3:16:29 Yeah.

3:16:30 And if you want to revisit these sessions later,

3:16:32 you'll find everything in the full playlist.

3:16:35 It's going to be@ah, aka aka.ms/cosmosconf26playlist.

3:16:39 It's in the chat right now.

3:16:41 Go ahead and keep it bookmarked for later.

3:16:45 Yeah, I know for me personally we're so busy hosting,

3:16:49 that's definitely something that I will check out afterwards.

3:16:51 Absolutely.

3:16:52 And if you have a minute please, we'd appreciate your feedback.

3:16:56 So if you want to share your thoughts, go to aka Ms.

3:17:02 aka.ms/cosmosconf2026survey.

3:17:04 Yeah, we want to know so we can keep

3:17:06 doing this show and making it great for you.

3:17:09 So up next we've got a really great one.

3:17:11 It's called Querying and Indexing in Azure Cosmos DB.

3:17:14 The Complete Guide.

3:17:16 Yeah, here's Dr.

3:17:17 James Codella.

3:17:20 Hey everyone.

3:17:21 Welcome to Querying and Indexing in Azure Cosmos DB.

3:17:24 I'm James Codella, Product manager for Querying AI in Cosmos DB.

3:17:28 Let's get started.

3:17:30 So in this session we're going to imagine a multi tenant eretail platform where

3:17:34 we have customers that can shop across

3:17:36 a portfolio of different products and outdoor brands.

3:17:40 We're going to take a look at some common query workflows that such

3:17:44 a platform might use in order

3:17:45 to help their customers search for different products.

3:17:49 Now we're going to run the same queries against the same set of data,

3:17:52 but we're going to compare it in two different ways.

3:17:55 First is we have a container called

3:17:57 Products which has no indexing policy whatsoever.

3:18:01 And we're going to use some unoptimized queries to run These workflows.

3:18:05 Next we're going to run it the same

3:18:08 queries against a optimized collection that has

3:18:11 the best indexing policies and using the best

3:18:14 practices for creating queries in Cosmos DB.

3:18:17 And we're going to compare the performance and cost differences.

3:18:23 So what are we going to cover in this session?

3:18:26 We'll start briefly on partition keys, what they are and how to use them.

3:18:30 Then we'll dive into basic queries, queries that use a, composite index or more

3:18:35 complex patterns handling array data in queries,

3:18:39 geospatial data, computed properties,

3:18:43 and finally end with our advanced search capabilities including vector,

3:18:47 full text and hybrid search.

3:18:50 Plus we'll take a look at indexing metrics and query execution metrics

3:18:54 to help us make sure that we're preparing our queries are performing optimally.

3:19:00 So in all of these different scenarios we're going to see a few common patterns.

3:19:04 First is we'll see an example of a query

3:19:06 or two that is going to power the scenario.

3:19:09 Next we'll see the RU charge.

3:19:11 So how many request units that query execution cost.

3:19:14 We'll see the latency or the time it took

3:19:17 for the Cosmos DB container to execute that query.

3:19:20 And then we'll see the indexing metrics.

3:19:22 So this will show us which indexes were utilized by the query

3:19:26 and what potential indexes could have added benefit to the query if they were,

3:19:30 if they existed in the indexing policy.

3:19:33 So maybe better performance or better cost or both.

3:19:36 We'll also see query execution metrics which

3:19:38 show us exactly component wise where the query

3:19:41 took the most time and spent the most amount of work in the entire process.

3:19:47 Now, what does our data look like?

3:19:49 Well, imagine that we have a catalog of different products

3:19:52 for our E retail shop and we have different products as data items.

3:19:56 So you may have an id, a tenant or a store id.

3:19:59 So this is the Ridge Supply Store, a sku, a name for the product.

3:20:04 So in this case it's a camp stove and a bunch of other

3:20:07 metadata that you'd expect to see in a product catalog along with a description,

3:20:10 tags, some date timestamps as well as GeoJSON data that tells

3:20:16 us exactly where this product is located at that specific store.

3:20:20 As well as a embedding field that contains a vector

3:20:23 embedding that describes the description and name of the product.

3:20:26 And we'll use this later on for a vector similarity search.

3:20:33 So let's start with partitioning.

3:20:35 Partitioning Azure Cosmos DB determines how data is

3:20:39 distributed in routed Partition key provides a hint

3:20:42 to Cosmos DB to which logical partition

3:20:45 the data should be written to or read from.

3:20:49 So in our scenario we imagine that we have three

3:20:51 retailers or three shops and we're going to partition on that.

3:20:54 So most of our eretail queries can be scoped to a single

3:20:58 retailer unless the user wants to search across different retailers.

3:21:02 Now in general a good pattern to use is

3:21:04 queries should really include the partition keys whenever possible.

3:21:08 This allows your queries to be routed to the exact logical

3:21:11 partition or small set of logical partitions where your data lives.

3:21:15 This helps you keep RU costs and latencies

3:21:17 low and avoids potentially costly cross partition queries.

3:21:23 Okay, and we'll see an example of that in our first couple of queries.

3:21:26 So let's start with the basic query operations.

3:21:28 Right, so when I say basic queries we're going to use equality filters.

3:21:32 So quality filters, range predicates, single property sorts or order buys,

3:21:37 all these in combination to create a very basic common query pattern.

3:21:42 So for an example in our E retail shop maybe we would need

3:21:45 to filter on a category or do some simple sorts like sorting products by price.

3:21:51 That's useful really every single sort of product grid

3:21:54 or carousel that you'd see in an E retail shop.

3:21:57 And it's important to note as we'll see a query that returns

3:22:00 the correct results is not the same as a query that scales.

3:22:03 So we'll want to make sure that we're using best practices

3:22:05 and the right indexing policy to make

3:22:07 sure that our query can execute performantly.

3:22:12 Okay, so I'm going to switch over here to my terminal

3:22:17 and the first thing I'm going to do is I

3:22:21 have some sample runner here where I'm going to run some

3:22:24 sample queries against my data set of about 10,000 different products.

3:22:28 And so I'll go ahead and run this project and we'll get started.

3:22:33 So we're going to see two scenarios.

3:22:34 We're going to see a basic filter query and a basic filter plus order by query.

3:22:40 So here we can see the example of the basic filter.

3:22:44 So we're selecting a few of the different property paths

3:22:47 of our data items and we're adding where clause and two conditions.

3:22:51 Right?

3:22:52 So we're going to scope it to a particular

3:22:53 tenant or shop and then we're going to make

3:22:55 sure that the category name is targeted

3:22:58 to the category that we're really interested in looking for.

3:23:01 So maybe this user really wants to look at camping tents and backpacking tents.

3:23:06 So we're going to press enter and run the first query.

3:23:09 And so now we're going to spin up our job

3:23:10 and we're going to go execute a query from my laptop,

3:23:13 hitting the Cosmos DB backend from my Cosmos

3:23:16 DB resource deployed on the west coast

3:23:17 and come all the way back here to New York and show me the results.

3:23:21 And so here I can see the results for my query and I can see,

3:23:23 you know, a couple different products that came up.

3:23:25 And as I scroll up here, I'm going to see a few things also.

3:23:29 Right, so here are the index metrics that we had talked about.

3:23:32 So this first query is running on our container

3:23:35 products container that has no indexing policy whatsoever.

3:23:38 So the index metrics says that it wasn't able to utilize any index,

3:23:43 which makes sense because we don't have an indexing policy.

3:23:45 But it was looking for potential indexes.

3:23:47 So it was looking for a category name and it was looking for a tenant id.

3:23:51 So these are both things that we filtered on in the query itself.

3:23:54 And the impact score is high.

3:23:56 So if we add these to our indexing policy like we're really

3:23:59 going to get much better performance and better RU cost characteristics as well.

3:24:04 Scrolling up we can see the query metrics.

3:24:06 So this tells me how many documents were retrieved,

3:24:09 what the size was of all those documents and other

3:24:12 metadata that helps me understand the total query expense,

3:24:14 execution time or where the query spent a lot of time in processing.

3:24:19 And then I also print out some summary metrics

3:24:21 like how many rows were returned in this query.

3:24:23 So 136 match the query filters and then the RU charge.

3:24:27 So 343, almost 344 RU.

3:24:30 So this is not a cheap query.

3:24:33 Okay, so let's go and run this on the optimized container now.

3:24:36 So we'll now run this on the container that has an indexing policy,

3:24:39 that will make this query more performant.

3:24:42 And we can see the same results come back.

3:24:45 And this time in the index metrics we see that it

3:24:47 was able to utilize the indexes that are now in the policy.

3:24:51 And you know there are no indexes that were

3:24:53 not in the policy that it was looking for.

3:24:54 So that's great.

3:24:55 So we're effectively using all the indexes that we

3:24:57 can and we can see that you know, RU charge is significantly less.

3:25:01 Actually let's do a side by side comparison.

3:25:03 So we can see that without the index it was about 240

3:25:06 milliseconds of execution time for this query

3:25:09 with the index about 38 milliseconds.

3:25:12 So more than a 6x improvement on latency and then more than

3:25:17 a 7x increase or sorry 7x improvements on RU charges for this query.

3:25:22 So really demonstrating the power of having proper index even just with a simple

3:25:27 filter with just a couple of query where clauses in the query.

3:25:33 Okay so now we're going to run a basic filter and an order by clause.

3:25:37 So we're going to go ahead and run this again using our.

3:25:40 NET SDK running sort of a default query.

3:25:42 And we get back like this error actually

3:25:45 wasn't able to operate because it was looking for an index to do the order

3:25:49 by on the container itself and we didn't have that there.

3:25:52 So let's go ahead and do the same query on the optimize path.

3:25:57 So an optimized collection where we actually do have the indexing policy.

3:26:01 It was able to leverage the indexes and actually it's telling us that I forgot

3:26:06 to add a composite index here but we'll see an example of that maybe next.

3:26:10 So we can see that it was able to execute.

3:26:12 It's actually relatively performant comes back

3:26:15 27 milliseconds of query execution time.

3:26:17 That's a pretty fast query.

3:26:21 All right, so now let's go back to our slides and we'll go to the next scenario.

3:26:26 So we saw basic query operations with filtering and order by clause.

3:26:30 Next what we'll do is look at composite indexes.

3:26:33 Composite indexes are really useful when I have

3:26:37 where clauses or filters that cover multiple different

3:26:40 properties and or an order by with multiple

3:26:44 different properties in the order by as well.

3:26:46 Right.

3:26:47 So adding complexity to my queries really makes

3:26:49 composite indexes a really useful pattern for me.

3:26:52 Right.

3:26:52 So maybe if I'm searching through my product catalog for you know,

3:26:56 the latest backpacking tents.

3:26:57 So I want to search for some keywords and I also want

3:26:59 to have an order by clause in there for the most recent items.

3:27:03 Or I want some you know price order by along with my query filters.

3:27:07 There's really great use cases for composite

3:27:10 indexes and you know the composite index needs

3:27:14 to match the order and sort direction in which

3:27:16 you're going to leverage it in the query.

3:27:18 So that, just keep that in mind.

3:27:19 And of course you have to add composite indexes.

3:27:21 It's not index by default in Cosmos db

3:27:24 Okay so let's clear our screen here and what we're going to do is we're going

3:27:32 to go run the demo scenario with composite indexes.

3:27:39 So let's spin this up and if you're curious,

3:27:41 this tool that I'm using in the terminal, it's going to be available on GitHub.

3:27:45 There's actually a link and a QR code at the end of this, this presentation.

3:27:49 So you know, don't need to bother taking notes or anything like that.

3:27:52 You have access to all the code at the end of this talk.

3:27:55 Okay, so we're going to test a, couple different scenarios here.

3:27:58 We're going to look at composite,

3:27:59 equality and range filters and then we're going to look

3:28:01 at composite filtering and composite sorting in the same query.

3:28:05 So in this example here we're going

3:28:07 to select some properties and we want to make

3:28:09 sure that we're scoping to a particular shop

3:28:11 and scoping to a particular category for those products.

3:28:15 And we want the price, you know, to be greater than you know, something.

3:28:20 Right?

3:28:21 So that something's going to be 100 so 100 US dollars, right?

3:28:25 And so let's go ahead and execute that query.

3:28:27 And again with all these scenarios,

3:28:28 we're first executing on the unoptimized collection with the collection without

3:28:32 any indexing policy and we can see that the results come back,

3:28:36 maybe it took some time for us to run, you know,

3:28:39 we can scroll up and see the metrics and it took

3:28:41 about 350 RUs to execute that query, which is not cheap.

3:28:44 And we'll go and run this on the optimized container and it's going to execute

3:28:48 much much faster and we're going to skip

3:28:50 directly to the side by side comparison.

3:28:52 Look, there's no, there's no contest here, right?

3:28:56 13x more than 13x improvement in server latency, so query execution latency,

3:29:01 more than 7, almost 8x improvement in RU charges.

3:29:05 It's no brainer that you want to use composite indexes,

3:29:07 for these sort of complex query filter scenarios.

3:29:10 Now when we compare in the next scenario using a composite filter and sorting,

3:29:17 we'll have sort of the same issue that we saw last time.

3:29:19 Right.

3:29:20 For composite indexes I really excuse me, for composite order bys,

3:29:24 I really need a composite index to be able to execute that effectively.

3:29:27 So we're not able to execute

3:29:29 this on our container that has no indexing policy whatsoever.

3:29:32 But when we run this on the container that does have the indexing policy,

3:29:35 not just the standard indexing policy but also our composite indexes,

3:29:39 we're able to realize that it's still incredibly performant, right?

3:29:44 16 seconds.

3:29:45 It took 16 seconds for this query to run on over 10,000 data items.

3:29:49 And it cost about 47 RUs.

3:29:51 So much, much cheaper than we would expect, for any of these operations,

3:29:56 for not using sort of an optimal

3:29:59 way and best practice of leveraging composite indexes.

3:30:04 Okay, so let's talk about the next area.

3:30:07 So we covered basic queries and composite indexes.

3:30:10 Now let's jump into arrays.

3:30:12 So Cosmos DB can index the content in an array.

3:30:17 And then you can use query functions like array contains or you can

3:30:21 reference certain positions of the arrays or you can even do these joins

3:30:25 like intra document joins and you can use exist subqueries and all sorts

3:30:29 of great functions to resolve from the index

3:30:32 instead of scanning every single document, for the contents of those arrays.

3:30:36 Right.

3:30:37 So for example, maybe my products have different tags like a waterproof

3:30:41 tag or a tag on what seasons that product can be used for.

3:30:46 And these tags can power, you know, filter navigation.

3:30:49 So imagine that a customer is searching

3:30:50 for products they want to filter by other countries,

3:30:52 categories or their characteristics in this product.

3:30:56 Important thing to know is that Cosmos DB is schema free.

3:31:00 Right?

3:31:00 It's essentially JSON documents.

3:31:02 So instead of joining data across different

3:31:03 entities like you would in a relational database,

3:31:06 joins occur within a single item.

3:31:07 So there's scope to that item.

3:31:09 You don't do cross item or cross container joins.

3:31:13 It's really meant to sort of bring up and interact

3:31:16 with the objects or items that are contained within an array.

3:31:20 Okay, so let's see an example of this.

3:31:23 So we will switch back here to our terminal

3:31:28 and then we're going to run the demo scenario for arrays.

3:31:34 Let's go ahead and load this up.

3:31:36 And so we're going to test two different scenarios.

3:31:38 So we'll run a scenario where we're leveraging the array contains properties,

3:31:42 we're searching an array to see if it

3:31:43 contains some specific value or object and then

3:31:46 we're going to see an example of join and how that can be used as well.

3:31:49 Great.

3:31:51 So in this example, maybe we're filtering on the tenant and we want

3:31:54 to make sure that the tags array contains a specific tag like waterproof.

3:32:00 So we're going to go ahead and run that without any index

3:32:02 whatsoever on our products collection and we'll see the results that come up.

3:32:08 And we see the results and of course you know,

3:32:11 we see that it was looking for indexes

3:32:12 that it couldn't find in the indexing policy

3:32:14 because there is no Indexing policy and we can see it took about 411 RUs to run.

3:32:20 So we'll go ahead and run this on the container with the index

3:32:24 defined and we'll see that the same results come up here.

3:32:29 And we'll see that this is actually a much, much cheaper and efficient query.

3:32:34 Right?

3:32:34 So we did use the indexes that we have defined and it cost you know,

3:32:40 significantly less ru.

3:32:41 So let's do a side by side comparison.

3:32:43 So over you know, 1.4x improvement in latency

3:32:46 and over 3x improvement in RU chart.

3:32:48 Right.

3:32:49 And these performance gains that you get from adding the index

3:32:52 only grow when the size of the data grows, right?

3:32:55 If I went, if I had instead of 10,000 plus documents,

3:32:57 if I had 100,000 or a million or 100 million documents, right,

3:33:01 you can imagine that the improvements that I'm seeing here would actually be

3:33:05 many fold more than what I'm getting just from like the small scenario.

3:33:11 Next let's see an example of join.

3:33:14 So we have select, have a couple properties like

3:33:15 our name and description and then also the tag, right?

3:33:19 So we're getting the tag, we're getting tags from the C tag.

3:33:23 So C tags, tags is the array of different tags for that product.

3:33:28 So we're scoping again to the tenant ID and we're looking for a particular tag.

3:33:31 So this time we're going to again look

3:33:33 for waterproof as a particular tag using a join syntax.

3:33:37 So again we're going to run without

3:33:38 the index and then once this executes we're going

3:33:42 to just go right ahead and run it

3:33:44 on the optimized collection that does have indexes that were,

3:33:47 you know that, that we chose specifically where that Actually

3:33:51 to be honest this is the standard indexing policy, right?

3:33:54 By default Cosmos DB indexes all the properties

3:33:56 by default for range and inequality inequality filters.

3:34:00 And so I actually don't need to add anything specific for array operations here.

3:34:05 The one in the queries that I'm showing you,

3:34:07 this is part of the default indexing policy.

3:34:09 And so we can see that without the index versus with the index.

3:34:12 So with the index you get over a 5x

3:34:14 improvement on latency and over 3x improvement on RU charge.

3:34:18 So again just demonstrating the power of having an index policy for my arrays.

3:34:29 And now what we're going to do is

3:34:30 we'll move on next to talk about computed properties.

3:34:33 So computed properties let you define derived

3:34:35 fields from query expressions on the container Properties,

3:34:38 So they're not materialized, at read, time or query time.

3:34:43 They're not computed at query time,

3:34:44 but they're materialized and indexed at break time.

3:34:47 Right.

3:34:47 And the properties aren't persisted in the items themselves,

3:34:50 which is really makes it really nice for efficient storage.

3:34:52 Right.

3:34:53 So imagine that in our product catalog, imagine that we're storing,

3:34:56 you know, the different category path in a sort of compound string.

3:35:00 Like maybe our string is camping and tents and backpack tents.

3:35:04 And what we want to do is maybe we want to filter on some,

3:35:07 some, you know, subset of these.

3:35:08 Right.

3:35:09 Some, some, some subset of, of, of these categories that are stored in a string.

3:35:13 So without a computer property,

3:35:14 maybe I'd write a query that separates these out at runtime in the query itself.

3:35:19 And that could lead to, you know, when I'm doing a lot of queries,

3:35:22 I'm just doing that operation over and over and over again in the query,

3:35:25 which is not the most efficient way.

3:35:27 So in a computer property, I can do this, I can write it once and then

3:35:30 I don't have to worry about it at query time.

3:35:32 And it's important to note computer properties.

3:35:35 Like you have to define these yourselves and they

3:35:36 need to be explicitly added to the indexing policy.

3:35:39 But once you do this, then you can reference them

3:35:41 just as you would any other property in, your document.

3:35:44 Except they're not property of your document, they're a computer property.

3:35:48 So let's see what that looks like.

3:35:52 So we'll go back to our terminal and we'll clear the screen,

3:35:58 we'll go ahead and run this computed properties workflow.

3:36:01 Okay, so a bunch of stuff's popping up here.

3:36:03 So we're going to look at a derived field lookup.

3:36:07 So, here's the query that we're going to run without computer properties.

3:36:10 We want to project a couple different properties of the products.

3:36:14 And then of course, we're going to scope to our, tenant or shop.

3:36:18 And then we have, in addition,

3:36:19 we have this additional part of the where clause where we're going

3:36:21 to look at the category name and we're going to take a substring,

3:36:26 and take the first, element of it.

3:36:28 So, right, we're going to trim, this as well.

3:36:31 And so what we're doing is we're basically getting

3:36:33 the first category that's listed in that string, right?

3:36:36 So we're going to call this the primary category.

3:36:38 So maybe the first category of this product is the most important,

3:36:42 the main category of this product.

3:36:43 Unfortunately, all the products categories are Maybe you know,

3:36:47 because of my data model or the way that my application is designed,

3:36:50 it's stored in a way that is in a string rather than a array.

3:36:55 And so maybe I want to search for a particular,

3:36:57 just what this primary category might be.

3:37:00 So without computer properties,

3:37:01 this is what my query would look like with computer properties.

3:37:05 I can actually just reference a computed property that I create

3:37:08 and I can reference it like any other property in my document.

3:37:11 C cprimarycategory.

3:37:14 And what I can see here is that when

3:37:15 I define a computer property from my collection,

3:37:17 I can give it a name and then basically I

3:37:20 can give it a query that defines a computer property.

3:37:24 Right?

3:37:24 So the logic that we had here that was previously

3:37:27 in our unoptimized query is now in a computed property.

3:37:31 So now this can compute it and store it

3:37:34 in the index and makes it really efficient for search and retrieval.

3:37:37 And I don't have to rewrite this in every single query.

3:37:40 I don't have to rerun this logic every single time, during query execution.

3:37:44 So it should make our queries more performant.

3:37:46 So let's go ahead and run this.

3:37:47 So we'll run without the computer properties.

3:37:50 We'll go ahead and execute again on our product container.

3:37:56 And again so we're going to actually,

3:37:58 so in this example we're actually going to show

3:38:00 the difference between a computer property and no computer property.

3:38:03 So we're still going to leverage the index in both of these scenarios.

3:38:06 But one time we'll be using,

3:38:08 in this first one we're not using computer properties.

3:38:11 And then, Right, and so we can see the characteristics of this.

3:38:15 Right, so we can see there's 378 RUs.

3:38:18 And then in this optimized version we're

3:38:20 going to be using the computed property itself.

3:38:22 So again the first query did not use computer property,

3:38:26 but it still used the indexing policies.

3:38:29 Now in this second run we're going to use the computer properties and the index

3:38:35 computer property as well and we'll look at the side by side comparison.

3:38:39 Okay, so the execution time of the query was 88 milliseconds.

3:38:44 88.77 milliseconds with computed property.

3:38:46 So over 3x performance improvement with computed property in terms of latency,

3:38:51 1.45x improvement from RU charge.

3:38:54 And again as your data set scales this benefit only increases.

3:38:58 So really, really interesting application of computer properties here.

3:39:02 Right.

3:39:02 It's a really great feature that if you're

3:39:04 doing this complex logic over and over again,

3:39:06 in your queries you can really simplify your workflow,

3:39:09 and get better performance out, of your queries by using the computer property.

3:39:17 All right, so what do we have next?

3:39:19 Geospatial.

3:39:20 Great.

3:39:20 Geospatial data enables you to do all sorts of cool filtering

3:39:24 or bounding or sorting by physical distance by having locations like longitude,

3:39:29 latitude locations, on Earth.

3:39:31 This allows you to use all sorts of complex

3:39:34 geospatial functions that we have in the cosmos.

3:39:36 DB Query syntax.

3:39:37 So in this example we're going to use the stdistance function.

3:39:40 So a query common example that you find in e retailers, right?

3:39:44 You are looking for products and you want to make sure that you maybe you can

3:39:48 find a store nearby that has this product in stock that you want to go pick up.

3:39:52 Right?

3:39:53 And again, you know, spatial data here needs to be valid GeoJSON.

3:39:57 And of course you need to add it to your indexing policy.

3:40:00 By default, custom, DB doesn't index geospatial data,

3:40:03 but it's really simple just to add it to your indexing policy wherever

3:40:06 you're going to store whatever property you're going to store your GeoJSON data.

3:40:10 Okay, now let's see this in practice in our demo code.

3:40:14 So first we're going to clear a terminal.

3:40:17 So we'll go ahead and do that.

3:40:22 Then we're going to go and run the geospatial scenario.

3:40:26 So what are we going to take a look at in this scenario?

3:40:29 Well, we're going to look at a couple of things.

3:40:32 So first we're going to do this distance filter.

3:40:34 So maybe in this scenario we want to make sure that the product

3:40:39 that the customer is looking for is in stock in a store that's close to them.

3:40:43 Right?

3:40:43 So we are going to have a couple things in our query, right?

3:40:48 So we're going to project the id, the name, the description and the distance,

3:40:52 or the geospatial distance as this distance meters, alias.

3:40:57 And then we're going to have a where clause on the tenant id.

3:41:00 So we're going to look for a specific, tenant or a specific shop.

3:41:04 And then we're going to, we want

3:41:05 to make sure that it's within this maximum distance.

3:41:08 So we want to make sure it's within, you know, say 50, kilometers, of our user.

3:41:14 And then we're in order by ST distance.

3:41:16 Right.

3:41:16 So what this is going to do is going

3:41:18 to rank the results in order by their geospatial distance,

3:41:23 from the user to the stores.

3:41:25 So let's go ahead and run this query

3:41:27 and we're going to run it again first without

3:41:29 the index and we can see some of the results

3:41:32 that come back and some of the characteristics.

3:41:34 But let's just skip this and we'll go ahead and run it

3:41:37 with the index and then we'll look at the side by side comparison.

3:41:41 So side by side comparison shows something really interesting.

3:41:43 Right?

3:41:45 Without the index we see about 89 millisecond of execution time on the server.

3:41:49 With the index 55.88 milliseconds.

3:41:52 So it's a 1.61x improvement with the index.

3:41:56 And then of course the RU charge as well,

3:41:58 we get a 2.34x improvement on RU charges.

3:42:02 Right.

3:42:03 And again this only increases, right?

3:42:05 The improvement level only increases as the scale of my scenario.

3:42:09 So the more data that I have, especially if I have really high throughput

3:42:13 account or account with lots of different partitions.

3:42:15 Right.

3:42:16 So we get a lot of benefit by using

3:42:18 the index and also by specifying the partition key,

3:42:23 in that where clause as well.

3:42:24 So now that we talked quickly about geospatial,

3:42:26 let's look at some of our search capabilities.

3:42:29 So vector, full text and hybrid search.

3:42:31 Let's start with vector search.

3:42:32 So vector search uses vector embedding and our vector

3:42:36 distance function combined with a vector index.

3:42:38 So in Cosmos DB we have multiple vector indices including disk ann,

3:42:41 which allows you to do semantic search with very

3:42:44 high accuracy and low latency really at any scale,

3:42:48 even up to billions of vectors.

3:42:50 So imagine that a shopper wants to search for something,

3:42:54 you know, a product and uses some, you know, text phrase, right?

3:42:58 Like lightweight tent for rainy hikes.

3:43:00 So the vector search will allow us to find semantically related products,

3:43:03 not just products that match these keywords or phrases and support a note.

3:43:09 So vector search is really good to add

3:43:11 a vector embedding policy and vector index.

3:43:13 So you could add like DiskANN index.

3:43:15 And if you want to do multi tenant search, like for example here,

3:43:18 if I wanted to search specifically

3:43:19 on a particular tenant or our particular shop,

3:43:23 I could shard our vector index to that, you know,

3:43:26 by the different shops or tenant IDs.

3:43:29 And so I could isolate my vector search just

3:43:31 to those shops and you can get really performant,

3:43:34 you know, tenant isolated search that way.

3:43:37 So let's see an example of vector search and practice.

3:43:40 Okay, so we'll clear screen again.

3:43:47 And we'll do a run of the Vector scenario.

3:43:52 So here our index.

3:43:54 Excuse me, here our query is very straightforward.

3:43:56 We're selecting the top five most similar items and we're going to project

3:44:00 the vector similarity score and then we'll also order by the vector distance.

3:44:04 Right?

3:44:05 So we'll order by the vector similarity score from most similar

3:44:08 to least some more and we'll run it with and without the index.

3:44:19 So we ran it with and without the index.

3:44:22 We'll skip right to the comparison and we can see the server time.

3:44:25 So almost 2024, almost 25x improvement in server side latency using

3:44:31 the vector index compared to not using a vector vector index.

3:44:34 And then the RU charge is also much much lower.

3:44:37 Right?

3:44:37 Not even 1615 RUs on vector search on over 10,000 documents.

3:44:44 So really, really performant using that to scan in vector index.

3:44:49 Going back to talk about full text search real quick.

3:44:52 So full text search uses different linguistic

3:44:55 analyzers so tokenization and stemming and functions like

3:44:58 full text contains and full text score

3:45:01 to match and rank documents by keyword relevance.

3:45:04 Right.

3:45:04 So I want to make sure that certain

3:45:07 products contain a certain phrase or keywords.

3:45:09 I can use full text contains if I want to search for Items using this BM25

3:45:13 scoring which is a measure of frequency of whether

3:45:17 these terms are in the documents or not.

3:45:20 I can do an order by rank full text score, clause as well.

3:45:24 So we'll see examples of that.

3:45:26 Right.

3:45:26 So an example here is when a customer wants to do a search

3:45:28 and you want to search by those text keywords or by those text phrases.

3:45:33 This is really powerful tool for enabling that.

3:45:35 Okay, so we'll run this scenario as well.

3:45:37 So we'll clear a terminal again and go ahead and run the full text demo.

3:45:45 So here's what we're going to test the full

3:45:47 text contains and full text scoring using BM25.

3:45:50 So we'll do a simple query here where we have we want

3:45:52 to make sure multiple terms are contained

3:45:54 within the products that we're looking for.

3:45:56 We want to make sure that it contains waterproof and conditions and windproof.

3:46:01 And so we're to run this without any index whatsoever.

3:46:04 And then we'll run this with a full text index

3:46:07 defined on that collection and we can see that you know,

3:46:11 it's 2.3x more performant and uses far fewer RUs.

3:46:15 And again the difference here the benefit

3:46:17 only increases as your data size scales.

3:46:20 Next we'll run the full text scoring.

3:46:22 Right.

3:46:23 So what we're going to do is we're going to search

3:46:24 for these items and we're going to order by their BM25 score.

3:46:28 So how relevant the products are by how often these keywords

3:46:32 appear in the descriptions or the names, of these items.

3:46:36 So again, running without the index versus with the index,

3:46:40 a 30x improvement in server latency using the index

3:46:44 and a 116x improvement on RU charge by using the index.

3:46:48 Right.

3:46:48 So, this really demonstrates that, you know,

3:46:52 we're able to really effectively use the index

3:46:54 once we do these advanced search or ranking,

3:46:57 methods for lexical or full text search.

3:47:00 The index is incredibly powerful for doing these very efficiently at scale.

3:47:07 And finally we have hybrid search.

3:47:09 Hybrid search combines multiple different measures like vector similarity

3:47:13 and full text scores into a single ranked result.

3:47:17 So for example, if you want your shoppers to be

3:47:20 able to make sure that maybe your searches contain certain keywords,

3:47:24 but you also want to capture the semantic

3:47:26 similarity from vector search to find related products.

3:47:29 Hybrid search could be a really great approach.

3:47:31 And again, you have to define both

3:47:32 the full text and vector index to utilize this.

3:47:35 And then the RRF function merges, the two scoring methods together.

3:47:43 Okay, so we'll clear the screen and run our final demo scenario here.

3:47:52 Okay, so we're going to see an example of the hybrid search.

3:47:55 So we're going to select the top five items

3:47:58 from our product catalog and we'll also project the vector similarity score.

3:48:02 This time we're going to order by rank.

3:48:04 We're going to use the RF function to fuse together

3:48:07 results from full text score and from vector similarity score.

3:48:11 And we're going to do this without using an index and with using an index.

3:48:14 So we'll go ahead and run this first without using the index.

3:48:21 And then we'll run this again while using the index.

3:48:25 And side by side comparison we see,

3:48:28 the server time 1 point 1 3x so faster and then ru charge is 2.8,

3:48:34 almost 3x, cheaper to run.

3:48:37 Right.

3:48:37 And again, these improvements, you know,

3:48:40 the server time looks relatively modest,

3:48:41 but you have to keep in mind that this is just on 10,000 data items.

3:48:44 As I keep mentioning, as the size of your data set grows,

3:48:48 benefit of using the index also increases.

3:48:50 The benefit of scoping to a particular partition also increases.

3:48:53 Right.

3:48:54 Especially if you're using any sort of search

3:48:56 like vector similarity search or the hybrid search.

3:48:58 Right.

3:48:59 Our vector indexing algorithms,

3:49:00 allow you to search in a way that scales sublinearly.

3:49:04 So as your data set grows, maybe your dataset grows 200 times in size,

3:49:08 your vector similarity search won't even increase by 2x in latency.

3:49:13 It's that efficient.

3:49:14 It really scales sublinearly with the size of your data.

3:49:17 So keep that in mind that as your scenario grows,

3:49:21 as your E retail shop becomes more and more popular and you

3:49:24 have more and more products and deal with more and more customers,

3:49:26 you're going to really retire the performance benefits

3:49:28 of using proper indexing techniques and query techniques.

3:49:35 Okay, so that's all we have today.

3:49:38 There are some references here to get you started.

3:49:40 All the material from this session, including the slides and the code examples

3:49:43 that I ran along with the sample data, are available at aka Ms.

3:49:48 Cosmos DB conf 26 query.

3:49:51 If you're interested in learning more about anything we talked about here,

3:49:54 whether it's queries, vector searches, full text search, hybrid search,

3:49:57 there's also some great learn documentation to get started.

3:50:00 And we also have a QR code that if you want

3:50:01 to scan that with your phone or scan it with your browser,

3:50:04 and it'll take you immediately to the GitHub repository

3:50:06 where we have all this content available for you.

3:50:09 All right, that's it.

3:50:11 Hope you enjoyed the content.

3:50:13 Thanks a lot folks.

3:50:14 See you next time.

3:50:18 Hi everyone, I'm Jess Ramos, or you might know me as Jess Ramos.

3:50:21 Data Online.

3:50:23 I'm the founder of Big Data Energy and I create data

3:50:26 and AI content for over half a million followers on social media.

3:50:30 Right now, everyone's talking about how good AI is.

3:50:33 You know, which model is faster, smarter, cheaper.

3:50:36 But for me personally, I can't stop thinking about the data.

3:50:40 We've all heard the saying garbage in, garbage out.

3:50:44 And this has never been more true than it is right now,

3:50:47 because you can have the most powerful model in the world

3:50:51 and still get completely unreliable outputs if your data layer is not solid.

3:50:57 AI doesn't run on neat structured tables.

3:51:00 It runs on vectors, documents in real time,

3:51:04 streaming data that's messy by nature and massive by scale.

3:51:09 And that's exactly where NoSQL shines.

3:51:12 We're talking AI search, real time personalization,

3:51:15 rag pipelines and agentic apps that need to read and write data fast.

3:51:21 And this is the stuff that actually

3:51:22 breaks your infrastructure if it's not built right.

3:51:26 Cosmos DB is built for exactly that.

3:51:28 Global distribution, single digit millisecond latency,

3:51:31 and both operational NoSQL and vector data.

3:51:35 And I'm so excited to keep learning more about this@ Cosmos DB conf.

3:51:39 Thanks so much for coming and let me know on LinkedIn what you learned today.

3:51:42 Bye hey, welcome back.

3:51:47 We are so glad you're here with us.

3:51:49 We've had a great start.

3:51:51 Well, we got about an hour to go.

3:51:55 Wow.

3:51:55 Yeah, thank you so much for sticking around and we have so

3:51:58 much more to show you as well as just the on demand sessions.

3:52:02 And you know what, we've seen so many things,

3:52:06 so many demos to help us move along,

3:52:08 we've got another amazing live session with none other than Andrew Lu.

3:52:13 You might have seen him earlier today.

3:52:14 Yeah.

3:52:15 Andrew is going to help us learn even more

3:52:17 about how Azure Cosmos DB works under the hood.

3:52:20 Andrew, welcome back to the stage.

3:52:22 Cool.

3:52:22 Hey, nice to be here.

3:52:24 Hey everyone, we're going to do a bit of an experiment.

3:52:28 We're going to do a whiteboarding session and we're going to go right into this.

3:52:32 So have you ever wondered like what happens when you go to the portal,

3:52:36 you hit Create Cosmos DB and just what happens behind the hood?

3:52:40 That's the intention of this session.

3:52:43 So let's start with our SDK and taking that perspective

3:52:49 from this point of the view of the world.

3:52:54 So the SDK is going to be embedded either in an application or even when we

3:52:58 build things like the Azure Portal we view

3:52:59 that as a client from the databases perspective.

3:53:02 So think of this as the app or the portal.

3:53:05 This is wrapping our SDK.

3:53:07 We actually have multiple flavors of the SDK beyond just programming languages.

3:53:10 We have a management SDK as well as a data plan SDK.

3:53:14 And what ends up happening is when we do from a client server,

3:53:17 perspective behind the scenes what this is going to go and talk to is a control

3:53:21 plane and a data plane and we're going

3:53:23 to think of these two things as very separate.

3:53:26 The control plane is going to be wrapped over arm.

3:53:33 This stands for Azure Resource Manager.

3:53:35 So whenever you make REST calls over into a DNS that looks something

3:53:39 like management.uhazure.com what happens behind the scenes

3:53:44 is you can think of arm, one mental model is like a big reverse proxy

3:53:49 and that reverse proxy when it goes hey, architecture, okay,

3:53:51 you want to create a Cosmos DB or add a region

3:53:54 or failover to a different region or scale your RUs.

3:53:57 What that's going to do is it's going to go

3:53:59 route those requests into what's called a resource provider.

3:54:02 So Cosmos DB has a resource provider

3:54:04 and of course every other Azure service out there also

3:54:06 has its own resource provider Arm's job is

3:54:10 to go and route to the correct Resource provider.

3:54:13 Now on the data plane this is going to be a different DNS entry.

3:54:16 This is going to be Cosmos Azure.

3:54:20 And the data plan is responsible for what you think of a database actually does.

3:54:24 Right?

3:54:25 This is going to be writing your data,

3:54:27 indexing it and then making it available for fast efficient queries.

3:54:32 So when we go and look at that path,

3:54:40 All of these requests will first flow into our front end server.

3:54:44 In Cosmos DB we like to call this our gateway.

3:54:48 The gateway's job is to do request routing because what happens is

3:54:51 we have a distributed backend behind the scenes in the back end.

3:54:59 What this will map to is effectively big clusters that we run behind the scenes.

3:55:04 They are multi tenant federations or clusters.

3:55:09 We like to use those words interchangeably within those clusters.

3:55:15 This is going to be the deployment model.

3:55:20 In Cosmos DB so if we go look

3:55:22 from a physical to a logical mapping on the logical

3:55:25 layer you're making this thing called a database

3:55:28 account down to a database down to a container.

3:55:33 I might use different words

3:55:34 like container or collection interchangeably container.

3:55:37 We use that as the API agnostic way of saying that.

3:55:40 And then depending on some of the Cosmos DB APIs,

3:55:42 you might have seen the word collection especially for any

3:55:45 of the document oriented data models within that we store our data.

3:55:50 Everything within this container is what is being

3:55:53 powered by the backend on the data plane.

3:55:56 Behind the scenes what this container is, it's not a single server.

3:56:01 What we're doing is we're stitching together replica sets.

3:56:05 So think of I have a little replica set powered by service fabric,

3:56:12 where I have a primary and a few different secondaries.

3:56:15 Each of these is a database, replica a database replica is going to be the core

3:56:22 building block for what's powering the software behind the scenes.

3:56:26 This is going to have the query engine, it's going to have the indexing layer,

3:56:29 it's going to have the, the storage and beyond.

3:56:32 Just doing the bread and butter of what a database does.

3:56:35 It's also going to have other things like admission control

3:56:37 and that's where we're going to have things like our resource governance.

3:56:41 So think of this what we're doing is we take this software,

3:56:45 we stitch it together into a replica set and on that replica

3:56:50 set this is what's going to form a physical partition.

3:56:56 And what happens is when you scale a container

3:56:59 we're going to create one or more physical partitions.

3:57:02 Each physical partition today is able to Support up

3:57:06 to 10 thousand request units per second worth of bandwidth.

3:57:09 It's also able to support up to 50gb worth of storage.

3:57:13 So when we talk about infinite scale, like millions of RUs,

3:57:16 what we're really doing is we're just taking

3:57:18 that container and we're sharding it aggressively where we're creating

3:57:21 tons of these, these different physical partitions behind

3:57:23 the scenes now these get placed onto these big federations.

3:57:27 So back to from a federation perspective,

3:57:29 think of this as in the data center I have a bunch of racks.

3:57:33 Each of these racks are going to have some

3:57:36 shared infrastructure like a top of rack network switch.

3:57:39 And we can also think of those in terms of helping us

3:57:42 compose what will become a set of fault domains or update domains.

3:57:47 So when we set up out these federations,

3:57:48 that go and host all these different database replicas,

3:57:52 what we're doing is on these different nodes,

3:57:55 think of these nodes as basically spinning

3:57:58 up these replicas and having them on standby.

3:58:00 And so when you go through the control plan saying hey,

3:58:02 I need to go scale my Cosmos db give me some more ru.

3:58:05 Behind the scenes what that resource provider is

3:58:07 doing with our control plan is saying hey,

3:58:10 this workload needs to scale and you some more partitions.

3:58:12 Do we have some bindable partitions?

3:58:14 Let's go take a look at this federation.

3:58:16 We might have a replica stitched together

3:58:19 across different fault domains and update domains.

3:58:22 So let's say I have a couple of different nodes here.

3:58:28 We're going to stitch together a replica from each one of these into a bindable

3:58:32 partition and we're going to give that partition over to that container.

3:58:35 Now these federations don't think of as an account

3:58:39 as having to live on a single federation.

3:58:41 We span across many federations, for infinite scale.

3:58:44 And why we do that is we decouple the physical and logical

3:58:48 layer and this allows us to always stay current on our hardware.

3:58:51 Speaking of the hardware, this is a nice shout out to our sponsor amd.

3:58:56 We use especially in the newer generations of our federations,

3:59:00 a massive amount of AMD CPUs.

3:59:02 Why we love the AMD CPUs is because they're incredibly power efficient.

3:59:06 They give really good watts, for the unit of compute,

3:59:09 that we give as well as that helps us drive the cost down.

3:59:12 And as we drive the cost down this is what we're also able to go

3:59:16 and Turn that into cost savings for our users

3:59:19 like you in the form of new features.

3:59:22 So things like the dynamic auto scale,

3:59:23 the RU pooling with the new Cosmos DB fleets the Cosmos DB serverless,

3:59:28 those Effectively what we're doing is,

3:59:31 is we're using elasticity to floor the TCO of operating a database.

3:59:36 Every workload is going to be bursty, any transactional workload is like that.

3:59:41 What we're doing is we're passing all of the savings we

3:59:43 get from the AMD into nice elastic features for our users.

3:59:48 Now on the federation,

3:59:49 as we go and stitch together these replicas we can jump into our first I

3:59:55 think bigger meteor topic and that is going

3:59:57 to be how we look at the partitioning.

4:00:08 So Cosmos DB uses a technique called consistent hashing.

4:00:11 It's a fairly ubiquitous concept for any kind of sharded database.

4:00:17 It tends to be the industry standard and I'll

4:00:19 kind of walk through that in terms a little bit

4:00:21 just to help understand it and compare it to what

4:00:25 maybe folks will think of as like a naive hash.

4:00:27 So in a naive hash think of it as I'm going to go partition my data

4:00:31 and behind the scenes that container now

4:00:33 has a ton of these physical partitions, right?

4:00:36 Like a container has let's say 50,000 RUs.

4:00:40 We know each one can support 10k each.

4:00:43 So this is going to go and create

4:00:46 five of these physical partitions behind the scenes.

4:00:48 Means it's going to go then route let's say some data onto this node,

4:00:53 some data over into this or this physical partition this physical partition.

4:00:57 How does that know?

4:00:58 And to deterministically route if I go and query

4:01:01 for let's say Alice's data and a user based system,

4:01:05 I need to be able to actually hit the partition hosting Alice's data.

4:01:08 Otherwise when I do that query I'm not going to find anything.

4:01:12 So what happens is we ask Hugh for, for a partition key.

4:01:17 So let's put in some values here just to give you an idea of what is happening.

4:01:25 So we have a logical partition key.

4:01:27 This is the partition key that you're supplying

4:01:29 and that's going to go hashed to some numeric value.

4:01:33 Behind the scenes we can use different hashing algorithms.

4:01:37 In this case Cosmos DB uses this algorithm

4:01:40 called Murmur hash because it's very compute efficient.

4:01:45 Let's look what happens if you do a naive hash like in a naive approach.

4:01:49 Let's say if I only had two partitions.

4:01:51 The challenge you're going to run into in a naive hash is if I,

4:01:54 let's say go and let's say mod two.

4:01:57 Well what I'm going to see is a distribution that looks like 1 2, 1 2, 1 2.

4:02:01 Now what happens if I need to grow

4:02:04 that on an elastic workload I'm scaling that over to three partitions

4:02:08 and right away you're going to go see what

4:02:11 the problem is and why a consistent hash is so valuable.

4:02:15 So when I do a mod 3 I'm going to get a distribution

4:02:17 that looks a little bit more like 123-123- Notice what's happening here.

4:02:24 If I am mapping data, let's say Alice,

4:02:27 Bob, Carroll, let's see, put Danny, Eve, et cetera.

4:02:32 Notice this challenge where if Danny Mapped to partition 2

4:02:35 when we mod 2 but now maps to partition 1,

4:02:39 we're doing effectively a huge shuffle.

4:02:41 A lot of data would have to go from the second partition into the first,

4:02:45 from the first into the second,

4:02:48 from the first into the third, the second to the third.

4:02:50 That's not very efficient.

4:02:51 So this is why a consistent hashing scheme is a much

4:02:55 better way of doing this than rather than a naive hash.

4:02:58 But this will give you the intuition in terms of how this works out.

4:03:02 So a consistent hash you can think of as a range

4:03:05 partitioning on top of a hashed based scheme.

4:03:11 So if I go and think about the hash values

4:03:13 where this is the min value on a number line,

4:03:17 and I do this as a circle since everything is going to go map onto this ring,

4:03:22 if I look at min value on where it hashes to 000 to ffff.

4:03:27 If I have two partitions on day one then

4:03:31 the way of scaling this is actually very easy on mapping.

4:03:36 What we can do is we can just go chop the number line, right?

4:03:40 And so from zero to approximately half and then

4:03:43 from half to let's say plus one over to max value,

4:03:46 I can go map this into physical partition one

4:03:49 and this can go map to physical partition two.

4:03:51 Now what's neat here is in a consistent hash,

4:03:53 rather than doing any of this shuffling,

4:03:55 what happens if the first partition is filling up

4:03:57 or what happens if the second partition is filling up?

4:04:00 Well then we can go and dynamically allocate more partitions by splitting them.

4:04:04 So what happens is let's say I need to do a split

4:04:07 where I need to go and map Now a third partition.

4:04:10 So we can go from zero to let's say quarter

4:04:14 and then quarter to quarter plus one over to half.

4:04:19 Well notice that I can now relocate, let's say, or not even relocate any data.

4:04:23 I can just go in on the hash values, directly.

4:04:27 If Alice was here, Bob was here and Carol was here,

4:04:32 rather than having to shuffle things around, I can just say hey,

4:04:35 Alice and Bob used to be co located on partition one,

4:04:38 but partition one has now turned turn into partition one,

4:04:41 let's say prime and partition one double prime.

4:04:44 All we've done is split that data.

4:04:46 What's awesome about this is Cosmos DB can do

4:04:48 this as an online operation where we'll handle the splitting or merging

4:04:51 of the data on allocating the partitions and then dynamically cut

4:04:56 over whenever you need to go and perform one of these operations.

4:05:00 So what Cosmos DB does is really the runtime aspects of partitioning,

4:05:05 so that you have auto partitioning.

4:05:07 Really as a user, what you're left with is the design time aspect,

4:05:10 which is choosing a good partition key.

4:05:12 And this is why it's so important to choose a good partition key.

4:05:16 Because ultimately if you're partitioning, let's say by a user id,

4:05:18 then what we're doing is we're using that user

4:05:20 ID to go deterministically route which piece of data,

4:05:24 let's say Alice's piece of data into partition one prime,

4:05:27 Bob's data into partition one double prime and Carol into partition two.

4:05:33 Now this is one aspect of scale out.

4:05:35 This is the partitioning aspect.

4:05:36 The other thing I'd like to cover is the replication side.

4:05:40 What replication is really therefore is a, little bit different.

4:05:43 Partitioning is what gets you ability to scale your request traffic.

4:05:47 In particular it does really well at scaling your write traffic.

4:05:50 It can also scale your read traffic.

4:05:52 If I have millions of QPS,

4:05:55 well I cannot have an infinitely number of CPU on a single box on a single vm.

4:06:00 There's still the physical aspect of it.

4:06:02 We just put you on more computers and so what we're doing is we can power many,

4:06:07 many CPUs as long as we divide that data set and partition it up.

4:06:11 Now what replication is doing a little bit different is copying that data.

4:06:14 Copying that data gives us two different things.

4:06:16 One that's similar and one that's very different.

4:06:18 The similar thing is going to be it's one way of scaling your read throughput.

4:06:22 The other thing that it can do is it gives you a lot of Redundancy.

4:06:25 And this is how Cosmos DB gives you five nines of availability.

4:06:30 And we can think of this in multiple grains.

4:06:33 So let me go and get some board real estate.

4:06:41 On the replication side.

4:06:42 When we zoom in on any of these partitions,

4:06:45 remember what we have is a replica set.

4:06:50 What we're doing is we're using a concept of called

4:06:53 quorums or quorum replication to go and provide that availability.

4:06:56 Now how we do those replica sets is going to be

4:06:58 a little bit different in terms of within a region versus across region.

4:07:02 When you do across regions, what we're effectively doing is we're spawning

4:07:05 another partition in each of the different regions.

4:07:08 And then what happens is we can have multi write regions.

4:07:12 What happens is one of these secondaries will get designated as a forwarder.

4:07:17 We might have three and we might have four.

4:07:19 This is a bit black boxed.

4:07:21 But what we do is we dynamically size a quorum across a replica set.

4:07:25 And this can be in let's say region two,

4:07:27 this can be in region one, within a region.

4:07:33 We rely on surface fabric heavily to do the within a region, replication.

4:07:40 And what's beautiful about this is it handles

4:07:43 a lot of the things like leader replication.

4:07:45 For us, let's say a primary is going down,

4:07:47 whether it's willingly, intentionally.

4:07:50 Let's say we're deploying a database software update.

4:07:52 Have you noticed Cosmos DB doesn't really have maintenance periods.

4:07:55 It's just always online.

4:07:56 You've never seen us say hey,

4:07:57 it's time for a Windows update or it's time for an application update.

4:08:01 That's because behind the scenes what

4:08:03 we're doing is we're just dynamically changing.

4:08:05 Okay, let me take down a replica that's

4:08:07 on a different update domain or a different fault domain,

4:08:09 and let me designate and promote one of these secondaries

4:08:12 over to a primary and then shift that traffic over.

4:08:16 What we're doing is we're live doing rolling updates.

4:08:19 And that way when this replica is brought back online,

4:08:22 what we can do is we're just dynamically writing to a rolling set of replicas,

4:08:28 behind that replica set.

4:08:29 Now what a quorum, is a little bit

4:08:32 different from your traditional primary versus secondary replication.

4:08:37 In a traditional database like let's say a relational database,

4:08:41 what you're typically working with is

4:08:44 either synchronous or asynchronous replication.

4:08:48 Those come with pretty steep trade offs.

4:08:50 When you do synchronous replication.

4:08:52 If your goal was high availability, if the primary goes down, that's fine,

4:08:56 you can go and fail over to a secondary,

4:08:59 but actually goes and has some risks as well.

4:09:04 What if the instead of the first failing, it's the second one that went down?

4:09:09 Right?

4:09:09 Well, if you're synchronously replicating,

4:09:11 well even if the primary is healthy now I have

4:09:13 to go and cut that secondary in order to restore,

4:09:16 work back onto bring the right availability back.

4:09:21 So what synchronous replication is doing is it's giving,

4:09:25 giving you really strong consistency.

4:09:27 And in terms of disaster recovery terms,

4:09:30 what you're really prioritizing is a low recovery point objective, right?

4:09:35 Like if I go from the primary to the secondary,

4:09:37 since everything was synchronously replicated,

4:09:39 if I rotate the primary, I know it's in the secondary.

4:09:42 So I have a recovery point objective of effectively zero.

4:09:45 Now where synchronous replication is really

4:09:47 bad at is the recovery time objective.

4:09:49 If you are having a coordinate effect failover.

4:09:51 Failovers take time.

4:09:53 So this is where you're going to be less happy, as opposed to more happy.

4:09:59 And then asynchronous is the exact opposite, right?

4:10:01 In an asynchronous model,

4:10:03 what you're going to see is your recovery time objectives.

4:10:06 Well, hey, if the secondary fails, that's fine.

4:10:09 I can just keep serving out of the primary.

4:10:11 I'll just let the secondary lag behind.

4:10:15 But how far is it lacking behind, Behind.

4:10:17 If I do a cutover to the secondary, is the data that I just wrote even there?

4:10:23 So what you're going to see is the recovery point objective is going to be,

4:10:26 well it's an eventually consistent model.

4:10:27 So you're going to have a lack of consistency,

4:10:30 but you're trading up the recovery time objective.

4:10:33 Now where quorum replication,

4:10:35 innovates on this is rather than thinking of the world as hey,

4:10:39 let me just have one additional redundant copy or N

4:10:41 number of redundant where I have a primary secondary model,

4:10:44 what I can do is I can always write to a quorum of replica.

4:10:48 So imagine that if I have four replicas here,

4:10:51 a majority quorum would be something like writing

4:10:54 to at least three out of four of those replicas.

4:10:58 That means that if any replica were to fail, in this example, let's say,

4:11:03 this fourth replica bombed out,

4:11:05 and I was able to write to the primary and to these secondaries,

4:11:09 or even if that was primary dies,

4:11:12 we go and promote One of the other replicas to be the new

4:11:15 primary and we get three out of four on the new topology.

4:11:18 Any node can fail and we're able to maintain that availability.

4:11:24 Now on the read side this is what is really interesting is if

4:11:27 I go and circle any two replicas I can do a minority quorum.

4:11:33 Well because I'm allowing fault tolerance for, for any

4:11:36 replica to go down on the write path.

4:11:38 On the read path as long as I read

4:11:40 at least two different replicas I can still guarantee strong consistency.

4:11:45 Even though this is in a fault tolerant

4:11:46 system where I allow a replica to go down,

4:11:49 I'm always going to have replica overlap.

4:11:52 If replica D goes down well A and B will of course both have the data.

4:11:56 B and C would both have the data C and D.

4:11:59 The D would miss, but the C would still

4:12:03 because it's further forward in time it's always going

4:12:06 to win that vote that you're able to get

4:12:09 strongly consistent reads in the midst of a downed replica.

4:12:14 Now that's what we're doing within a region and then across regions.

4:12:20 What we do is we designate one of these secondary

4:12:23 replicas to go and replicate over to another region.

4:12:29 There's some slight caveats to that.

4:12:31 Depending on whether you're using strong consistency versus

4:12:33 any of the other Cosmos DB consistency levels,

4:12:35 if you're using any of the weaker consistency levels you can think

4:12:39 of it as we're doing an asynchronous replication over to the secondary region.

4:12:43 And this is what gives you well higher availability,

4:12:47 faster RTOs at the trade off of of consistency

4:12:56 or if you're using global strong consistency.

4:12:58 What we're doing is you can think

4:13:00 of it as synchronously writing to multiple regions.

4:13:03 Now there's a bunch of more things that go into that.

4:13:05 For example we can do dynamic quorums where if we have a third region

4:13:08 we can vote two out of three regions to give you even strong consistency.

4:13:12 That allows a region to go down.

4:13:15 And what's really cool is Cosmos DB allows it's not a single primary

4:13:19 system where able to put primaries in each of these two different regions.

4:13:24 On the SDK side you never have to really wait for a failover

4:13:28 if you're any of the weaker consistency levels using multi region writes.

4:13:32 If I'm writing to region one and region one goes down that's fine.

4:13:36 The SDK is actually smart enough when you instantiate the SDK,

4:13:39 yeah, that connection policy, you can set what the preferred location is.

4:13:43 What it's doing is it's going to do a live reading redirection

4:13:46 to the secondary region without even waiting for a server side failover.

4:13:50 This is how you can collapse that recovery time objective down to zero,

4:13:54 even in a strong consistency manner,

4:13:58 we're able to go and minimize the RTO where we now

4:14:01 have the ability of doing fast failovers in a strongly consistent setup.

4:14:06 We are very happy about.

4:14:08 Recently we had launched per partition automatic failover.

4:14:12 What it does is rather than looking at hey, if one partition's unhealthy,

4:14:15 another partition is healthy and trying to have

4:14:18 a live ops team go hey, should I fail over?

4:14:20 Should I not fail over?

4:14:21 That tends to be the bottleneck for any failover.

4:14:24 We don't have to go and do any

4:14:25 decision making in a strongly consistent system with ppaf.

4:14:29 What we can do is we can also

4:14:30 then trigger that one partition to independently trigger

4:14:33 a failover in this big distributed database without

4:14:36 having to impact all of the other healthy partitions.

4:14:39 So this is how we're able to go and try

4:14:40 to give you just softer trade offs rather than thinking

4:14:43 of the world as extremes in terms of eventual or strongly

4:14:46 consistent systems with like hey do I want RTO or rpo?

4:14:50 You can get both by using quorum topologies for replication,

4:14:57 and then doing a lot of magic around redirections.

4:15:01 Now a third thing I want to go and show you that I think is

4:15:04 good for every developer to understand in Cosmos

4:15:07 DB is the resource governance aspect, right?

4:15:10 Like when you go and create a Cosmos DB container you're going to see like hey,

4:15:12 what is this thing, request units.

4:15:14 This is what allows us to go and do paper request models around serverless.

4:15:19 It allows us to give you guaranteed throughput rather than just best effort.

4:15:24 Hey, we have some nodes with some cores,

4:15:26 we're able to give you predictable performance.

4:15:28 So what happens is each of those replicas, when I look at the replica,

4:15:32 where behind the scenes you can think of it as we have a resource

4:15:36 governor and you can think of that as a token bucket algorithm, right?

4:15:40 If you haven't seen token bucket algorithm before,

4:15:43 think of it as like of a bucket

4:15:44 of tokens and the request units are those tokens.

4:15:48 When you go and do a database read,

4:15:51 a one KB read is actually the baseline charge.

4:15:53 That's one request unit.

4:15:55 What we'll do is we'll take a token out of the bucket.

4:15:58 So I'll subtract one request unit and then

4:16:01 on the resource governor in order to go and say hey,

4:16:03 is the system overloaded, is it not overloaded?

4:16:05 We want to give predictable latency so let's

4:16:08 try to avoid anything like overly saturated cpu.

4:16:13 What we'll do is we'll govern that based off

4:16:14 of the relative CPU memory and IO footprint of that request.

4:16:18 So each second what we're doing is that we're

4:16:20 then replenishing those tokens back in on a timer interval.

4:16:25 As long as I have greater than 0 RUs, we're able to keep serving requests.

4:16:31 Now one innovation that we had that's a bit new is

4:16:36 the RU pooling functionality that we did with the Cosmos DB Fleets.

4:16:42 So what happens is we have one

4:16:44 of these token bucket resource governors on each individual replica.

4:16:48 So on that, that partition, so I have a partition,

4:16:52 the RUs right now first is a shared nothing counter.

4:16:55 We have this on each individual replica.

4:17:00 This is now a partition and then we go and divide that evenly across partitions.

4:17:06 What happens if I have a hotter partition,

4:17:08 and then I have let's say a colder partition.

4:17:11 Can we go and donate budget for from a quarter

4:17:15 partition over to a hot partition in a responsible manner.

4:17:19 So that's what the RU pooling is all about.

4:17:21 We're able to pool across different accounts up to we

4:17:24 have also a node level governor but on the local

4:17:28 governor what's happening is historically if we just use

4:17:32 a shared nothing counter we have a few different trade offs.

4:17:35 Right?

4:17:35 The benefit of having a shared nothing counter is let's say

4:17:38 if one node goes down it's not a global rate limiter.

4:17:42 I can go and continue to serve requests and the system is highly resilient.

4:17:46 But the trade off here is well what

4:17:49 happens if I'm overly spread thin across many replicas,

4:17:53 many partitions is all of these now also

4:17:57 report into we call a global rate limiter,

4:18:00 or we internally like to call it a controller service.

4:18:03 And so think of it as a heartbeat where it's not

4:18:05 on the critical path where each of the local replicas is

4:18:09 going and saying hey here's my utilization on how many RUs

4:18:12 I am consuming and here's the leftover budget that I have.

4:18:16 All of that leftover budget is then sent over to the global

4:18:19 rate limiter where the global rate limiter on the response

4:18:22 can also go heartbeat back saying hey you Just use

4:18:25 all of these RUs and you have some leftover budget.

4:18:28 Meanwhile this other partition is running hot.

4:18:31 But because we have leftover budget somewhere else.

4:18:34 Let's say if this is the one that's

4:18:39 running blazing hot and this one's got leftover budget,

4:18:42 what we're able to do is the global rate limiter can say,

4:18:46 hey, let's relax some limits.

4:18:48 I can give and donate some RUs back over

4:18:50 to the other partition and help you pull that.

4:18:53 So these are some of the really cool

4:18:54 new capabilities where we can give you leftover budget,

4:18:57 give you higher burst ability and do so

4:18:59 in a responsible manner that still preserves a shared nothing architecture.

4:19:03 On the front door.

4:19:05 Now, this is a lot to take in.

4:19:07 Let me go and zoom out.

4:19:11 We talked a little bit on partitioning,

4:19:12 on replication and on request units on what

4:19:15 happens behind the scenes in Cosmos DB.

4:19:17 Let us know what you think about this format if this works out for you.

4:19:20 There's a lot of other topics we can do in the future like indexing,

4:19:23 there's query engine, et cetera.

4:19:24 Let us know in the comments on what you

4:19:26 think and please do fill out our form evaluations.

4:19:29 Otherwise, thank you so much.

4:19:32 Thank you so much.

4:19:33 Thank you, Andrew.

4:19:34 You know what's crazy is that that was

4:19:37 such a great behind the scenes look and you

4:19:40 gave me something like this, I think maybe the first day I started on the team.

4:19:45 So it's always great to kind of see this again.

4:19:47 It's a refresher and keeps me really up to speed on what we're doing.

4:19:52 Yeah, I mean I figure we keep doing this as an onboarding thing internally,

4:19:56 but honestly it helps our developers so much

4:19:58 when they know what's happening behind the scenes too.

4:20:00 On rationalizing, hey, why is the system behaving the way it is?

4:20:04 And that's the intent of also disclosing this out broadly here.

4:20:07 Fantastic.

4:20:08 Well thanks Andrew.

4:20:10 We appreciate it.

4:20:11 And we now know a little bit more about how Azure Cosmos DB operates, at scale.

4:20:17 And you know what, we had so much information shared with us and I

4:20:21 want to go ahead and share some things that we've had from our community.

4:20:25 Nicholas Chang, he's a really,

4:20:27 really wonderful member of the global AI community,

4:20:30 working to help get everybody.

4:20:32 You might have seen him on one of these AI tours that we've got going right now,

4:20:36 but he actually said that if Damian Brady,

4:20:38 who's a buddy of mine who's over at GitHub in their DevRel.

4:20:43 He said it was good to go.

4:20:44 I think you should go.

4:20:45 Patty.

4:20:46 What else do we have as far as programming language is a concern?

4:20:48 Yeah.

4:20:49 We also released a poll about what programming language used the most

4:20:52 and it was great to hear that 45% of udevelopers use.

4:20:58 Net.

4:20:59 I personally am a Python developer,

4:21:01 so I'm glad that I'm joined by 32% of the audience.

4:21:04 And we're followed by Java, JavaScript, TypeScript and Java.

4:21:08 So really exciting stuff.

4:21:09 Thank you for filling out our poll.

4:21:11 Absolutely.

4:21:11 We really, really appreciate it.

4:21:13 Yeah.

4:21:13 But now we're moving into one

4:21:15 of the most critical topics which is data modeling.

4:21:18 Yeah.

4:21:18 The decisions that you make early can define your application success

4:21:21 and it's a great reason to keep learning and sharpening your fundamentals.

4:21:25 I know one of the things that I always hear from Andrew,

4:21:28 who we just heard from is making these decisions early

4:21:32 are going to help you save yourself from headaches later.

4:21:36 No, 100% preparation really is key and if

4:21:39 you want to lock in those fundamentals.

4:21:41 The Cloud Skills Challenge, again it's a great follow on.

4:21:44 So data modeling is one of those topics and again it runs through May 8th.

4:21:49 Yeah.

4:21:49 You complete the modules and you can enter

4:21:51 to win one of 500 free DP-420 exam certification vouchers.

4:21:56 So start here at aka Ms.

4:21:59 Cosmos DBConf challenge.

4:22:02 So next we have an awesome talk

4:22:04 from one of my favorite community champions, Hassan Sovereign.

4:22:07 Yeah, and we're going to be talking about data modeling decisions,

4:22:10 Jay and that they will make or break your Cosmos DB application.

4:22:14 Let's go.

4:22:18 Hi everyone.

4:22:19 Hi Azure Cosmos DB Community, I am Hassan Savran.

4:22:23 I truly enjoy meeting many of you at test conference and I cannot

4:22:27 wait to make new connections and talk about Cosmos DB at upcoming events.

4:22:33 Today I'm going to talk about the data modeling with you and I'm

4:22:37 going to try to make it interactive as much as I can.

4:22:40 So I have no slides for you.

4:22:42 Today we are going to work on the VS code, as much as we can.

4:22:46 And you know, we are going to actually look at the Stack Overflow data model

4:22:50 and try to move it to Cosmos DB and make it scalable and very affordable.

4:22:56 Currently it's in the relational database.

4:22:59 So we are going to take it from there

4:23:01 and move it to Cosmos DB by using different VSCode, extensions and tools.

4:23:06 So I just want to kind of show you how

4:23:08 I guess I handle when I get a project like this.

4:23:13 So I guess before, we Move to the VS code.

4:23:17 Let's look at what we are dealing with.

4:23:18 Right.

4:23:19 So we have limited time so we cannot

4:23:21 kind of look at all the tech overflow tables.

4:23:24 So we are going to actually focus on the busiest

4:23:27 part of the application which is we display the questions and we are going to go

4:23:32 and click on a question and get all the answers.

4:23:36 Right.

4:23:37 Sounds simple but as you can see there's all things that are going on here.

4:23:41 So we have like questions, we have answers,

4:23:44 we have comments, we have badges, we have votes.

4:23:48 So we have potentially really a large application page to display here.

4:23:54 So this is what we are going to tackle today.

4:23:57 So let me go back to my VS code and let's

4:24:00 start from the relational database where the data is actually currently at.

4:24:05 So these are the tables we are going to tackle today.

4:24:08 So we have the post table here.

4:24:11 As you can see there's a bunch of kind of like a columns here.

4:24:16 We have each post can have multiple comments and multiple votes.

4:24:21 We have the user's information and users can have number of batches.

4:24:25 So in a relational database you can easily join this table

4:24:28 together and really deliver whatever the query is asking for.

4:24:33 But we are going to go to the Cosmos DB, which we don't have any more join.

4:24:37 So what are we going to do?

4:24:39 Well before we jump what we are going to do, let's talk about what we should,

4:24:43 we shouldn't just take this as it is and move it to Cosmos.

4:24:49 This is not going to work because we don't have joins in the database side.

4:24:53 So you have to kind of go and join each of these, you know,

4:24:57 data in your front end or in the middle layer which is not going to be easy.

4:25:02 So let me actually show you if we will actually

4:25:06 take this as it is and put in Cosmos db, what's going to happen?

4:25:10 So I have a bunch of the different databases here.

4:25:13 So what the application we are seeing here, this is the one that I created.

4:25:17 It's available in VSCOPE extension, it's just for the Azure Cosmos DB community.

4:25:22 It's free, you can download it.

4:25:24 So what I'm doing here is I'm just picking my database.

4:25:28 As you can see all these tables that I just show you, they are here available.

4:25:32 So we are going to go and start with the probably the posts table.

4:25:37 So first we have to go get the questions

4:25:38 and answers and the post table has both of them.

4:25:41 So I'm just going to go and actually go pick this one before I click to execute.

4:25:48 Actually there's A good function here.

4:25:50 I can.

4:25:51 So I'm just going to run my query analyzer here so

4:25:54 I can track how much everything is costing to me later.

4:25:58 So the first we are going to go get the questions and answers.

4:26:01 As you can see it looks like we have three.

4:26:03 I'm guessing I have one question and two answers here.

4:26:07 As you can see it cost me three requests.

4:26:10 Okay, not bad.

4:26:11 But we are not done yet because we have many other

4:26:14 stuff that we actually have to go and get like comments.

4:26:17 So to get the comments we kind of need to use

4:26:19 the in here because we are dealing with three objects.

4:26:22 So we have to pass their post ID and get this you know, information.

4:26:27 So let's go and get that too.

4:26:30 That cost me 337.

4:26:31 Okay, not bad.

4:26:33 But now we are going to continue.

4:26:35 You might need to get the votes or maybe try

4:26:38 to count the votes and display them on the page.

4:26:41 So we need this information too.

4:26:43 Another three then as you saw on the, you know,

4:26:48 the page there are some user information.

4:26:50 So I need to go and get the users,

4:26:52 for these questions and answers I'm going to go and get

4:26:56 those another three that I have to get, you know, paid for.

4:27:00 Then I'm going to go and get the badges.

4:27:02 I may need to maybe count the badges of each of these users,

4:27:05 display emojis or you know, whatever depending on what they have.

4:27:09 Plan to get that information.

4:27:12 Well, what's happening here?

4:27:13 Looks like I have 277 badges and it actually cost me 16 request units.

4:27:19 If I'm going to go back to my query analyzer, this is really what's happening.

4:27:23 Just to display one page.

4:27:25 I need to sum it up.

4:27:27 All these request units.

4:27:29 I couldn't join any of this data in database level.

4:27:32 So that's going to be my problem to solve in front end or in the middle.

4:27:36 So that's why you don't want to get a relational

4:27:40 data model and move to Cosmos DB as a base.

4:27:43 It's not going to be you know, good for feature for today.

4:27:48 As we can see it's kind of as the data is going to get larger,

4:27:51 if the traffic is going to get larger,

4:27:52 it's going to get, you know, cost you more and more.

4:27:56 So what do we do?

4:27:57 Well, what do we do is it's pretty

4:27:59 clear we need to reduce the number of containers.

4:28:04 So for that let's actually look at this diagram here.

4:28:07 That's where we are right now.

4:28:09 We want to try to write and read data, but currently we have five containers.

4:28:16 We only look at the read, we didn't even look at the right.

4:28:19 So if you actually.

4:28:19 I want to look at the right, I can quickly do that here.

4:28:23 So let's actually try an update.

4:28:25 So I'm going to just update a post.

4:28:27 I'm just increasing those numbers here by using the patch operation.

4:28:31 Let's see how much this is going to cost me.

4:28:33 15 requests and that's only one container.

4:28:37 So you know, things can get really expensive really quick with this.

4:28:42 So as I said before, we want to, you know,

4:28:45 reduce these numbers, the number of the containers here.

4:28:49 So what are we going to do here is.

4:28:52 Well, we have the post ID is the partition

4:28:55 key in each of these, contexts as you can see.

4:29:01 Now the good, news is we have the post ID for, you know,

4:29:05 getting shared by these three containers and the user

4:29:08 ID is getting shared by this too.

4:29:10 Since they share the same partition key,

4:29:12 why don't we just put them in the same container,

4:29:15 just like the version two suggests?

4:29:18 Right.

4:29:19 So we have a post.

4:29:19 Our partition key is post ID and each comments

4:29:22 and was is going to be under the same post document.

4:29:27 Sounds good on paper.

4:29:28 We can try it.

4:29:30 Also that applies to the users.

4:29:32 So rather than we have five containers, we end up with two containers.

4:29:36 So let's try that.

4:29:37 Let's see how this is going to work.

4:29:40 So I'm going to go and change my data model to two here.

4:29:44 Go back to results.

4:29:46 And as you can see, I have only two containers now.

4:29:49 So let's start with the post.

4:29:51 Well, for that I'm gonna, run or execute the same exact, query that I just ran.

4:30:00 So let's see what's gonna happen here.

4:30:01 Execute this one.

4:30:04 It cost me three request units, but this time, I got all the information.

4:30:09 I have the questions, I have the answers,

4:30:11 I have the comments, and I have the votes.

4:30:13 And if you look at the document size 5.71, it's really not that bad.

4:30:18 The good part is if I actually go scroll down,

4:30:21 you're gonna see everything is already mapped.

4:30:23 So I know that these words are for this question,

4:30:26 these comments are for this question.

4:30:28 So everything is really kind of ready to render, in the front end.

4:30:32 So I kind of like that.

4:30:34 The next one, this is, I think the answer.

4:30:37 There's no what's and comments.

4:30:38 Great.

4:30:40 The third one, let's see what's going on here.

4:30:42 Oh, that one kind of gives me some red alerts here.

4:30:46 Why?

4:30:47 Well, look at the number of comments we have.

4:30:49 21.

4:30:51 So which tells me this, the Relationship between the post

4:30:54 and comment is not one to few, it's one to many.

4:30:57 Which means if comments are going to grow and we don't have a limit for it

4:31:01 then that means that we are going to have an issue in feature for sure.

4:31:06 This is great for today but you know

4:31:08 you really need to worry about the feature too.

4:31:10 When you start to have more data, more traffic comes in your database.

4:31:15 So I know this cost great in all three but I don't think that this idea is

4:31:21 going to actually work for this data model

4:31:25 because in feature you know we will have issues.

4:31:28 Well let's look at the users table too I guess that was the only post.

4:31:31 So this is my user table.

4:31:33 Let's execute it.

4:31:35 Let's see what's happening here again.

4:31:37 299 I have looks like one user here but look at this badges.

4:31:42 I have 54 badges which kind of tells me

4:31:46 that maybe users can get the same batch multiple times.

4:31:50 Then this is not going to be easy to kind

4:31:52 of keep them under the one object like that.

4:31:55 So this is a one to many relationship too and really just doesn't

4:32:00 make sense to me especially for feature to put them in the same place.

4:32:06 The mainly because it's going to actually increase our data

4:32:09 model and there's going to be no limit for it.

4:32:12 So I'm going to actually go back to my diagram

4:32:15 and try to maybe talk about the sizes.

4:32:19 The size is important.

4:32:21 As you can see the data model size, you have a limit for it.

4:32:24 Two megabyte is the limit.

4:32:25 It cannot be larger than that.

4:32:27 And your data model, as, as long as your data model gets getting large and large

4:32:32 that means that your cost of the query is getting you know,

4:32:36 expensive and expensive.

4:32:37 Same with the right because well Cosmos DB needs to have you know,

4:32:41 handle many more properties and it's going to cost you more.

4:32:46 So whenever you create a data model you want to be you know,

4:32:49 like give some kind of estimate how big this data model is going to be.

4:32:53 So if your data model is going to continue growing

4:32:55 and growing probably you're going to hit that two megabyte

4:32:58 limit and your solution is going to get more expensive

4:33:01 and more expensive every year as data grows and traffic grows.

4:33:05 So because of that, yeah it is a good

4:33:09 kind of it give me a good result right now.

4:33:12 But for feature, I don't think that I can

4:33:14 end all this under the two documents or two containers.

4:33:20 But what do we do?

4:33:22 Well we know one thing we are going to roll back,

4:33:25 but we really cannot roll back to that because it's not really cheap either.

4:33:30 The good news is, well we are still,

4:33:32 you know, I guess sharing the same partition keys here.

4:33:37 And I'm not sure if you notice one thing here,

4:33:39 there is actually a nightmare here waiting to happen too.

4:33:43 So if you look at our select statement,

4:33:45 it says Post ID is this number or Parent ID is this number.

4:33:50 So the first one is giving you the question,

4:33:52 the second one is giving you the answers because

4:33:54 both of them are actually in the same post.

4:33:57 Table.

4:33:58 The problem is the post ID is the partition key, Parent ID is not.

4:34:03 So as soon as, right now we have all data in one physical partition.

4:34:08 As soon as you start to distribute this data,

4:34:11 this query is going to become a cross partition

4:34:14 query and it's going to get ready expensive really quickly.

4:34:18 So we need to do something about that too.

4:34:21 So what do we do from here?

4:34:23 Well we can try the document type data modeling.

4:34:29 Document type data modeling is working a bit different.

4:34:35 So we still have two containers here, Post documents and user documents.

4:34:42 But rather than saving only one entity type in the, you know, one container,

4:34:46 we want to save different types and we want

4:34:49 to control that with a new property named document type.

4:34:53 So I can easily figure out which post, this question,

4:34:56 which post is answered by looking the post type id.

4:34:59 So I just, you know,

4:35:01 change that and add question and answer document types in this post.

4:35:05 So I can easily, for example when you click the question,

4:35:09 tell me your question id.

4:35:10 And by knowing the question id I should be able to get all the questions,

4:35:14 all the answers, all the comments and all whatever really from this container.

4:35:19 So this might actually work very well for the situation.

4:35:23 Same with the users.

4:35:24 You know, you can just get the user information

4:35:26 or you can get the badge or you can get all.

4:35:29 And the other part that I really like about this one, this is open so you,

4:35:32 if you have different kind of, you know, document types later you can continue

4:35:36 adding here rather than creating new containers.

4:35:40 So let's see how this is going to work.

4:35:42 I'm going to go back to my Cosmos DB NoSQL studio,

4:35:45 change my data model to data model 4

4:35:49 and let's start with the post document first.

4:35:54 Now as you can see I don't have the post ID anymore.

4:35:58 My partition key is question id and if I want

4:36:01 to get only the question for this, for example question,

4:36:04 I can run this and just execute this one

4:36:10 here and this will give me only the question.

4:36:15 2.94 not bad.

4:36:17 I can get the only the answer here too.

4:36:20 I'll execute this one.

4:36:22 296 as you can see I have only the answers

4:36:25 because there are two answers here or I can get everything.

4:36:30 I don't care.

4:36:30 It's a question answer comment at the end.

4:36:32 I need to render this in some way in my front end.

4:36:36 So I can just do that and look at it.

4:36:39 This is beautiful.

4:36:40 You know what is beautiful?

4:36:41 This is beautiful here.

4:36:43 The cost is 3.61 and I was able

4:36:46 to get everything 29 items with that much request.

4:36:51 So this is really very scalable and very affordable solution.

4:36:56 I am really happy about this one.

4:36:58 So as you can see I have some answers here, I have some comments here.

4:37:02 And you can easily look at the document type and you

4:37:05 know render that in the correct places in the front end.

4:37:08 Or if you want to kind of use a factory

4:37:10 design and kind of put them in your, you know,

4:37:13 the C, sharp object or anything like that, you can do both.

4:37:16 So this is, this is working great.

4:37:17 I'm really happy about this one.

4:37:21 Now next is usually developers will kind of stop here.

4:37:28 Why?

4:37:28 Because probably that was the requirement.

4:37:31 Create a scalable affordable data model.

4:37:35 Which we did.

4:37:36 Right.

4:37:37 Well I would suggest you not to stop there because there's

4:37:41 all kind of other things that you can still work on.

4:37:44 And those are what about the feature?

4:37:47 What is going to happen in feature?

4:37:49 How can I make this data model better so it can

4:37:52 tackle the problems is going to come in feature for the system.

4:37:57 So for that there are a couple of options you can actually look at.

4:38:01 The first thing I can look at is probably the tenants, right?

4:38:05 So yeah, we have on the stack overflow of questions

4:38:08 today but as you know stack overflow got much bigger.

4:38:12 There's a stack exchange.

4:38:13 You can actually take this, you know,

4:38:15 questions and answers and apply to different domains and they

4:38:18 can use talk about the totally different questions and answers.

4:38:22 So therefore it will be a different ui but in the back end it will be the same.

4:38:27 So we can really easily make that ready if

4:38:30 we actually add the tenant ID to our data model.

4:38:36 But we can do that.

4:38:37 And if you're going to do that then that means

4:38:39 we can actually use the hierarchy of partition key.

4:38:41 Currently we are using only regular partition key which is the question id.

4:38:45 But if I'm going to do something like that, what is that going to look like?

4:38:48 Well all I need to do is I need to go and create

4:38:52 a tenant ID in here and whenever I'm asking the question here.

4:38:56 You know I'm going to Add the tenant ID1 and question ID is this.

4:38:59 So really that's what my select statement is going to change into.

4:39:03 And when I run that issue, it should run just fine.

4:39:07 Let me go data model 5 first.

4:39:12 As you can see I have the data so it's still three or five.

4:39:16 And if you actually go to Cosmos DB info tab here and look at the oral.

4:39:21 So look at the partitioning here.

4:39:23 My partition key is the tenant ID and question id.

4:39:26 So whenever I actually I ask this question id let's say we have two tenants.

4:39:30 One is about the C sharp.

4:39:32 Other one is maybe the farmers who's you

4:39:34 know asking a question how they're your chickens are, you know, can do better.

4:39:39 So in this case if you're going to tenant

4:39:41 ID 1 we are only search the C sharp question.

4:39:44 We are not going to search for you know, farmer's question in that case.

4:39:48 So you can easily you know change where they actually

4:39:53 are looking by you know changing the tenant ID here.

4:39:58 So I like this one.

4:40:00 This will be actually make the system ready or multi tenant system.

4:40:05 What else we can do?

4:40:07 Next one is find questions.

4:40:09 Now you're going to say well this is great

4:40:11 but how are we going to find these questions?

4:40:13 Right?

4:40:14 So the question is already in the page, you click on it.

4:40:17 But users needs to find these questions in some way.

4:40:20 So there are two things I will suggest for that.

4:40:22 In older days we used to have like categories,

4:40:25 we used to have tags user supposed to go and click the category,

4:40:29 find the products, what they need to find.

4:40:32 But well the things start change very quickly, right?

4:40:35 So right now everybody likes the chat, everybody likes to write,

4:40:38 you know what they need and they are supposed

4:40:40 to go and find whatever they are looking for.

4:40:43 So for that case what we can do here is

4:40:46 Cosmos DB has two options that you can easily set up.

4:40:49 The first one is full text search.

4:40:51 In the full text search all you have to do here

4:40:54 is you have to go and look at your full text policies.

4:40:58 You just need to define which language the data is in.

4:41:02 In post body that's where my data is for that description,

4:41:06 about the questions or the answers.

4:41:08 And I just create a full text index 1 with the post body.

4:41:12 After that, after re indexing is done,

4:41:15 all I have to do here is I can use the full text contains.

4:41:19 That's one of the functions for full text.

4:41:22 But this is one of the easiest ones.

4:41:24 So I'm looking for the questions.

4:41:27 I'M looking the post body and I'm looking

4:41:29 for if it has any SQL Server in the text.

4:41:32 So as you can see the post body is a pretty large text.

4:41:36 So if I'm going to go run this one

4:41:38 and if the user goes and writes SQL Server and search,

4:41:42 this is what's going to happen.

4:41:43 This is going to find top 10.

4:41:45 And as you can see the SQL server is in the post body out there.

4:41:49 I'm just giving the title here.

4:41:50 And the best part is look at the cost.

4:41:52 It really didn't cost us that much to make a search like that.

4:41:58 Next one is the vector search.

4:42:00 So the vector search in that one

4:42:04 user is really not looking specific words anymore.

4:42:06 It's looking maybe like looking

4:42:08 for a relational database operation problems for example.

4:42:12 So you have to Go and find MySQL, PostgreSQL and all kind of stuff.

4:42:15 So if that is the case,

4:42:17 the first thing you need to do for that is you just need to create no vectors.

4:42:21 And as you can see I created policies here.

4:42:24 You just need to give some information about the model that you pick.

4:42:28 And then this is my indexes,

4:42:32 interestingly there's other actual options here which I don't show it,

4:42:35 but you can actually see in the Azure portal

4:42:37 that is the shard ID for the vector index.

4:42:40 So in that case we have tenant one and two, right.

4:42:43 So you can easily create those indexes

4:42:46 and the indexes will distribute it by the tenant id.

4:42:49 You can actually write that.

4:42:50 So whenever you make the vector search,

4:42:52 it can search only that tenant ID vectors.

4:42:55 So it will make it much cheaper that way.

4:42:58 And DSCan is actually supporting that.

4:43:00 So that's why I kind of like the D scan indexes.

4:43:03 Anyway, before we go, you know more details, I go more complex.

4:43:06 Let's see.

4:43:08 Also this application is going to make the vector search much easier

4:43:12 because before you make the vector search you need to generate the vector,

4:43:16 sorry vector itself.

4:43:17 So as you can see, if you have a vector policy,

4:43:19 I have this vector search option here.

4:43:22 You just need to go and this is going

4:43:24 to go and discover all the AI policies you have.

4:43:28 And I have this embedding model for example.

4:43:30 And this is the data I want

4:43:32 to vectorize or search for relational database performance issues.

4:43:36 You click vectorize it, it calls and creates the embedding for you.

4:43:41 You can actually easily see it here.

4:43:43 So this is the text and this is the.

4:43:45 Well all that numbers is the vector I'm searching for.

4:43:49 So I can easily use this embedding parameter Before I

4:43:53 click on running that, I hope you notice one thing here.

4:43:56 Look at the document type.

4:43:58 It is post vector.

4:43:59 So I am not searching the whole database.

4:44:03 I'm only searching for document type.

4:44:04 So that actually makes my search much lighter.

4:44:07 And I'm passing the tenant ID too.

4:44:09 So I won't kind of go and search for farmer stuff I guess in this case.

4:44:14 So click on execute and this is going to cost a little

4:44:18 bit more because this is much more complex than looking for the word.

4:44:21 As you can see it find MySQL it find

4:44:25 WSQLDesigner insert problems dynamic SQL store procedures in Oracle.

4:44:30 It's much more complex than the other one that I just showed.

4:44:35 But it's available and you can easily use that too.

4:44:40 All right, the other three things that I

4:44:42 can just talk quickly is schema version.

4:44:46 You should have a property for the schema versioning out there

4:44:49 and every time the schema changes you can just increase that.

4:44:52 Then you can use that property to, if the business workflow changes

4:44:57 from schema version to other schema

4:44:59 version you can actually handle that retention.

4:45:02 You should really look into that because

4:45:04 Cosmos Signature have only the hot data.

4:45:07 If you don't care about that data anymore,

4:45:10 you can just take it out and put it somewhere much, you know, cheaper.

4:45:14 So that's going to make your data size smaller and that's going

4:45:16 to make your storage size much smaller and things going to get go,

4:45:20 you know, work much better for that.

4:45:22 You can use a TTL function of Cosmos DB for example.

4:45:26 If you don't care about the data after two years,

4:45:29 well you can easily make the TTL two years and Cosmos DB

4:45:33 automatically will delete your data after two years for free for you.

4:45:37 So that will be a great thing to ask to all the requirements,

4:45:42 whoever give you the requirement, what is your retention policies for that?

4:45:45 The last one is the analyzing data.

4:45:47 We create a great operational database,

4:45:50 data model but at the end somebody needs to analyze

4:45:54 this data so you will see how the business is doing.

4:45:58 So for that Cosmos DB has the mirroring data to fabric

4:46:03 and you can easily turn that on and data analyst

4:46:06 or BI developers can easily kind of go and look

4:46:09 at the data and analyze the data and create reports from there.

4:46:13 So those are the things that you should really

4:46:15 kind of consider and think about in the data model.

4:46:19 Well that's all I have for you today.

4:46:21 I hope everybody you know learns something new today.

4:46:24 If you have any questions,

4:46:26 you know how you can find me on the LinkedIn and Twitter or you know,

4:46:30 and other social media social platforms.

4:46:33 And I'll be more than happy to help you from those areas.

4:46:37 Thank you everyone.

4:46:38 Thank you for coming to my session.

4:46:46 At the intersection of science, artistry and design, Pantone gives designers,

4:46:51 brands and manufacturers a shared language to turn creative ideas into reality.

4:46:58 I'm Sky Kelly and I'm the president of Pantone.

4:47:01 Going beyond color standards, Pantone is about the business of creativity

4:47:05 and the trust and precision behind it.

4:47:08 Pantone harnesses decades of research,

4:47:11 being expertise and deep ties with the design community

4:47:14 to understand and predict what's new and what's next in design.

4:47:20 We really bring life to trends that shape culture.

4:47:24 When Pantone and Microsoft teamed up,

4:47:26 we built a system that translates those insights into actionable inspiration.

4:47:31 With the help of Azure AI Foundry, Azure Cosmos DB and GitHub Copilot,

4:47:36 we built AI agents that launch a new way

4:47:39 to engage with our platform and instantly unlock our expertise.

Study with Looplines Download Captions Watch on YouTube