Azure Cosmos DB Conf 2026 | Live Stream
Microsoft Developer
0:03 Hey everybody.
0:04 Welcome to azure Cosmos DB conf.
0:06 Hello.
0:07 Hello.
0:08 Hi, welcome.
0:09 We've got so much great stuff to show you.
0:11 Hi.
0:12 Hi.
0:12 Hello everyone.
0:13 Hey.
0:14 Hello, hello, hello.
0:16 Welcome.
0:16 Welcome to azure Cosmos DB conf 2026.
0:30 All right.
0:30 Hello everybod.
0:32 Welcome to Azure Cosmos DB Conf 2026,
0:34 a virtual event for developers building modern apps with Azure Cosmos DB.
0:39 I'm Patty Chow, product manager for DocumentDB
0:42 and I'm joined by my amazing co host,
0:45 Jay Gordon, Senior Program Manager for Azure Cosmos DB.
0:48 You know what today is all about?
0:50 Practical guidance that you can use right away.
0:53 So we're talking about, architecture patterns,
0:56 best practices and real stories from teams building
1:00 production systems with Azure Cosmos DB, DocumentDB and AI.
1:05 Yeah, that's right.
1:06 You'll see a mix of deep technical sessions,
1:08 quick interviews and demos and this will
1:11 cover everything from getting started to performance,
1:14 scaling and building AI powered experiences on Cosmos DB and DocumentDB.
1:20 And you know what, whether you're joining us live,
1:23 and I, wanna say thank you to our live audience.
1:25 I'm gonna give you all thank you.
1:29 Thank you.
1:29 Whether you're joining us live or watch Demand later, I know I will be.
1:34 We're excited to have you as part of the Azure Cosmos DB community.
1:38 And without our community, we can't have an event like this.
1:43 But before we get started, let's get the big hand emoji.
1:47 We want to say a big thank you to amd.
1:49 Yeah.
1:50 AMD is sponsoring Azure Cosmos DB Conf.
1:53 This year and their partnership has helped us significantly expand the show.
1:57 You know what that means, Patty?
1:58 What does that mean?
1:59 Well, you know what that means?
2:01 More sessions, deeper technical content,
2:03 which I know is important to all of you.
2:06 And, we'll have a broader set of speakers
2:08 and topics that we'll be covering today.
2:10 Yeah.
2:11 And this whole expanded programming really
2:13 wouldn't be possible without their support.
2:15 So thank you, amd.
2:17 Yes, thank you so much.
2:18 We really appreciate it.
2:19 Your support has just really made this into something huge.
2:24 I'm loving just seeing the LED wall.
2:26 I love seeing all these people here.
2:28 Thank you so much, amd.
2:29 Yeah.
2:34 Now what you're seeing live today is just part of that experience,
2:38 you know, that's right.
2:39 You know what, Patty?
2:40 There is a full library of sessions available live today and on demand.
2:44 You'll see them all go live around 1:30pm, Pacific.
2:49 And they're going to be covering everything
2:51 from AI architectures to design patterns, performance tuning.
2:55 We all need to do that.
2:56 And real world cases.
2:58 Yes.
2:59 So let's say that we're presenting here and if we can't catch everything today,
3:03 or you can't catch everything today,
3:05 you can go deeper and explore it all at your own place, at your own pace.
3:09 And you know what else we're doing, Patty?
3:11 And this is a big one.
3:12 We're also running the Azure Cosmos DB Conf.
3:15 Cloud Skills Challenge.
3:17 This is something.
3:18 You know what, Patty, I know you're excited about.
3:20 Yeah, I'm really excited about this, Jay,
3:22 because it is designed to help you build with hands
3:25 on skills using Azure Cosmos DB and it runs through May 8th.
3:29 Sounds put in your calendar.
3:30 Yeah, yeah, absolutely.
3:32 And I just hope that everybody who is watching you take advantage
3:36 of it because the first 500 people to complete the modules will get.
3:40 And here we go.
3:41 Patty, this is my favorite part.
3:42 What's your favorite part?
3:43 It is a, 100% free voucher to take the DP-420 exam.
3:50 So you know what?
3:51 Get certified.
3:51 Get certified for free.
3:54 Get learning.
3:55 Yeah, you know, I love free too.
3:57 And if you want to get a head start, visit AK Mscosmos CB CompChallenge.
4:03 That's right.
4:03 And like I said, our community,
4:05 the huge part of what we do to make sure that you're all involved,
4:09 you're all included and you know, you can take part in this event.
4:13 Yeah.
4:13 I mean, yeah, truly.
4:15 I think one of the things that makes Cosmos DB and Cosmos DB Conf.
4:19 So special is the people behind it.
4:22 And that's what makes it so special.
4:23 Like you all in our audience, thank you for being part.
4:29 And so we want to thank our incredible speakers,
4:32 our customers and partners who are sharing real world experiences today.
4:37 And that includes teams from OpenAI, Vercel, Walmart, AMD,
4:42 of course, the ODP Group, Southworks, Next, and many, many more.
4:48 And you may have gone on one of these sites today.
4:50 You might have bought something.
4:52 They're part of everyday life.
4:54 And these teams that we'll be,
4:56 sharing today that are building production systems at scale.
4:59 Yeah.
5:00 Today they will be sharing what's been working for them.
5:04 Sure.
5:04 And if you want to explore everything today that's happening,
5:07 you can head over to aka Ms.
5:11 Azurecosmos dbconf.
5:13 Yeah.
5:13 Just a quick note.
5:15 The YouTube description has all of today's links in one place.
5:18 There are over 30 sessions total,
5:21 with every session going live before the end of today's show.
5:24 Plus, if you're in the YouTube chat,
5:25 please tell us where you're watching from and ask questions as we go.
5:29 Sorry, I was working at my own pace.
5:30 No, it's okay.
5:31 We want to hear from you.
5:32 We want to hear from you too.
5:34 Well, you know what?
5:34 I think we've set you all up and it's now to get you to our, opening keynote.
5:39 Yeah, that's right.
5:40 Here's Kirill Gavrylyuk, Vice President of Azure Cosmos DB at Microsoft.
5:45 Developers.
5:45 Developers.
5:46 Developers.
5:48 Welcome to Azure Cosmos DB Conference 2026.
5:52 We are really grateful to have you here.
5:54 And let's take a look back at the year since
5:57 the last conference and what a world year has it been.
6:02 Of course we're all building AI apps at amazing pace, right?
6:07 And for Cosmos DB perspective, we invested heavily in semantic search.
6:11 That is core part of any AI app you have now.
6:14 Full text search, hybrid search, vector search,
6:17 semantic data, ranker, plethora of capabilities.
6:20 With confidence we can say if you're using Cosmos DB,
6:22 you do not need a separate search system,
6:24 just use built in capabilities and take advantage of cost efficiency,
6:29 performance, scale reliability and precision.
6:33 And of course more than half of our customers
6:36 are using coding agents today to build apps on Cosmos.
6:40 And for that we offered great agent kits
6:44 with skills to make your coding agent world class expert.
6:48 We added MCP servers and other capabilities to make this possible.
6:53 Now, AI does not remove the need for reliability,
6:57 security, performance and that is our priority number one.
7:01 This is where we invest the most of our energy.
7:04 We've provided throughput management,
7:06 fleet management for customers with large fleets of Cosmos DB.
7:10 We've invested with simplified partitioning with global security indexes,
7:14 hierarchical partition keys.
7:16 We've invested in per partition auto failover.
7:19 And now Cosmos DB is the only database that offers you 5,
7:22 9 reliability for any consistency mode.
7:25 And of course security is number one priority
7:28 and a number of improvements in security space.
7:32 Last year we've added second NoSQL database to Azure DocumentDB.
7:38 Unlike Cosmos, DocumentDB is fully open source.
7:41 It's based on an open source project at Linux foundation where Microsoft,
7:45 Amazon, Google and 12 other vendors are collaborating.
7:49 DocumentDB has a number of advantages over other document databases
7:54 and notably more than 40% cheaper than any other document database,
7:59 including MongoDB Atlas on AWS.
8:01 So please come over to Azure, use DocumentDB,
8:04 take advantage of performance and low cost.
8:08 Now AI transformation is on top of everyone's mind, right?
8:12 And for databases really means, in my opinion three things.
8:16 One is AI increase the importance of flexible data.
8:22 The data AI processes and emits is all semi structured, right?
8:26 It's prompts, it's memory, context, that's what AI is focused on.
8:30 And Cosmos DB provides the first class support for semi structured data.
8:37 It proudly presents flexible data model.
8:41 Now, AI accelerated the pace at which we develop applications.
8:46 Flexible data is paramount for you to enable the space.
8:50 Developers cannot be locked down by strict schemas,
8:53 they need schema flexibility which is again number one capability of Cosmos DB.
8:58 That's why we built a database is
8:59 to provide you schema less flexibility in data modeling.
9:03 Now AI brought semantic search as a first class operator
9:09 into queries and it's true for any database that matters.
9:13 And of course in Cosmos DB we invested far and beyond of that.
9:17 We added semantic search,
9:19 full text search vectors, hybrids, semantic grid anchor.
9:23 We invested in GraphRAG and we're about to announce
9:26 agentic retrieval which takes a graph rack to the next
9:29 level where the graph knowledge graph can be inferred
9:32 by LLM for you without you having to manually construct it.
9:36 Now the third important thing about AI transformation is
9:39 we all now are accompanied by our coding agents.
9:44 That's our favorite friends now, right?
9:46 Coding agents speed up development and for coding agents
9:50 to be efficient world class experts in databases, they need skills.
9:54 And that's why we invested in Cosmos DB Agent Kit.
9:57 We'll talk about it later.
9:59 That provides skills to your coding agent no matter which coding agent you use.
10:04 Along with MCP servers plugins for any favorite coding agents,
10:09 that you may choose.
10:11 Now let me introduce our guest speaker, Guillermo Rauch,
10:18 founder and CEO of Vercel, a popular AI app development cloud.
10:23 And Guillermo Rauch knows a thing or two about what AI apps need from databases.
10:30 In order, to have with us CEO and founder of Vercel, Guillermo Rauch.
10:37 Guillermo, you have built libraries like Socket.IO and Next.js
10:42 that form the modern web as we know it today.
10:45 On top of it, you've built a highly successful company, Vercel,
10:49 that has seen a really good growth of developers, that provides developer cloud.
10:57 What do you credit Vercel?
10:59 Success.
11:01 I think if I look back on the principles that we
11:04 created Vercel on, it was really about the ease of use.
11:08 At the time it was ease of use for humans.
11:10 Now lately I've been calling it agent ergonomics or ease of use for agents.
11:15 It's about, you know,
11:16 you have an idea and you need to bring it online and you actually
11:20 want it to scale and work really well and be a really high quality product.
11:25 So if, if the, if it was, if it could be summarized as something simple,
11:28 it would be the focus of ease of use,
11:31 but in the service of quality, not just, you know, getting slop as we call it,
11:36 this taste online that crashes and pages you at midnight makes sense.
11:41 As you transition to agents as your customers.
11:45 What, surprised you is what's happening in the industry.
11:49 It's really the scale.
11:53 As an entrepreneur, as a founder,
11:55 you start doing market sizing estimates when you're preparing
11:59 your early decks for your investors and would be investors.
12:04 The common wisdom was that even being generous,
12:08 Maybe there were 30 million people that could deploy applications because
12:12 we targeted the easiest programming language that we could think of, JavaScript,
12:17 TypeScript, trying to make the cloud super accessible.
12:20 But it was still tens of millions of people.
12:23 And Next.js also lowered a barrier to entry to react into this universe.
12:29 But when you think about agents,
12:31 they're empowering everyone on the planet to create software.
12:35 They might not even be realizing that they're creating software.
12:38 They might just ask for a solution to a problem.
12:41 And in the process, the agent writes an application and deploys it on our cloud.
12:47 So a lot of the growth that we're seeing these days is literally,
12:50 by the way, happened yesterday.
12:51 A high school friend of mine came to visit me.
12:53 He'd never written a line of code in his life and he's like, oh, by the way,
12:57 I know what Vercel does now because Claude code deployed
13:01 a Vercel and told me the application was now online.
13:05 So I could have never imagined that we would
13:08 see an increase in, let's call it a hundred times,
13:12 the addressable market of creators of software.
13:16 Maybe even higher.
13:18 I think we'll be soon in the future where, software is like water,
13:23 like it's just a normal thing to create, consume, sell, and it's everywhere.
13:30 That's amazing.
13:32 And as the platform that is, so frequently used by the agents,
13:36 what do you expect of the modern database?
13:40 How can we help you?
13:42 Well, I'm very thankful.
13:43 I always tell the team, because part of growing up as a company is
13:46 telling the stories and scars of the early days.
13:50 We started out using a database that was very managed by me in the early team.
13:56 And it was a horrifying experience.
13:58 And I reached out to you and I'm very thankful,
14:02 Carol, through an introduction from Nat Friedman,
14:04 because I was looking for a solution that would help
14:07 me focus on growing Vercel and not being a dba.
14:11 We had very few people on staff.
14:13 This is literally the garage story and one of the convictions that I had
14:18 that helped me choose Cosmos and be savvy about this because I'll tell you,
14:23 some of the less familiar developers look at Query languages.
14:27 And they might have a little bit of culture shock sometimes,
14:30 but one of the things that convinced me is that compute was becoming serverless.
14:35 The value of Vercel was you deploy it scales to zero,
14:40 it scales to billions of people.
14:42 I literally am supporting applications these days that were vibe
14:45 coded and go viral all over the Internet and have traffic
14:49 that even teams of engineers would have never dreamed of and can
14:53 now be sustained by someone that barely knows how to prompt.
14:57 And the magic there is that it was that bet on serverless.
15:00 Because I also tell you the reality is
15:02 that a lot of this software is quite ephemeral,
15:05 that's being written by the agents.
15:07 Maybe it's useful for a day, maybe it's useful for a presentation.
15:10 We're seeing a lot of sales engineers, including people at Microsoft,
15:13 use v0 to vibe code an app for a sales pitch,
15:18 for a demo for pre and post sales to demonstrate an integration of a product.
15:23 And so if software is ephemeral, serverless,
15:26 scale to zero and flexibility there is extremely useful.
15:29 So I'm very thankful that we made the right bet.
15:32 The other thing that I wanted to be super mindful
15:35 of that serverless compute helps a lot with is fault isolation.
15:41 So I didn't want to get a machine hot that could create noisy neighbor problems.
15:46 Imagine if the misbehavior of one tenant could impact the reliability,
15:51 uptime, availability of other of other customers.
15:54 But what I just described is the everyday database.
15:57 It's kind of nuts that we still.
15:59 There's so many people that kind of live in the shadows,
16:01 I think because they put systems online whose behavior they cannot predict,
16:07 they might become a wave of queries or users or whatnot that just ruins
16:12 the day for everybody and people start citing slow query logs and all that junk.
16:16 And so I wanted a system that gave me an economical
16:22 thinking where the developer writes a query and they understand its cost.
16:26 We can project out its cost.
16:29 A lot of what I think has made Cosmos successful is
16:31 that it kind of makes sense in the token world, right?
16:35 Like when I use a coding agent,
16:38 I know that certain traces of reasoning cost more compute.
16:42 When I make a cross partition query in Cosmos,
16:44 I get back an RU count of like how much I spent on that query.
16:48 And so it allowed me to scale Vercel from a few
16:52 developers that didn't want to have a DBA on staff
16:55 and to actually hundreds of engineers that are supporting
17:00 a growth in usage that would have never imagined with deployments,
17:06 which is one of the key collections that we Host on Cosmos Having grown,
17:10 our weekly Deployment count is 3x'd in a few months.
17:14 And I think there's probably no database system
17:17 in the world that I could have homegrown,
17:20 for example, that could have absorbed that kind
17:23 of growth that agents have gotten us.
17:27 Thank you.
17:27 Well, we're very grateful for you to be our customer.
17:31 We have many developers online right now watching this that are early in career.
17:37 What would be your advice to them?
17:41 Well, one of the things is that the agent is the new computer and this is
17:46 advice that I give myself every day is that I really have to update my priors.
17:51 Every day I give a presentation to the whole company saying look,
17:56 these are the things that I used to believe AI couldn't do.
17:59 And I was wrong.
18:00 I'm coming out and saying it like I
18:03 thought AI couldn't contribute to large code bases.
18:07 I was wrong.
18:08 I thought AI couldn't have, couldn't do good designs.
18:11 I was wrong.
18:12 And so it's very important to realize that the things
18:17 that you've learned you have to hold on to.
18:20 Of course you have to appreciate the skills that you've developed and be proud.
18:24 But also just be very light on your feet is my advice.
18:29 I like to think of the engineer of the future.
18:33 I'm someone that is very full stack in their understanding of the business.
18:38 When I hire an engineer I don't expect them
18:40 to just be typing code or even prompts for that matter.
18:43 I wanted to understand the full problem.
18:47 Again it's a little bit like I mentioned that economy system that Cosmos
18:51 has given us where I was able to shift the burden of understanding,
18:55 you know, what our database performance and cost
18:58 is going to be to the everyday developer.
19:01 They didn't toss the problem out to.
19:03 Oh, some other team will worry about scaling my database, right?
19:06 No, you are participating in that process.
19:08 I think the original ethos of DevOps was very much in that direction.
19:12 I think the difference now is that everyone is super full stack.
19:16 I do design, I generate videos, I generate SVGs,
19:22 I create front end code bases, I create backend code bases, I write clis.
19:27 During the holiday break I vibe coded my own Swift menu bar app.
19:32 And so I think the antidote here or like the way to really
19:38 stay relevant is to be super open minded and super full stack.
19:45 Thank you Guillermo.
19:46 Really grateful for us to have you here.
19:49 Appreciate everything what you're doing, love your posts,
19:53 love everything that vercel companies likewise.
19:55 It's been a great Journey and excited to continue collaborating
19:58 with you all and do more with Vercel in the future.
20:12 Thank you Guillermo.
20:13 Now Guillermo just covered few of the reasons why they chose Azure Cosmos DB.
20:17 But as these are the same reasons why thousands of other
20:20 companies choose Cosmos as a database for their agentic apps.
20:24 Its schema free flexibility enables your team to run fast iterate fast.
20:29 It's search built in into your query engine so that you
20:33 don't have to deploy different systems for search for database et cetera.
20:38 It's reliability, low latency and serverless
20:42 elasticity to enable your software to spike,
20:45 to enable your app to grow with demand without you touching anything.
20:50 Of course it's 5, 9 reliability that makes your sleep sound
20:56 while your software runs and serves millions and billions of people.
21:01 Now when it comes to AI use cases we have many
21:05 but the common ones with Cosmos DB of course include vector database.
21:09 Right.
21:09 That's why we invested in vectors, multimodal vectors,
21:12 the speed and cost effectiveness of vector search, semantic retrieval,
21:18 billing on top of vectors offers you
21:20 ability to do full text search, hybrid search,
21:23 precise increased precision of your retrieval,
21:26 taking advantage of the knowledge graph of your system.
21:31 But a recent trend is new use case that tripled
21:35 in size over the last six months is agentic memories.
21:39 We have more and more systems choosing Cosmos DB as agentic
21:43 memory for their apps and there is a reason for it.
21:46 Memory is semi structured, memory is dynamic,
21:50 memory needs to be reconciled over time and all
21:54 of these capabilities are built into Cosmos so you
21:57 can focus on what you app needs to do
21:59 and less about the infrastructure that it runs on.
22:05 Now let's get a couple words about the search in Cosmos DB.
22:10 The vector search is based on DiskANN algorithm built by Microsoft Research.
22:15 And the two key things about it is one, it's extremely cost efficient compared
22:19 to the traditional separate standalone search systems.
22:23 It's up to 50 times cheaper as you can see on the charts.
22:28 It is also low latency and scales no matter the scale of your app.
22:32 It provides you the same millisecond latency
22:35 for your vector search and full text search queries.
22:39 It scales well from few vectors to billions of vectors in your data set.
22:45 Now when you think about an AI app built on Cosmos,
22:48 of course OpenAI ChatGPT comes to mind.
22:51 That is well known.
22:53 It's a tremendous one of a kind application.
22:56 It executes more than 1.4 trillion transactions daily against Cosmos DB.
23:02 It stores more than 45 petabytes in Cosmos by now and at this scale,
23:07 it's the fastest growing app on the planet.
23:08 It still grows tenfold every year.
23:12 And to dive deeper into what enables OpenAI
23:15 to scale that fast and maintains a trap, it's a pace of innovation.
23:23 Let me invite John Lee, staff engineer from OpenAI,
23:26 to share a few thoughts with us.
23:30 Well, hi John, it's really great to have you here.
23:33 Would you mind introducing yourself?
23:36 Yeah, of course, it's good to be here.
23:38 My name's John, I'm on the online storage team at OpenAI.
23:42 Awesome.
23:43 Well, OpenAI is a unique company in many ways.
23:47 Right.
23:47 But one of the things that everyone knows about it is
23:50 that, and it's really visible is that it moves really fast.
23:54 You guys are doing many things concurrently and very fast.
23:59 What does OpenAI look for from databases
24:02 to enable users fast pace of innovation?
24:06 Yeah, that's a great question.
24:08 I think there are lots of standard database things that we care about, at scale.
24:13 So predictability, being able to scale to petabyte sizes,
24:16 being very flexible with how our use cases can be built,
24:21 and then of course observability is very important.
24:24 But I think from the scale perspective and being able to move really fast,
24:29 the most important thing here is being able
24:34 to scale from zero to millions of QPS,
24:37 being able to scale from zero bytes to petabytes,
24:39 features are launched on the regular, and they go immediately from no usage
24:45 to being used by hundreds of millions of people, on a daily basis kind of thing.
24:49 Right.
24:49 You know, so the way that you might design a normal database,
24:54 and then scale it out as your use case grows,
24:56 you know, that just doesn't work at this kind of velocity.
25:01 And I think we have thousands of developers that are actively building products.
25:05 So with Codex now things are iterating even faster.
25:09 It's really important to make it easy to onboard to databases really fast.
25:13 So we have thousands of tables, that we have to back by Cosmos DB.
25:20 And we have a system in front of Cosmos that helps
25:24 you be able to do sort of schema less design,
25:28 and allow these product developers to onboard very quickly.
25:33 Amazing.
25:35 Could you share any of the interesting patterns
25:37 that you guys have done on top of Cosmos DB?
25:40 Yeah, I think predictability is really key for how we use Cosmos DB.
25:46 We have a very well defined API that we have and expose
25:52 and all of these clients they build on top of this.
25:56 And that allows us to keep the system stable.
26:01 We don't want Every client to be writing their own Cosmos DB queries,
26:04 creating their own Cosmos DB accounts, creating their own tables.
26:08 So we try to build a multi tenant system on top of Cosmos, interestingly,
26:12 sort of at this global scale where we have
26:15 a lot of Cosmos accounts in a lot of regions, region failure is a big problem.
26:20 Right?
26:21 It doesn't happen often, but with dozens of regions it will happen.
26:26 And so we definitely use Cosmos DB multi region write.
26:31 We replicate our accounts, so that we can always shift reads
26:34 and writes to other regions in case of outages.
26:38 But generally also for good performance,
26:40 you need to be able to allocate your data
26:43 in Cosmos DB accounts that are close by.
26:46 One particularly interesting feature that we use for Cosmos db,
26:49 is that we have tiers of Cosmos DB accounts where some
26:53 of these accounts that are being used to store globally accessed metadata.
26:57 So like things like account, information and user settings,
27:01 we actually replicate that to dozens of regions,
27:05 as opposed to just a single region.
27:09 That's fascinating.
27:11 You've done a lot of innovations on top of Cosmos DB.
27:13 What's your favorite?
27:18 I would say that, being able to move data across Cosmos DB.
27:24 So we have a lot of products that will start,
27:30 on a single Cosmos DB account, and in a single container.
27:35 But as they grow or as their needs change,
27:38 iteration is very important at OpenAI.
27:41 As needs develop, maybe we realize
27:44 that they have much higher performance, constraints.
27:48 And so being able to transparently migrate their data from one Cosmos
27:52 DB account to a different Cosmos DB account that is globally replicated,
27:55 that is something that we've spent a lot of time trying to build around.
28:00 And the Cosmos DB team has given us a lot
28:03 of the building blocks that we need to be able to do that.
28:06 That's awesome.
28:07 We should at some point join the publishers.
28:10 A lot of viewers would love to see us as well.
28:14 Well, we have many app developers that are
28:17 early in career or potentially are, still in college.
28:21 What would be your advice to them?
28:24 Yeah, you know, I think this is evident also
28:26 in the way that our system has been built, to accelerate product teams.
28:31 There's a big culture at OpenAI about shipping things,
28:35 and shipping things quickly.
28:37 I think it's important to have like a vision of what you're working towards,
28:40 but don't over build solutions to problems.
28:43 Try and get something working into the hands of your users,
28:46 and then you can evolve and iterate on it.
28:49 Almost Never.
28:51 You will get the first iteration correct, but you'll learn from that experience.
28:57 Thank you so much John.
28:58 And thank you so much for being our customer
29:00 and thank you for joining us at this conference.
29:05 Thanks for the invite me.
29:10 Thank you, John.
29:10 Now what you noticed John mentioned Codex use in OpenAI,
29:15 everyone uses coding agents, right?
29:17 As I said, based on telemetry,
29:19 more than half of our customers are now using coding agents.
29:24 And for you to get started, let's say you don't know about Cosmos DB yet, right?
29:31 The best thing is to install our Agent Kit.
29:34 This is a set of the skills that gets you agent
29:37 up and running and becoming the world class expert in Cosmos DB.
29:40 It has 80 plus skills in growing encompassing best
29:44 practices with Cosmos DB across the entire app lifecycle.
29:48 From data modeling to partitioning to queries,
29:51 indexing, throughput management, high availability and monitoring.
29:56 Now to see it in action, let me invite our fearless product leader,
30:00 Andrew Liu to show Agent Kit in a demo.
30:05 Hi Kirill.
30:08 So to set the stage, I am building an app here.
30:12 It's a travel planner.
30:13 It's a multi agent Cosmos DB application and it's going to help me plan a trip.
30:19 So here I'm planning a five day family trip
30:22 to LA that includes my 30 year old daughter.
30:24 Now what's important about this is that it's not just helping me plan the trip.
30:28 It remembers where I left off.
30:32 This is done through by storing memory in Cosmos DB.
30:35 This includes short term memory, long term memory, and vector.
30:38 Now here's the thing.
30:39 It's not about the application, but rather, I'm new to building AI apps.
30:43 And the thing I constantly question myself is did I build it right?
30:47 Because at scale, getting it not right gets expensive.
30:52 What if I pick the wrong partition key?
30:54 What if I have the wrong data model, the wrong index?
30:57 I mean that's burning RUs in production.
30:59 That's the kind of thing that you don't want to surprise you when you launch.
31:03 And scale.
31:05 So let's jump over into what this looks like in vs, code.
31:10 First, thing I want to show you is
31:11 actually our new VS code extension or relatively new.
31:14 Many of you guys have seen the portal.
31:17 We get it.
31:18 We've also gotten a lot of the feedback on.
31:19 We really want to go and improve the desktop tooling.
31:21 So we've been putting a lot of effort into the VS code extension.
31:25 What you'll find here is you'll get a lot of the same operations but you
31:28 can do it now in line with your code without having to context switch.
31:32 So here I have my Cosmos DB account, and it works just the same way.
31:36 I can go and query the database right behind this application and what
31:41 you'll see is I'm storing a bunch of data in Cosmos DB.
31:44 I have my short term memory in the form of messages.
31:46 I have my long term memory, in the form in this memories container where I
31:50 have things like my declarative memory which is durable facts,
31:55 procedural memory where I have behavioral preferences and episodic memory,
32:00 which is going to be things that are a bit more trip specific.
32:03 Now let's go into the magic on Agent Kit.
32:08 Did I even get this right?
32:10 So installing this is actually very easy.
32:13 Just1 CLI command away.
32:15 You can use npx skills add and go and add the Cosmos DB Agent Kit.
32:20 It's going to go through a few different questions,
32:23 like setting it up for the project scope or doing it globally on your computer.
32:31 But when it sets it up, what it's going to do is it's going to go
32:34 and create this on the project scope a agents container,
32:38 or rather directory and it'll load all of my skills in here.
32:44 Now that I have my agent skills set up,
32:47 let's go and go check out the quick application.
32:54 So I am going to go take a look here and I'm going
32:56 to have it just go review my Cosmos DB data model for performance issues.
33:02 Now this isn't just a generic LLM looking at the code.
33:05 The Agent Skills is a package set of specialized expertise modules.
33:09 Right?
33:09 We're taking a lot of context about Cosmos
33:12 DB and bringing that into modules, just for this.
33:17 Now there's no services, no accounts.
33:19 It's a repo of skills so the agent gets smarter about Cosmos.
33:22 It's teaching the agent how Cosmos DB actually works.
33:25 It'll include data modeling,
33:27 best practices, partitioning, indexing, RU economics,
33:30 and make sure that that actually matches up with my data access patterns.
33:34 These are things that it took us decades
33:36 to learn over building real world Cosmos DB applications.
33:41 Now you can get that directly in the form of agent skills.
33:47 Now this is what's really cool about this.
33:50 It's not just doing some generic pattern
33:52 matching and giving me generic NoSQL advice.
33:55 It's reading my actual code.
33:57 It's taking a look at My code, my infrastructure as code, Bicep templates,
34:01 correlating my partition key, my index policy,
34:05 my document shapes, to the different data access patterns.
34:09 It knows what good looks like.
34:12 Specifically for Cosmos DB,
34:14 this is expertise that used to live in one senior architect's head on my team.
34:20 But it was hard to get a hold of those individuals.
34:23 I had to book them a week out.
34:25 Now it's just in my editor here on demand.
34:30 So let's take a look at the summary here.
34:32 It was able to go and find an order of priority and highest roi.
34:38 Oh shoot, I missed some composite indexes for my order by query
34:42 that can shave off a bunch of RUs for each of my different queries.
34:46 And when I'm running at scale with high QPS, that's really going to add up.
34:50 It's also going to go and prioritize additional things like hey,
34:54 I should go and fix my memory's partition key once again.
34:58 Hey, if I get the wrong partition key,
35:01 as this application scales to many partitions,
35:03 if I have a lot of cross partition queries that's going to go add up.
35:07 So what I'm really happy about this is
35:09 it's catching this before I hit live production.
35:16 So there's two key takeaways I'd like
35:19 to for everyone here in the audience number one, if you're a new developer,
35:24 Cosmos DB doesn't have to feel foreign or scary like
35:29 all the cognitive load of partitioning data modeling right here.
35:33 I can go and interact here with my Cosmos DB Agent Kit and go and get a live
35:39 code review of my system and make sure
35:41 that what I'm doing is well architected by default.
35:45 Also if you're an existing Cosmos DB developer,
35:47 well this also lets you go and scan your code,
35:51 scan your past projects and go look for opportunities for optimizing
35:55 the performance as well as cost characteristics of that application.
35:59 Now back to you Kirill.
36:07 Wow, thank you Andrew.
36:08 I think the last point resonated with me so much.
36:11 This is not just allows you to increase your productivity and build apps faster.
36:16 It's actually world class expert in Cosmos DB as your code reviewer.
36:21 You can plug it in into your CI CD processes in your code,
36:25 review flows and it's right there catching any bugs that you
36:29 may introduce or helping you tune your code to be more efficient.
36:34 Now we have great coding agents skilled on Cosmos DB.
36:38 We have fantastic AI apps,
36:41 but with all of that we need systems to continue to be reliable,
36:45 secure and performant.
36:46 And Cosmos DB is built for that.
36:49 That is the number one priority for Cosmos DB as a database.
36:53 Cosmos DB is a database that can take your applications anytime,
36:56 from gigabytes to petabytes,
36:58 from hundred transactions per day to trillions of transactions per day.
37:02 You don't need to worry about it if you hit on overnight success at some point.
37:08 Cosmos DB is the only database in the cloud
37:10 that offers you guarantees not only five nines down,
37:14 uptime across any consistency modes, including the strong consistency.
37:20 It also provides you money back guarantees for zero data loss.
37:23 Thanks to the consistency SLA's financially backed.
37:29 Finally, it's the only database that gives you full autonomous resiliency,
37:33 zero touch for your system.
37:36 Now this scale and reliability is made
37:38 available to you not only just by software, but also by an awesome hardware.
37:43 And to discuss how we run the fleet of Cosmos DB,
37:47 I would like to invite Steve Berg,
37:50 Corporate Vice President and General Manager of the Service
37:53 CPU Cloud Business Group at AMD on stage.
37:56 Welcome Steve.
37:57 Thanks for having me.
37:59 Glad to be here.
38:00 Great to have you.
38:03 Now, Steve, from your perspective,
38:04 how does AMD and Microsoft partnership impact how customers experience Azure?
38:09 That's a great question.
38:10 Before I answer the partnership piece, let me give you a little bit
38:13 of history about our working together with Microsoft.
38:16 Back in 2017 we launched our first EPYC processor, Naples,
38:20 and Microsoft was actually the first cloud provider
38:23 to offer it to make it available in the cloud.
38:26 That created a really unique bond between us and Microsoft.
38:30 Over the past decade,
38:31 we've worked together to launch five more generations of of EPYC instances
38:35 that are available in about 60 different compute families across the globe,
38:39 with our most recent one being Turing, which is an excellent product.
38:43 So we're constantly working to optimize our hardware
38:46 along with your software to be more performant,
38:49 more efficient and provide new capabilities with each generation.
38:53 We do this with a deep engineer
38:55 to engineer collaboration on making roadmap changes
38:57 that improve our products and help them
39:00 get tuned better for Microsoft Azure needs.
39:03 So some of these examples include things like confidential computing,
39:06 3D V cache and high performance bandwidth that's used on some of your VMs.
39:13 This partnership has led to more Azure first parties adopting our hardware.
39:20 And that is what the partnership provides is Cosmos DB.
39:24 And users on Cosmos DB get all the benefits of decades of work
39:28 that we do together and all the hard work that happens within our collaboration.
39:32 I could agree more.
39:34 We actually use AMD quite heavily in our Cosmos DB fleet and that's what
39:38 enables us to scale with demands like
39:40 OpenAI and other super fast growing applications.
39:43 Yeah, exactly.
39:44 OpenAI and all those just get all the benefits
39:46 of us working together with everything under the hood.
39:49 So it's really, that's what the partnership really brings.
39:53 That's awesome.
39:53 Now, in today's world with global situation, energy efficiency is top of mind,
39:59 how should builders think about balancing
40:01 the performance needs and sustainability and energy efficiency?
40:05 Yeah, we absolutely focus on that.
40:07 Our product teams and our engineers are
40:09 always looking at what's energy efficiency and sustainability.
40:13 And fortunately, builders who choose to run Azure get all of this benefit
40:17 without even really having to think much about it at all.
40:20 Because we do all of the for them.
40:22 We work to deliver generation after generation EPYC CPUs with more performance,
40:27 better efficiency per watt and improved TCO,
40:30 which directly translates into users and builders being able to get
40:35 benefits of cost and maximizing their performance per watt per dollar.
40:40 That's the critical metric.
40:42 So when you consider the global reach
40:43 of azure with over 70 regions and counting,
40:46 it's absolutely imperative that we maximize performance
40:49 for those users in a sustainable way.
40:51 So focusing on that performance per dollar is where we concentrate on.
40:57 And when Cosmos DB team rolls out new improvements to your SaaS offerings,
41:01 builders implicitly get all those benefits
41:04 just by using the modern infrastructure.
41:06 And when end users roll their own infrastructure as a service,
41:11 if they gravitate towards amd, epyc, vms, they also receive those benefits.
41:16 That's amazing.
41:17 That's exactly what we all need to do, right?
41:19 We don't want to think about problems that we can't solve.
41:22 We want you to solve them for us.
41:23 We want to make it easy for you.
41:24 Yes.
41:25 And semiconductors has always fascinated me
41:27 as a really tough industry because you
41:29 have to plan ahead five years ahead for what CPUs customers might need.
41:35 How do you do that?
41:36 Yeah, it's interesting you asked that because I don't think anyone could
41:41 have predicted what was happening with AI in the past year or so.
41:45 And it's really fascinating to see what you're doing with agentic AI today.
41:49 I think the rapid pace of AI took everyone many by surprise and no
41:54 one would have guessed the importance that CPU provides the role in this market.
41:58 Right.
41:58 CPUs are becoming what I like to refer to as the new bacon.
42:02 The inference, workloads, orchestration, pre and post processing, agentic AI,
42:07 general purpose, all continue to use CPUs in a very greater detail.
42:14 And so cloud infrastructure will continue to move from general
42:18 purpose Commodity type machines
42:19 to increasingly customize differentiated platforms.
42:23 Now we're going to focus on what's important,
42:25 that performance per dollar per watt that is so critical in our infrastructure.
42:30 So we already do a lot of this custom work today.
42:32 I provided some of the with Microsoft,
42:34 I presented some of those examples earlier and we'll
42:37 continue to expand in this space in the coming years.
42:40 So if I leave you with a prediction
42:42 of what's going to happen in the five year span,
42:44 I can tell you that we're going to launch
42:47 more and more VMs with Azure and Cosmos DB
42:51 and user developers in the planet will be able
42:55 to get benefits from AMD and our collaboration together.
42:59 That's exciting.
43:00 So Cosmos DB customers can look forward to more transactions per
43:03 dollar thanks to the work that we at a better cost efficiency.
43:08 Absolutely awesome.
43:09 Well, thank you so much Steve.
43:10 Thank you for our partnerships.
43:12 Thank you for sponsoring this conference
43:14 and made this wonderful sessions available.
43:16 Appreciate you having me and thank you all for the time.
43:20 Thank you.
43:27 That was amazing.
43:29 Now Cosmos DB is not the only NoSQL database we have on Azure.
43:34 Last year we launched our second NoSQL database, Azure DocumentDB.
43:38 And unlike Cosmos, DocumentDB started from an open source project.
43:42 This is a DocumentDB project at Linux foundation that Microsoft,
43:46 Amazon, Google, Yugabyte and 12 other partners are collaborating on.
43:53 Both Amazon and us are offering fully managed services.
43:57 Amazon DocumentDB and Azure DocumentDB on top of this open source project,
44:01 this is a fully managed document database that has
44:06 a number of advantages over other document databases.
44:09 It's faster, it has faster AI search for example.
44:12 You can deploy it anywhere in any cloud or on premises.
44:17 And it's more than 40% cheaper than any other document database on the cloud,
44:23 including AWS MongoDB Atlas, for example.
44:27 So let's take a look at an example of this.
44:30 Let's say I want to send 1 million
44:32 writes per second to my database on Azure DocumentDB.
44:37 A single instance, single shard 80 allows me to do that no sweat.
44:45 An equivalent scenario,
44:47 equivalent setup on AWS with MongoDB Atlas with the same SKU M80
44:53 would only give me 800k transactions per second according to AWS documentation.
44:59 We have the codes if you want to run it and see for yourself what it does.
45:02 But according to docs, it's only 800k and it cost me 60% more.
45:09 So all in all, I can get twice as many
45:14 transactions per dollar with Azure DocumentDB for every SKU.
45:19 It's at least 40% cheaper on skewer SKU basis overall,
45:23 as you scale your application load more data,
45:26 you get more than 85 reduction in TCO for your application.
45:31 So DocumentDB is really cost effective.
45:35 Take it for a spin and get the cost efficiency.
45:39 Now when should I use which?
45:43 If you are having a lot of simple queries,
45:46 if you want to have a mission critical application,
45:51 if you want to scale limitlessly, if you like serverless,
45:55 form factor, you like global distribution, Cosmos DB is for you.
45:59 This is the only database in the cloud that can do all that.
46:02 But if you prefer open source, if you like MongoDB API compatibility,
46:06 if you want to potentially port your application to other
46:10 clouds or on prem and you like the core architecture,
46:14 then DocumentDB is a good choice for you.
46:19 Multi cloud, open Source, Complex Queries,
46:24 DocumentDB, Lots of Scale, Serverless,
46:27 5, 9 Reliability, Cosmos DB now what if I want both?
46:32 What If I want 5, 9 resilience and I want cross cloud portability?
46:39 As of today you can do it both.
46:42 We created an open source project, we launched it today called MultiCloudDB.
46:48 It's fully open source,
46:49 MIT licensed and it provides you cross cloud portability across Azure AWS
46:56 and GCP by providing one data access layer on top of Cosmos DB AWS,
47:03 DynamoDB and Google Spanner, we provide you cross cloud disaster recovery.
47:14 We will provide you cross cloud replication with sync
47:16 as part of this project you don't have to choose.
47:20 You can have five ninth availability and cross cloud portability
47:24 at the same time using this project and with Cosmos,
47:28 Dynamo and Spanner from each of the hyperscalers.
47:34 Now these are just some of the topics we will cover in the next few hours.
47:38 And while we have your attention, join us for Skill Challenge to win free Cosmos
47:44 DB certification and stay tuned for more exciting sessions today.
47:50 Install Cosmos DB Agent Kit, look for more sessions,
47:54 build an app and maybe you can talk
47:56 about it at Our next Cosmos DB conference 2027.
48:00 Enjoy the rest of the event.
48:06 Thank you so much Kirill.
48:08 Hey, wasn't that a great keynote?
48:12 Thank you so much Kiril.
48:15 Thank you Andrew.
48:16 Yeah, thank you so much.
48:17 There is so much to take in, from AI to real world applications.
48:23 Even with the interview with AMD that was so interesting.
48:27 There's so much that you can do with Azure Cosmos DB.
48:30 Yeah, and you know what, the demo really brought it to life.
48:33 I love AgentKit.
48:34 I want to say hi to Theo and Saji.
48:37 They're doing such amazing work.
48:39 Yeah.
48:39 And if you're enjoying the event so far, please let us know in the YouTube chat.
48:43 We love to hear what you think.
48:46 Up next, we've got a quick story from one of our partners.
48:49 Let's take a look.
48:53 Hi, my name is Mick Feller,
48:54 I'm a distinguished software engineer for the ODP Group.
48:57 But you might better know us as Office Depot.
49:05 We are currently using Cosmos DB mainly, in our homegrown AI,
49:10 application that we built for our enterprise called the Personal Assistant.
49:13 We use it for time series to do analytics,
49:16 but more importantly we use it for, for HR profiles.
49:19 So we store each employee's profile so we can give more context
49:23 to the AI models to be able to answer questions better in contextual form.
49:32 The main benefits we've gained from using Cosmos DB is the hands off management.
49:38 It scales automatically.
49:39 We have full control over how much we use, how much we don't use.
49:43 And it's been just amazing around the whole management of the actual
49:46 tool because the scalability we can control cost to fine grain level.
49:51 And that's why we chose Cosmos DB.
49:58 We've seen almost 100% year over year increase
50:01 in usage and big driver of that was Cosmos DB.
50:05 Because it scaled so flawlessly with our needs,
50:08 we were able to seamlessly scale to that level of usage for the customers.
50:13 That's our success story with Azure Cosmos DB.
50:16 And I can't wait to hear about yours.
50:21 Yeah, welcome back and if you're just joining us,
50:25 yes, you missed a fantastic keynote.
50:29 I can't believe they missed it.
50:30 I know, but don't worry, we've got you.
50:33 Everything from today will be on demand,
50:35 including exclusive sessions so you can catch up at any time.
50:40 So grab the full playlist here at aka Ms.
50:44 Cosmosconf26 playlist.
50:47 Let's give some shout outs.
50:51 Do you want to just say to anybody who's watching,
50:53 I saw some great people from Sweden.
50:56 Yeah, I saw people from the uk, Mexico, all around.
51:01 So tell us where you're watching from.
51:03 We want to give you shout outs if you got
51:06 some interesting stories also about what you're doing with Cosmos DB,
51:09 we may just throw here.
51:10 Oh, we've got a few of our.
51:12 So, well, we've got Mike, from Norwich, thank you so much.
51:18 And then one of our great influencers, Ravi Jain.
51:21 Thank you so much for all the support.
51:23 Yeah, no, thank you so much for being
51:25 here and please keep it coming in the chat.
51:28 So if you can tell us what you're building with Cosmos DB,
51:31 we'd love to shout you out.
51:34 We promise we Won't ask for your partition key unless you bring it up, which.
51:39 Okay.
51:39 Hey, that's why we have that Agent Kit, right?
51:41 Yeah, exactly.
51:42 Andrew showed us.
51:43 If you want to keep connected, aside from the YouTube chat,
51:47 after the event, you can join us on the NoSQL channel on Discord.
51:52 You can go to aka Ms.
51:55 Discord Cosmos Conf.
51:57 Yeah.
51:57 And also the Cloud Skills challenge.
52:00 Yeah, absolutely.
52:01 We want you to get, that ability to win a, 500.
52:07 Excuse me.
52:07 The first 500 of our entrants will be able to enter
52:11 to win one of 500 free certificate vouchers for the DP-420 exam.
52:16 So don't miss out on that opportunity.
52:18 Yeah.
52:18 And if you want to get started, visit AKA Ms.
52:22 aka.ms/cosmosdbconfchallenge.
52:24 And you know what, quick feedback helps a ton.
52:27 So make sure that you fill out our evaluation survey.
52:30 We want to know what you thought, so go to aka Ms.
52:33 Cosmosconf2026survey.
52:36 So mouthful of all those URLs I know.
52:40 All right, well, it's back to keynote programming and it's the last part of it.
52:44 And then we're going to get into all the other sessions.
52:46 But we wanted to do a really great featured session.
52:50 I used to watch the Jetsons as a kid
52:52 and this one kind of near and dear to my heart.
52:55 Jade, stop this crazy thing.
52:56 Cosmos DB Best Practices A lesson from Spacely Sprockets.
53:01 Oh, wow.
53:02 Wait, that sounds amazing.
53:04 Here's Sid Anand from Walmart and welcome to my talk.
53:09 My name is Sid Anand.
53:11 I will be talking today about
53:13 Cosmos DB Best practices and architectural patterns
53:16 in the context of building a fictional E
53:19 commerce site for a company called Spacely Sprockets.
53:23 First, a little bit about me.
53:26 My name is Sid Anand.
53:27 As I mentioned earlier, I'm a technical fellow in the data platform at Walmart.
53:32 I've spent about 25 years in the SF Bay Area working on end
53:37 to end data infra at scale at a variety of companies big and small.
53:44 Let's dive into a, product overview of the site we're trying to build.
53:49 First of all, since it's an E commerce site,
53:51 we expect to see some core shopping experience features such
53:55 as a homepage with some sort of integrated search landing page.
54:00 On this page you'd expect to see some product carousels and a search box.
54:05 A user may search for items or click on one of the personalize
54:09 recommended items and that will take him or her to a product details page.
54:13 If the user likes what he sees,
54:16 he can click on add to cart to add that item to his or her cart.
54:21 And later on he or she might decide to check out and that would
54:26 kick off order processing flows that would result in delivery of this item.
54:32 Last but not least, after the user has bought multiple items from the site,
54:38 he or she may want to go back and look at current and past orders.
54:43 And also, like other modern e commerce sites,
54:46 we'll also have other features like profile and preference,
54:51 payment and billing, authentication and account management.
54:54 All of these things are important to have,
54:56 but they're outside of the scope of this talk.
54:59 We're going to focus primarily on the core shopping experience.
55:04 So let's start with a hundred foot view of what we need in engineering.
55:08 We'll start with a classic three tier
55:11 architecture made up of a presentation tier,
55:14 an application tier and a persistence tier.
55:17 There are many options and choices
55:19 for stacks for presentation tiers and application tiers,
55:23 so we'll just go with what we know best.
55:25 We'll pick React Vue and Angular for the presentation tier,
55:29 and we'll pick a spring boot based, platform for the application tier.
55:36 The next big question we need to ask ourselves is what will be
55:40 the wire protocol and format for data
55:42 that is transacted between our different tiers?
55:46 When making a decision like this, we often have to consider multiple things.
55:51 For example semantic richness.
55:53 Does the wire format, have a rich type system that is natively supported?
56:01 And usually if it does have a rich type system,
56:04 it comes with schema enforcement and schema evolution.
56:10 But sometimes schema enforcement and evolution go against developer velocity.
56:15 So we have to decide which of those is more important for us.
56:18 And because we'll have incidents in production,
56:21 we'll need a way to debug the data flying back and forth between microservices.
56:26 So we'll want to also ensure that we have
56:28 a rich tool and ecosystem to support whatever we pick.
56:32 Most often teams are given two choices or they're faced with two choices.
56:38 Either they go with gRPC and protobuf, which is much more efficient,
56:42 then the other option which is REST over HTTP with JSON.
56:47 However, the rest case with JSON is more widely used
56:51 and it is what we'll pick for our talk today.
56:55 So now we've picked our top two tiers and the interchange format between them.
57:01 We have to make a decision about what
57:03 to pick for our database in the presentation tier.
57:07 While making this decision, we can consider either a tabular,
57:13 data model or a JSON based document model.
57:17 If we pick The JSON based document model.
57:19 It simplifies our mental model and our cognitive load.
57:23 It reduces our cognitive load because what we're transacting
57:26 between microservices is also what we're storing in the database.
57:30 We don't have to concern ourselves
57:32 with flattening from JSON into a tabular format.
57:38 And we also want to keep the SQL
57:40 as our query language because it's very expressive.
57:44 Therefore, we need a database that provides both
57:47 the flexibility of JSON and the expressiveness of SQL.
57:52 Luckily, Cosmos provides both as part of its NoSQL API.
57:57 So now we have a full stack.
58:00 What are some of the other benefits of Cosmos?
58:03 Well, we get scalability thanks to its support for horizontal partitioning.
58:09 It's globally distributed, it provides a range of consistency options.
58:15 So we get tunable consistency and through this we get support for transactions,
58:20 which is very important.
58:22 Finally, it gives us a flexible cost model.
58:25 Now being an enterprise database, it also supports change,
58:28 data capture, snapshot and point in time recovery and other features.
58:36 But this is all very good and what we need in a database.
58:40 Now that we have picked our stack, let's talk about the product.
58:45 While we have five different things to pick from, we're
58:48 going to focus on the three most challenging ones.
58:50 The shopping cart, order processing and also viewing a customer's order history.
58:56 Let's start with the shopping cart.
58:58 What are our engineering requirements for the shopping cart?
59:02 First of all, we need it to be globally distributed.
59:05 If a region has an outage, we don't want an outage in our shopping cart.
59:10 We want people to be able to add to their cart
59:13 and view cart no matter what is happening in a given cloud region.
59:18 Now that we have picked a database that works in multiple cloud regions,
59:22 we also care about the data being consistent.
59:26 If a user makes a change to a cart in one region,
59:29 he should see that exact same change reflected in other regions.
59:34 And last but not least,
59:35 we need all of these interactions to be low latency because any
59:39 type of latency friction will cause a drop off or conversion problem.
59:46 Luckily, Cosmos DB provides all of these.
59:49 So we can create a cart container
59:51 that operates in multiple regions and uses session consistency.
59:56 This provides lower latency and better
59:58 availability for writes than globally strong consistency.
1:00:03 Let's look at this in detail.
1:00:05 Let's say the user wants to add an item
1:00:08 to his cart after he's called that add to cart.
1:00:12 The presentation tier forwards the request to the application
1:00:16 tier which will call upsert on the Cosmos DB instance.
1:00:22 Once it's successful, the database will return,
1:00:26 200 or 201 error code, response code, along with a session token.
1:00:32 And that session token will flow all the way upstream back
1:00:35 to the user where it will be stored in a cookie.
1:00:38 Now let's say after this set of steps is finished, the region goes down.
1:00:44 Like we have a total region, one outage.
1:00:46 But now the user wants to view his cart
1:00:49 because we stored the session token in the cookie.
1:00:53 The user will present that session token when he calls View cart.
1:00:58 That will go to the presentation tier and get forwarded to the application tier.
1:01:03 The application tier will then call query
1:01:05 against the cart collection presenting that session token.
1:01:09 If the data has been replicated from the first region
1:01:13 to the second region by the time this call is made,
1:01:17 the item will be found and returned to the user
1:01:19 so the user can view a consistent cart.
1:01:23 But if there are any delays in replication,
1:01:26 rather than showing the user an inconsistent cart,
1:01:29 what will happen is the application tier will continue to retry until it,
1:01:34 syncs with the initial region, at which point it will show the final cart.
1:01:41 The design provides a globally consistent low latency shopping cart.
1:01:45 And it's exactly what we need for our application.
1:01:49 Now that we talked about shopping cart, let's talk about order processing.
1:01:53 Order processing starts when we're, checking out.
1:01:57 We convert the cart that we check out with into an order.
1:02:01 An order is made up of one or more order line
1:02:04 items where each line item encapsulates a product and its quantity.
1:02:09 Let's look at an example.
1:02:11 Let's say the user had a cart with ID 123 and checked out.
1:02:16 As you can see, in this cart there are three items.
1:02:19 The first one is for the foo sprocket, the second one is for the bar sprocket,
1:02:24 and the third one is for the monkey sprocket.
1:02:26 And each of them come with different quantities.
1:02:29 After checking out, the cart will be converted into an order with ID
1:02:34 A and we can see the order line items are carried over.
1:02:40 So what are the engineering requirements for order processing?
1:02:44 When an order is inserted into the database,
1:02:47 we want the order and all of its line items to be inserted, atomically.
1:02:54 It should also be possible to upsert line items
1:02:58 independent of each other and of the order itself.
1:03:01 On the read side, we should make it really easy for an application
1:03:06 to read and order and all of its line items in a single query.
1:03:10 But they may also want to read either
1:03:12 the order or align item independent of one another.
1:03:16 To achieve this, in Cosmos, we have two options for data models.
1:03:21 Option one gives us a Denormalized model,
1:03:24 also known as embedded or a single document model.
1:03:28 Option two gives us a normalized AKA reference, AKA multidoc model.
1:03:33 Let's have a look at these.
1:03:36 Let's say I pick option 1.
1:03:38 In option 1 I have an order JSON.
1:03:41 I see that the partition key is the order ID,
1:03:44 which happens to also be the ID of the document.
1:03:50 There are other, properties in this document that are relevant and important.
1:03:53 For example the status of the order.
1:03:56 Also, the shipping address and payment address are included,
1:03:58 as well as the customer's id.
1:04:00 And if you notice there is an array type called orderline items.
1:04:04 In this array we have each of the orderline items that are part of the order.
1:04:11 In option two, what we do is we create a parent order document and its
1:04:17 children order line item documents and we
1:04:20 store them together in the same container.
1:04:23 Because they use the same partition key,
1:04:25 they are also part of the same partition.
1:04:29 The order line items have their own line item IDs,
1:04:32 but they are all stored with its parent order in the same partition Now,
1:04:39 how do you choose which model to use?
1:04:42 Well, if you have orders that don't have too many line items,
1:04:47 if you know reliably that there will be less
1:04:49 than 50 of those, option one is a good choice.
1:04:52 It's also a good choice if your dominant access pattern
1:04:55 is always on the order with all of its line items,
1:04:58 such as reading the entire order.
1:05:01 And also if you want simplicity and lowest read ru cost,
1:05:05 option one is a good one for you.
1:05:08 But if the orders can be unbounded,
1:05:11 have unbounded number of line items, you must choose option two.
1:05:16 And option two gives you some other benefits as well.
1:05:18 It allows you to update individual line items.
1:05:22 Because you may want to update the fulfillment status,
1:05:24 for example it's been delivered, and also the quantity.
1:05:27 If you amend your order, you also may want to independently query the line
1:05:33 items from each other from the parent order.
1:05:35 And so the multi doc approach gives you this.
1:05:39 And how does it work?
1:05:40 Well, if you use option one whenever you want to read an item,
1:05:44 in this case read an order, you just call readitem for the order and it'll
1:05:49 return the order and all of its nested line items.
1:05:53 If you want to, modify or delete the order and its line items,
1:05:59 you can also issue point operations.
1:06:01 For those with option two, you get more flexibility.
1:06:05 For example, if I want to, read a parent and all of its children together,
1:06:11 I would use the query shown here.
1:06:13 I would select star from order where the partition
1:06:16 key is in this Case the order id.
1:06:18 But I'm also given the ability to do point reads and read
1:06:22 a single document like a single line item and update its quantity, for example.
1:06:27 And last but not least,
1:06:29 whenever I'm doing any changes across the order and its line items,
1:06:33 I can use transactional batch to either update them
1:06:36 or delete the whole order and the line items.
1:06:41 Now that we've solved two of the hard problems, let's get to the third one.
1:06:44 In this case, a customer wants to see its order history,
1:06:48 which means all of the orders that it
1:06:52 has completed and their associated line items.
1:06:55 And we want to do this with low latency and low cost.
1:06:58 Now, as the order container is partitioned by order id,
1:07:02 querying by the customer ID results in a cross partition query.
1:07:08 Cross partition queries incur higher cost and latency than point queries.
1:07:13 The greater the number of partitions,
1:07:15 the greater the latency and cost of cross partition queries.
1:07:20 Now, does Cosmos provide a way to efficiently query by customer id?
1:07:25 Yes it does.
1:07:26 And it's called Global Secondary Indices.
1:07:30 Let's look at how this works.
1:07:32 Let's start with our normal container, the order container.
1:07:36 As we mentioned earlier, it is Partitioned by Order ID,
1:07:39 but it contains other JSON properties like Customer ID.
1:07:43 Now, I define a new index container called CustomerOrders,
1:07:48 which is defined by a covering query shown here.
1:07:51 Select star from order and a partition key.
1:07:55 In this case, the partition key will be customer id.
1:07:59 Once I set this up, the customer ID will get
1:08:02 live updates from the order via the full fidelity change feed,
1:08:07 filtered by, the covering queries seen below.
1:08:13 Now, the customer orders container is called an index container.
1:08:17 It's a read only container.
1:08:19 Apps cannot write to it,
1:08:21 but writes can only come to it through the order container.
1:08:25 And the beauty of this is a user can query the order container
1:08:29 by order ID or against the customer
1:08:31 orders container by customer ID at both times.
1:08:35 They are point queries, so they're cheap and fast.
1:08:37 We avoid expensive cross partition queries.
1:08:42 Now that we've covered three of the main core pieces, of functionality,
1:08:46 let's talk about some optimizations in terms of latency and cost.
1:08:51 There are three types of optimizations worth considering.
1:08:54 We've already talked about the first one which is to use a gsi.
1:08:58 The second one is about partition scoping.
1:09:02 First of all, what is partition scoping?
1:09:05 Let's look at an example.
1:09:08 Let's pretend that we have a user collection.
1:09:11 That user collection stores documents as shown here.
1:09:14 User JSON and it's backed by a POJO class shown here on the Right user job.
1:09:22 If I want to select like George Jetson from the database,
1:09:27 I would provide a point query as shown here.
1:09:30 Select star from C where C ID equals a bind param.
1:09:34 There's one problem here.
1:09:36 Every time I issue this query, there's an extra round trip that is used
1:09:41 to parse the query and generate a query plan.
1:09:47 If, I execute this query every time,
1:09:51 if I use this code every time I execute the query,
1:09:55 the query will incur an extra round trip.
1:09:58 But if I add the line shown here where I
1:10:01 set the partition key on the Cosmos query request options,
1:10:05 then the first time will be the only time that the query will be,
1:10:11 will generate a query a plan,
1:10:12 and then it will be cached in whichever pod this application
1:10:16 is running in so that subsequent queries avoid an unnecessary round trip.
1:10:21 And this can cut your point query latency by up to 50%.
1:10:27 Last but not least, let's talk about payload size reductions.
1:10:32 What I'm showing here is how RU costs vary as a function
1:10:37 of of the payload size and also based on the type of operation you're doing.
1:10:44 As you see, as I go to larger and larger documents in the payload,
1:10:51 the cost in terms of RUs increase,
1:10:53 but they increase more dramatically for point reads.
1:10:57 Point reads, double in cost, as a power of 2 of the document size.
1:11:05 However, you'll notice that under 32 KB it's cheaper than point queries.
1:11:10 But because point queries have a very flat,
1:11:14 slope over time, they end up costing less.
1:11:19 So let's look at the example from before, where we had an order I.D.
1:11:24 in model one with its array of line items.
1:11:27 In this case, let's pretend I had 100 items in my cart when I checked out.
1:11:34 If, I want to save space, what I can do is compress the order
1:11:38 line items using an algorithm like zstandard and base64,
1:11:43 encode it so that I can store it as a string field, a string property.
1:11:48 I've also shown a, property called compression,
1:11:50 which has a bunch of metadata, but that is not needed.
1:11:53 Assuming we kept that metadata, what will happen?
1:11:57 If we use Zstandard, with its default compression level of 3,
1:12:03 we can save 75% of space,
1:12:06 and this will result in almost, up to 75% in RU reduction.
1:12:13 At this point, I'm, at the end of my talk and I
1:12:15 want to thank the following people for their guidance and support.
1:12:18 Thank you.
1:12:21 Hello everyone.
1:12:22 My name is Chander Dahl and I'm the CEO of Kasden,
1:12:26 an elite AI consulting firm delivering high impact Enterprise grade solutions,
1:12:32 including custom built AI chatbots and advanced
1:12:36 automation systems tailored to complex business needs.
1:12:41 Today we're going to talk about one of the most important concepts,
1:12:45 which is agent memory.
1:12:48 And by the way, it's the same for your chatbots.
1:12:51 And why is agent memory so important and how to get it done right.
1:12:58 So today I'm going to show you four strategies,
1:13:01 starting from a strategy in which we talk directly to LLMs.
1:13:06 And let's not forget LLMs are stateless
1:13:08 and they do not have any memory whatsoever.
1:13:12 And then we'll use the other three strategies,
1:13:15 sliding window, hierarchical and entity graph.
1:13:19 And one of them will lead to the highest precision
1:13:23 and that will be the high precision or recall winner.
1:13:27 However, that doesn't always mean that it's going to be hundred percent recall.
1:13:32 With that particular strategy,
1:13:34 we'll show you different use cases and from that you
1:13:38 can decide what's the best for your particular use case.
1:13:43 We will start with 60 seed messages,
1:13:46 and the reason we chose 60 seed messages is we saw
1:13:50 that in a lot of companies when they do AI testing,
1:13:54 especially on their chatbots,
1:13:56 a lot of times they go up to 10 to 15 conversations only.
1:14:01 And guess what?
1:14:03 All the three strategies provide 100% recall up to 30 turns in most cases.
1:14:10 And that's why by the time you end up in production,
1:14:13 you may not even know that you have less than 100% recall.
1:14:19 And that's a, big no, no.
1:14:22 We also will show you 10 recall questions.
1:14:26 And the first five questions are really easy
1:14:29 and all the three strategies get them right.
1:14:33 However, the next six questions are more nuanced
1:14:36 and that shows how you need to change your testing strategies,
1:14:41 especially if you're leaning more towards
1:14:43 the easier questions during your AI POCs.
1:14:48 So let's have some fun and define the problem.
1:14:52 If you started working like me before ChatGPT was launched on AI applications,
1:15:00 you may have noticed the context limits really small.
1:15:05 The token window limit would be the most important thing.
1:15:10 Before you even started creating these applications,
1:15:13 you would notice right up front that these LLMs actually are stateless.
1:15:19 So after ChatGPT was launched,
1:15:22 a lot of people assumed that LLMs are actually stateful.
1:15:26 And in some cases they even thought that it has a memory.
1:15:30 A lot of that has to be done by engineers like you and I.
1:15:34 You might have heard of the famous quote, I told you my name,
1:15:37 my budget, my team, my deadlines, and you forgot everything.
1:15:42 And that is why memory happens to be
1:15:44 the most important conversation in enterprise applications worldwide.
1:15:50 So now that we know that LLMs have this problem
1:15:53 that they are stateless and every API call starts from scratch.
1:15:58 How do we solve this problem?
1:16:01 Well, very easy.
1:16:02 We can send the full conversation history every single time.
1:16:07 Yes, except we know all LLMs have a limit.
1:16:11 All right, then maybe you can just summarize it.
1:16:15 Let's try it out in case you send the entire
1:16:19 conversation and let's assume that's within the limit an LLM expects.
1:16:26 In that case you're still having costs that are going to explode.
1:16:32 You're also going to slow down the LLM.
1:16:35 What if you start summarizing parts of it?
1:16:39 In that case you still have that issue
1:16:41 that you may be losing some context, maybe some fact,
1:16:47 anything real world that you really care about
1:16:50 that got compressed because all summarizations have a compression problem.
1:16:56 And then finally users will have to repeat themselves quite a bit.
1:17:02 So why does memory matter?
1:17:04 One of the most important things to remember is that more
1:17:07 than 70% of enterprise projects will require multi time agents.
1:17:13 And the number one gap cited by developers happens to be long term context.
1:17:19 And by the way, if you go with the LLM alone,
1:17:21 which really has no memory, and you go with the best memory strategy,
1:17:26 we're going to compare today,
1:17:27 you may be looking at a 20x token cost difference between these strategies.
1:17:33 And that's why memory isn't just a feature checkbox.
1:17:36 It's an architectural decision that determines cost,
1:17:39 recall quality and user experience.
1:17:42 So the first strategy happens to be sliding window which
1:17:45 keeps your recent messages and then summarizes the messages before it.
1:17:50 An analogy would be a security camera that keeps the last eight hours
1:17:54 of footage and it writes a one page summary of everything before that.
1:17:59 That's really easy to implement and that's a strength.
1:18:03 And it's also going to have a low token cost, in this case roughly 1100 tokens.
1:18:09 And it's great for short conversations, especially less than 30 turns.
1:18:14 And this is something to keep in mind.
1:18:16 Just because it's great for that particular use case does not mean it's going
1:18:22 to be great for another use case where we actually require a real world fact,
1:18:27 even the one that was hundred turns ago or a thousand turns ago.
1:18:33 In that particular case, recall is going to be lower.
1:18:36 And in today's demo it's got about a 60% recall.
1:18:40 Let's not forget you will have 100% recall if it is less than 30 turns.
1:18:46 So remember, it isn't about the strategy, it's more about the use case.
1:18:50 And based on that use case, we need to pick the strategy.
1:18:54 Strategy 2 hierarchical memory think of this as three different tiers,
1:18:59 the first year being hot.
1:19:02 For example, a company's knowledge management system will
1:19:05 have your team's messages as your first tier
1:19:08 and that's part of a conversation and you
1:19:10 want to send them as it is uncompressed.
1:19:14 Number two is your weekly meeting notes,
1:19:17 let's call that tier two and kind of compressed
1:19:21 summaries because tier two not as important as tier one.
1:19:26 And then finally the least used out of the three which could be
1:19:30 your company wiki and those are extracted facts in your long term storage.
1:19:36 So better recall than sliding window.
1:19:39 And of course it can do better than it, but not as good as something that will
1:19:45 always present the right fact every single time.
1:19:48 Because tier three, as you can notice it
1:19:51 will miss on certain facts here or there.
1:19:54 Finally we got entity graph and I love this analogy,
1:19:57 it's like a detective's case board where you have all
1:20:00 these different flags and then you got linkages between them.
1:20:03 Right as every single card with connections between them.
1:20:07 And that's entity graph.
1:20:09 So what does it store?
1:20:10 Stores all the entities, Anything real world,
1:20:13 A name, organization, anything real world.
1:20:17 Any entity that is real world gets stored facts.
1:20:21 These are key value pairs for that particular
1:20:24 entity and you can have any amount of these.
1:20:26 And finally embeddings you may have heard
1:20:29 of vectors and vector databases got created simply because
1:20:34 of these LLMs wanting to search basically do semantic
1:20:38 search on your data and get you that particular retrieval.
1:20:44 Now in this case what I love about Cosmos DB
1:20:46 is that you have the native vector search inside Azure,
1:20:50 Cosmos DB and then finally relationships which are connections between entities.
1:20:56 Now why do you get 100% recall in this particular case?
1:20:59 I mean the reason is simple because we don't really have any compression.
1:21:03 We're doing the vector search,
1:21:05 we're getting those entities and those entities have
1:21:08 the facts and we can get all of that back.
1:21:11 So in this particular case, for this particular data set,
1:21:14 we are able to get 100% recall except we're also using average 1660 tokens.
1:21:21 That's pretty much 50% more than the first strategy.
1:21:24 That's sliding window and remember that's cost.
1:21:27 So that's another consideration.
1:21:30 So even though I would say there are four approaches,
1:21:32 you might have noticed I start with zero.
1:21:35 Why is that?
1:21:36 That's because I'm an engineer and as engineers
1:21:39 we use arrays and we start with zero.
1:21:41 That's not the case.
1:21:43 I really don't think it's an apples to apples
1:21:45 comparison because LLMs really have no memory, no context.
1:21:49 What you're trying to do unless you provide it.
1:21:52 So for that reason I don't think it's a good comparison.
1:21:55 But I still want you to see that you may be able to get 92 tokens,
1:21:59 but you got a 0% recall.
1:22:02 However, the other three strategies, very important, as you can see,
1:22:06 we go from 1100 tokens all the way to 1660.
1:22:09 And in this particular scenario, you go from 60% recall to 100% recall
1:22:14 at a, one and a half times the token limit.
1:22:17 But here's what I love about Azure Cosmos DB.
1:22:20 Same database, same SDK, same partition key, different recall guarantees.
1:22:25 And I can use all these three strategies at the same exact time.
1:22:31 So that leads me to why Azure Cosmos DB For agent memory.
1:22:35 Well, that's because we are in Azure Cosmos DB conference.
1:22:39 That's not the case.
1:22:40 It's really because at the end of the day all my data is in Azure Cosmos DB.
1:22:46 And by the way, for someone like me who's been working on DocumentDB, that was,
1:22:50 I think it was 2016 onwards I started working
1:22:52 on it and then it became Azure Cosmos DB.
1:22:55 And I know we now have another DocumentDB which is
1:22:57 a little different than that DocumentDB that was there long back.
1:23:03 My data and my client's data lives on Azure Cosmos DB.
1:23:06 It used to be a big problem to now have another vector storage,
1:23:11 a completely new database just for the vectors.
1:23:15 Well, not anymore we can actually have that data.
1:23:18 Your embeddings, your vectors live right alongside your data.
1:23:21 I don't actually need to have another native vector search database.
1:23:26 The schema is flexible and we talked about that.
1:23:29 I, love the session isolation.
1:23:31 My partition key in this case happens to be session id,
1:23:34 so I have zero cross partition queries.
1:23:37 And then finally, for this demo, I'm not worried about global scale at all
1:23:42 because all we have is about 60 different messages.
1:23:44 But once I go into production,
1:23:47 I don't have to worry about global scale because now I
1:23:51 can scale as much as Azure Cosmos DB would allow me,
1:23:55 which is literally all over the world.
1:23:58 And that's why Azure Cosmos DB.
1:24:01 So let's talk about the benchmark where we have about 10 recall questions.
1:24:05 You got the core questions, basic recall.
1:24:08 Why do we have that?
1:24:09 A lot of times what we noticed was all these POCs
1:24:12 that people were doing had a bunch of the same kind of questions.
1:24:16 And a lot of times those questions were created by AI itself.
1:24:19 Well, but it was also 100% recall based on questions
1:24:24 that probably don't understand how these strategies actually work.
1:24:27 So what you'll notice is 100% recall
1:24:31 for all three strategies for the core questions.
1:24:35 We also notice that every single data set is different.
1:24:38 We work in healthcare, you know, manufacturing, finance,
1:24:42 fintech, tech, pretty much name any major domain.
1:24:46 And all these bigger corporations worldwide
1:24:48 in those domains have very nuanced data.
1:24:52 That's why I wanted to show you how much difference gets made in just five
1:24:56 questions the moment you start asking nuanced questions
1:25:00 on that particular data and how the Mrs.
1:25:04 Happen.
1:25:05 So for example, in this case,
1:25:06 sliding window only gets one out of those five correct.
1:25:10 So your recall happens to be only 20%.
1:25:13 But if you have the right use case,
1:25:15 sliding window could be at 100% and that's why it gets confusing.
1:25:20 Before we do the demo, I would like for you to ask yourself which strategy is
1:25:25 going to remember a, URL mentioned once 40 plus turns ago?
1:25:32 And I hope you get the answer right.
1:25:34 Let's dive into the demo.
1:25:36 As you can see here, I've got two providers.
1:25:38 Here's Azure OpenAI and you can also click OpenAI.
1:25:40 It'll work with both the keys.
1:25:42 So if you have Azure OpenAI,
1:25:44 assuming you have the same exact keys for whatever model you're trying to use,
1:25:49 here's two different demos.
1:25:50 One is Kasden and another one is default.
1:25:53 And remember, they're both fake data, so please feel free to use your own data.
1:25:58 I'm not going to ask anything from the Direct LLM because
1:26:01 at the end of the day we're not going to get any answers.
1:26:04 Now we're going to compare the other three strategies which is sliding window,
1:26:08 hierarchical and entity graph on the right hand side.
1:26:12 It's just a way to see what the results are and what to expect.
1:26:17 So let's ask the first question and see
1:26:20 what it does while it's answering the question.
1:26:23 I just want to show you that all three of these are going to pass.
1:26:27 So this is actually a really easy question.
1:26:30 No problem whatsoever.
1:26:31 All of these strategies will get you the same exact answer,
1:26:33 which is cast and offerings are consulting, training and recruiting.
1:26:39 Same exact answer pretty quick.
1:26:42 And you'll notice hierarchical was extremely fast,
1:26:45 but entity graph is actually slower still gets you the same three answers.
1:26:51 So now I'm going to skip the next four questions and go to number six.
1:26:56 The reason is because the next four questions we're
1:26:58 going to get the same result from all the strategies.
1:27:03 So here, here's a result which says about us in contact us.
1:27:08 All right, so let's go to round number six
1:27:11 and open this and you notice it's missing Consulting.
1:27:15 Let's go to Hierarchical.
1:27:18 Run this round.
1:27:20 It's got about us and it's got consulting, which is true.
1:27:23 And that's exactly what we were expecting.
1:27:25 On the right hand side, here's your Entity graph.
1:27:32 It's got about us and it's got, consulting.
1:27:34 It's got both.
1:27:36 The answer is correct.
1:27:38 All right, so now let's go back to Sliding
1:27:40 window and we're going to go to number seventh.
1:27:44 And here we're going to open the round and see what we were expecting.
1:27:47 All right, it's got a about, it's got webpage, it's missing products,
1:27:54 and it's also missing from our stored graph context, which is true.
1:28:00 Let's go to Hierarchical.
1:28:05 And as, you can see, it's got pretty much everything,
1:28:08 except it's still missing products and also
1:28:11 missing from our stored graph context.
1:28:15 Let's go to Entity Graph.
1:28:21 And you may have noticed Entity Graph is actually slower.
1:28:24 So that's another thing to keep in mind is what your, use case is.
1:28:30 So you've got, you've got actually everything.
1:28:32 You got products, you've got the About Chander Dahl, you've got the webp.
1:28:36 It's actually perfect.
1:28:38 And it also says from our Stored Graph context.
1:28:41 That's pretty amazing.
1:28:43 Let's go back to the eighth round.
1:28:45 I'm going to pass because they all get the training URLs.
1:28:48 It's the same exact answer for all three of them.
1:28:51 And now let's just show one more.
1:28:53 And let me go back to Sliding Window.
1:28:56 Let's go to number nine, very quick.
1:28:59 And it's got something.
1:29:01 Let's see what it is.
1:29:02 Okay, so it's actually missing presentations, which is true.
1:29:07 Let's go to Hierarchical.
1:29:09 It's got presentations, so hierarchicals got all of those.
1:29:13 And same with Entity Graph.
1:29:20 And then the 10th round, it's a very nuanced round.
1:29:24 All that is getting wrong is not the response.
1:29:28 They have all the responses correct, both of them,
1:29:30 except they're missing from our stored graph context.
1:29:37 All right, so how many of you got that right?
1:29:40 I hope it's all of you.
1:29:42 So what you just saw was directllm has no prior contacts.
1:29:46 We didn't do that part of the demo.
1:29:48 But the idea is you're not gonna get any response if you didn't.
1:29:51 If the LLM does not have that data, unless you fine tune it to have that data.
1:29:56 Whereas, Sliding Window, the first five questions correct.
1:29:59 The next six, it only got one right, which is the eighth one,
1:30:03 which we didn't run but it was really easy and it was going to run.
1:30:07 Now you've got hierarchical, which was three out of ten,
1:30:11 which means in this case about eight out of ten.
1:30:14 So you notice what's happening, even though it was three out of five.
1:30:17 Sorry, not three out of ten.
1:30:19 We had the first five really easy questions and you may think you have an 80%
1:30:25 recall where all you had was a 60%
1:30:28 recall because three out of five were correct.
1:30:31 An entity graph in all the cases was 10 out of 10.
1:30:34 And, one of the things I want you to remember is that if
1:30:36 you had a sliding window which never requires you to go beyond 30 turns,
1:30:42 you would have 100% recall and sliding window.
1:30:45 So it's really not the recall as much as it is the combination of your strategy,
1:30:49 your use case and your data done right?
1:30:52 So here's your scorecard.
1:30:55 If you remove the blue, which is the first five questions,
1:30:57 sliding window is only one question.
1:31:00 That's 20% recall.
1:31:01 Hierarchical, 3 out of 5, that's 60% recall in entity graph, 10 out of 10.
1:31:07 So.
1:31:07 And 5 out of 5 and 5 out of 5.
1:31:09 So that's 100% recall cost versus recall.
1:31:13 The trade off is right here.
1:31:14 And it's important because if your use case is literally less than 30 turns,
1:31:19 you're better off going with sliding window because the cost is
1:31:22 one and a half times more in terms of entity graph.
1:31:25 You're using the least amount of tokens.
1:31:27 You're hoping you get the response also faster.
1:31:30 So that's another cost to keep in mind is performance cost.
1:31:33 Otherwise, if you really care about higher recall, well,
1:31:37 in that particular scenario,
1:31:38 entity graph would be a much better deal altogether.
1:31:42 So an easy way to look at this would be sliding
1:31:44 window for less than 30 turns supports the chat recent contacts hierarchical,
1:31:50 about 30 to 100 turns.
1:31:52 And again, you know, this is more,
1:31:54 it'll depend on your data, but let's just say ballpark,
1:31:57 that's what you're looking for, especially if you're using it
1:32:01 for planning or consulting and you have to have some key facts.
1:32:05 And then entity graph, for example,
1:32:07 you have your CRM bot or something where every fact matters,
1:32:11 like you don't want a response
1:32:13 with a financial number that's actually not correct,
1:32:16 you know, and it can also scale to way more than 100 terms
1:32:20 because at the end of the day you're really getting the fact correctly.
1:32:25 So again, you can start with one, you can also upgrade later.
1:32:30 You can also do a hybrid, which is also not a bad idea.
1:32:34 And in some cases you can have a combination
1:32:36 of, let's say a sliding window as well as Entity Graph.
1:32:39 You know, for 90% of your cases you may have sliding window,
1:32:43 especially for less than 30 turns.
1:32:45 And then you can go to Entity Graph
1:32:48 whenever there's a premium use case behind it.
1:32:51 So again, memory is a spectrum and you need to choose by recall, not by height.
1:32:57 And you gotta remember what's your use case and what your data is like.
1:33:03 What we covered was the problem,
1:33:05 which is LLMs are stateless, they have no memory.
1:33:08 Then we really covered the baseline plus
1:33:11 the next three strategies and then the architecture,
1:33:14 which was really easy by the way.
1:33:16 You must be wondering why five Cosmos containers to remember for Entity Graph.
1:33:21 You will have different containers for Entity Graph.
1:33:24 And that's not at all a bad idea in this particular case.
1:33:28 That's why we did that.
1:33:29 You can, you can do it multiple different ways,
1:33:31 but for that reason we had more than just one container.
1:33:36 And then you've got the live demo which had 10 recall questions.
1:33:40 Then you had the evidence which was 100% recall in terms of Entity Graph.
1:33:45 And then the decision,
1:33:46 which isn't really about the recall numbers that we shared here,
1:33:49 but a lot more than that and we discussed every single of those.
1:33:54 Now just remember this is not production data.
1:33:58 And this is literally fake data.
1:34:00 But this is also not production level code.
1:34:02 However, it's a good POC if you want to run.
1:34:05 And that's your barcode.
1:34:06 If you click that barcode,
1:34:08 it will take you to this blog post that explains the nuances,
1:34:12 how the decisions were made,
1:34:13 and it also explains to you the different parts of the code.
1:34:17 I just want to show you this real quick for your code to remember.
1:34:23 Here's your direct LLM.
1:34:25 And if you're a dot net developer you can look at the right
1:34:27 hand side of the code and if you're a Python developer,
1:34:29 you can look at the left hand side of the code.
1:34:32 Directllm.
1:34:33 It's literally sending the system prompt and then
1:34:36 creating the message chain and then making a call.
1:34:40 Very simple.
1:34:41 That's how you've been making all your calls so far.
1:34:44 Next, the sliding window.
1:34:45 And as you can see here we have
1:34:47 the summary for anything beyond 30 messages and then
1:34:51 just 30 messages where everything is uncompressed and it's
1:34:55 part of the context we're sending to the LLM.
1:34:59 And as you can see here, the window size is 30 and this is your one hour,
1:35:03 which is we just chose that time to live.
1:35:06 And that could change for whatever you're trying to do.
1:35:10 And then you've got the hierarchical memory
1:35:11 and here is where all your tiers are.
1:35:14 And as you can see, here's your tier one size,
1:35:17 your tier two block and the max tier.
1:35:20 Same for both Python as well as Net.
1:35:23 And then you've got the entity graph.
1:35:24 This is actually really easy to produce because
1:35:27 your entity entities and your embeddings are living together.
1:35:32 So one thing to remember is that this is exactly how we anyways code
1:35:37 when we are doing programming because we have the same exact objects now we just
1:35:42 have a way to retrieve those objects and then make it part of of your LLM
1:35:48 answer by giving it the context which it needs to give you the response.
1:35:53 And finally here's your read path on entity graph.
1:35:57 And this is just a small blog post that I prepared only for the demo.
1:36:02 However, a lot more detailed blog post is right here.
1:36:06 If you want to do a more involved code walkthrough,
1:36:09 you can click that link and it goes through a lot more than what I just shared.
1:36:16 So let's not forget,
1:36:17 here's the link to that and I hope you take advantage of it.
1:36:21 Let me know if you have any feedback for me.
1:36:24 I had a blast and hope you have
1:36:26 a blast coding and using this in your applications.
1:36:29 Thank you for having me.
1:36:31 Have a great day and a great rest of the conference.
1:36:37 Welcome back.
1:36:38 Thank you so much to our speakers.
1:36:40 I believe that was Sid or Ann Chander.
1:36:43 They were both great, right?
1:36:45 Yeah.
1:36:45 Honestly that was a very fun and super practical look at Cosmos
1:36:49 DB's best practices through the story of get excited, spacey sprockets.
1:36:54 Jane, get me off this crazy thing.
1:36:57 Right?
1:36:58 Well, at least if you don't have a well architected partition key.
1:37:01 Right.
1:37:03 Always on availability,
1:37:05 low latency across the galaxy and keeping costs under control.
1:37:09 Those are all trade offs we all face when we build real systems.
1:37:14 But you know what, we're going to talk way more about
1:37:17 that and I want to just make sure that we're acknowledging you, the audience.
1:37:21 So let's take a quick look.
1:37:23 We've got some really, really nice comments, but I wanted to just show this one
1:37:29 the thought processes people go through when designing solutions.
1:37:34 You know, there's a lot you have to make.
1:37:36 So many considerations when you're designing a production application.
1:37:40 And you know what?
1:37:41 You all are here to learn about that.
1:37:42 Patty, what do you got?
1:37:43 Yeah, honestly, for me,
1:37:44 I've got nothing but love in this group chat or in the comments section.
1:37:49 So please, if you are around.
1:37:53 We'd love to hear from you.
1:37:54 Feel free to drop a comment, share what you're building or, just show some love.
1:38:00 We're back into the program.
1:38:02 Thank you for the love.
1:38:03 Oh, thank you.
1:38:04 Thank you for the love.
1:38:05 We're shifting into one of the most important
1:38:07 patterns in Azure Cosmos DB in our next session.
1:38:11 One of my favorite teammates.
1:38:12 I know yours too, Justine Cocchi takes us
1:38:15 deep into the Azure Cosmos DB Change feed.
1:38:17 How it works under the hood,
1:38:19 the patterns that hold up at scale and the live demo of debugging and lag,
1:38:24 recovering processing and production is going to be great.
1:38:27 Yeah, you know Jay, Justine is amazing.
1:38:30 And if you love Justine.
1:38:31 Yeah.
1:38:31 And if you haven't used Change Feed, every write becomes a durable ordered event
1:38:36 stream that you can fan out to services, search and analytics,
1:38:39 event driven architecture without staying on a message bus.
1:38:43 So let's hear it all from Justine.
1:38:46 Here's mastering the Azure Cosmos DB Change Feed patterns,
1:38:49 scaling and real world architectures.
1:38:51 We'll see you soon.
1:38:56 Hi, I'm Justine Cocchi and I'm a program manager on the Azure Cosmos DB team.
1:39:00 I'm really excited to talk to you about Change Feed.
1:39:03 Today we'll cover what it is, how to read it,
1:39:05 and some tips for debugging for real world applications.
1:39:10 So first, what is the Change feed?
1:39:12 It's a persistent record of all changes to items
1:39:15 in your container that's ordered by modification time.
1:39:18 As your client app is making writes to your items,
1:39:21 you can then read them in the Change feed with your consumer app.
1:39:25 Your Change Feed consumers are processing changes per partition.
1:39:29 So let's take a look at what that looks like under the hood.
1:39:32 Because Cosmos DB is a distributed database,
1:39:35 your data is distributed across multiple physical partitions in a container.
1:39:40 In this example, I've got three physical partitions.
1:39:43 Each physical partition can have its own independent change feed
1:39:46 reader that's reading the changes in order for that partition.
1:39:50 Order is guaranteed within a logical
1:39:53 partition key in your physical partition range.
1:39:57 Your continuation token tracks progress of that change
1:40:01 feed processing for a given container.
1:40:04 And this means that each of your partitions
1:40:06 can have independent continuation tokens and independent processing.
1:40:09 This is great for efficiency because it really allows you to scale your chain
1:40:13 tree processing even for extremely large
1:40:16 containers with a high volume up writes.
1:40:20 There's a couple of different Change feed modes that you can choose from.
1:40:23 The default Change Feed mode and what's enabled
1:40:26 on all of your accounts is latest version mode.
1:40:29 In this mode you get the latest create or replace
1:40:32 for every item items that are deleted from your container.
1:40:36 No longer appear in the feed, and this means you won't get a notification
1:40:40 for deleted items in the change feed itself.
1:40:43 Any item that still exists in the container,
1:40:45 you can read the latest version of it.
1:40:47 There's infinite retention of all of these changes
1:40:49 in your container and you can go back and reading
1:40:52 even from the very beginning of your container
1:40:54 to get all of the items and their updates.
1:40:58 This is really great for a lot of change
1:41:00 feed scenarios and powers most change feed apps today.
1:41:03 There's also all versions in deletes mode.
1:41:06 With all versions and deletes, you get every create,
1:41:09 update and delete including TTL expirations.
1:41:13 Items for TTL expirations will appear in the feed in order of their purge time,
1:41:18 which may be slightly later than their actual expire time.
1:41:22 Retention of changes in this feed is based
1:41:25 on the continuous backup window for your account.
1:41:27 So either seven days or 30 days, depending on what you've configured,
1:41:31 you can read all of these changes within that window.
1:41:34 This mode is really great for audit logs or real
1:41:37 time processing where you need to react to every change.
1:41:40 Consider an application that is replace heavy where you have a high
1:41:44 volume of updates to a single item within a short period of time.
1:41:48 For all versions in deletes mode, you would get every single update in the feed,
1:41:52 whereas latest version you would only get
1:41:54 the latest version per change feed pull.
1:41:57 Depending on the needs of your application, each mode may be a better fit.
1:42:01 So it's important to consider what your app is trying
1:42:04 to do and choose the mode appropriate for your app.
1:42:09 There's a couple of different ways to actually consume the change feed,
1:42:12 and we're going to cover three of these ways.
1:42:15 The first is the change feed processor.
1:42:17 This works on a monitored container
1:42:20 or the container that you're reading your changes from.
1:42:23 When you create a change feed processor,
1:42:24 you can create multiple instances to handle scaling of all the partitions.
1:42:30 And this is the compute environment that you're
1:42:34 actually deploying your change feed processor on.
1:42:37 So you can imagine this like AKS,
1:42:39 where you'd have multiple pods that represent multiple instances.
1:42:43 The code that's actually being executed when a batch
1:42:46 of changes is processed is your handler delegate.
1:42:49 This is like the business logic or the brains of your change be processor.
1:42:53 Here is where you handle what you're reacting to for the change.
1:42:58 Maybe it's writing to something downstream or maybe you're doing some analytics.
1:43:03 This is really where you write your change feed
1:43:07 logic and the Change feed processor is available in both.
1:43:10 Net and Java.
1:43:11 Now, as your changes are being processed,
1:43:14 the lease container helps to Maintain state and checkpoints for keeping track
1:43:20 of where you are in processing for all of your various partitions.
1:43:24 Let's take a look at some lease
1:43:25 management and really understand how our physical partitions
1:43:29 map to leases which map to the instances
1:43:32 we've already established as a distributed database.
1:43:35 Cosmos DB can scale out to multiple physical partitions
1:43:38 and each physical partition owns a range of partition key values.
1:43:43 That top row is our physical partitions.
1:43:47 In this example we have six physical partitions
1:43:49 and each of them have their own lease.
1:43:52 Leases are always one to one with partitions
1:43:54 and this is managing the state of where we are processing.
1:43:57 In terms of the modifications to that individual partition.
1:44:02 You'll notice they each have their own
1:44:04 continuation token which is what maintains the state.
1:44:08 Now your Change Feed processor is deployed across multiple
1:44:11 instances and each instance owns a number of leases.
1:44:15 These are evenly distributed across instances to ensure
1:44:19 that as you're processing you can keep up with these changes.
1:44:23 Assume you only had one instance that's handling all of your leases.
1:44:27 If that instance goes down, it might impact the resiliency of your application.
1:44:31 However, if you have more instances than the number of physical partitions,
1:44:34 you would have idle instances that aren't able to do any work.
1:44:38 The Change Feed processor automatically rebalances leases across
1:44:42 these instances to ensure the maximum efficiency of your processor.
1:44:46 Let's say one of your instances goes down.
1:44:49 The leases that are owned by that instance will
1:44:51 automatically be rebalanced across all of the other healthy options.
1:44:56 The next way to read change feed is through the Cosmos DB trigger.
1:45:00 This is an Azure functions trigger that is
1:45:04 a really simple way to read change feed.
1:45:06 It's built on Change Feed processor under the hood.
1:45:08 So all of the lease and distribution of partitions
1:45:10 that we just reviewed still applies in the Azure functions trigger.
1:45:15 However, the management is done for you.
1:45:18 Because it's built on Azure functions, it's serverless.
1:45:21 It has built in auto scaling to dynamically react to the number of instances
1:45:26 you really should be having without over
1:45:28 provisioning instances that are not actively being used.
1:45:32 This is a really easy way to build the change
1:45:35 feed and the logic of your Azure function is like
1:45:38 the delegate in the Change feed processor where you're just
1:45:41 responsible for the business logic of processing the changes themselves.
1:45:46 The last option is the Change Feed pull model.
1:45:49 This model gives you the most control over how to read the change feed,
1:45:53 but also means you need to have all the safeguards put in place yourself.
1:45:58 So there's no lease container for maintaining state.
1:46:01 You instead use continuation tokens where you're
1:46:04 responsible for storing and maintaining those continuation tokens.
1:46:08 You're responsible for any error handling.
1:46:10 There's no built in retries.
1:46:13 You need to handle all of that in your client code.
1:46:16 While it does give you full control, it's also more code to manage.
1:46:21 One additional benefit of the pull model say you
1:46:24 wanted to only process changes for a specific partition range.
1:46:27 This is possible in the pull model because again,
1:46:30 you have full control over what is actually being pulled.
1:46:34 Now that we've got our Change Feed app,
1:46:36 let's talk a little bit about monitoring and how we can debug
1:46:39 some common failure modes to ensure that our Change Feed app is resilient.
1:46:45 Some of the key metrics to monitor for Change Feed are the estimated lag.
1:46:50 This measures the gap between the latest change that actually occurred
1:46:54 in your container and the last processed item in your change feed.
1:46:58 This is going to be your most
1:46:59 important metric for monitoring your change feed health.
1:47:02 While this estimated lag gives you a number of outstanding changes,
1:47:06 it is intended to be an estimate and it's not
1:47:08 always the exact number of changes that are actually pending.
1:47:12 The key thing to measure here is not exactly what the number is,
1:47:15 but the slope of how the number changes over time.
1:47:18 Is your lag steadily increasing?
1:47:20 Is it decreasing?
1:47:21 Are you staying constant?
1:47:22 This will really help you understand if your Change Feed app is healthy.
1:47:27 Some other things you can take a look
1:47:28 at is the throughput of your change feed processing.
1:47:31 You can look at the number of items that are processed
1:47:33 per second and the RU consumption of your change feed processor.
1:47:38 It's important to look at our use not only of your source
1:47:41 container that you're monitoring the change
1:47:42 feed of, but also the lease container.
1:47:45 If you have a lot of lease rebalancing
1:47:47 and redistribution that can put strain on your lease container,
1:47:51 which typically is provisioned with a relatively low number of RUs.
1:47:54 Monitoring these things will ensure
1:47:56 that your change feed application is healthy.
1:47:59 Digging deeper into the leases, you can look at the owner distribution.
1:48:03 Are your leases evenly distributed across instances?
1:48:06 Is the last checkpoint increasing for each of them,
1:48:09 or is one of your leases stalled with processing?
1:48:13 For setting up the monitoring itself
1:48:15 you can use application insights or OpenTelemetry,
1:48:17 which is built directly into the SDK to make it really
1:48:21 easy for you to export this telemetry for your change feed handlers.
1:48:26 For alerting.
1:48:26 Some things that you might consider creating
1:48:28 alerts for is sustained lag growth or if
1:48:31 there's any partitions that have stalled
1:48:32 and are no longer making progress in processing.
1:48:36 We can look at a couple key failure
1:48:38 modes and Some resilience patterns to help combat these.
1:48:42 The first is as Changes arrive, they're processed in batches.
1:48:46 When everything is going smoothly in the happy path,
1:48:50 your process handler succeeds and your lease checkpoints are all updated.
1:48:55 However, because these changes are processed in a batch,
1:48:58 what if there's an error?
1:49:00 Unhandled exceptions can create issues in your processing.
1:49:05 While the change feed processor does have retry,
1:49:07 you don't want to infinitely retry these poison messages
1:49:10 that are not able to be handled after multiple attempts,
1:49:14 and instead consider writing them to something like a dead letter queue.
1:49:17 This will allow you to write the item
1:49:20 that is failing to take a look at it later,
1:49:23 instead of blocking the entire processing for that container for that partition,
1:49:27 because your partition would not be able to continue
1:49:29 making progress if you're getting a consistent error.
1:49:34 There may be other causes of silent stalls for a given partition,
1:49:39 and one of the common causes is outgoing calls.
1:49:42 So if you're making calls to an outgoing service and they're failing,
1:49:45 this can also cause errors in your processor that will stall your change feed.
1:49:51 You can implement the circuit breaker pattern to ensure
1:49:54 that you don't flood this downstream service and inhibit,
1:49:57 your ability to actually recover from these errors.
1:50:01 A lot of the key resilience patterns
1:50:03 for any application also apply to change feed,
1:50:05 and it's important to keep these in mind
1:50:07 for the most resilient application and processing.
1:50:11 Now, we also talked a little bit about RU throttling.
1:50:15 For consistent RU throttling,
1:50:16 you can consider implementing priority based execution
1:50:20 to ensure that your mainline transactions on your container,
1:50:25 maybe your writes and your reads,
1:50:27 perhaps they have a higher priority than your change feed processor.
1:50:31 With priority based execution, you can specify that in the client.
1:50:35 To ensure that the change feed is not
1:50:37 accidentally consuming extra RUs that you'd maybe prefer,
1:50:41 go to the rest of your application.
1:50:44 There's also throughput buckets that will
1:50:46 help you indicate the percentage of RUs that you want to distribute across
1:50:50 your various applications and including your change feed.
1:50:55 Now, while you may see ru, throttles across your entire container,
1:50:59 it's also important to consider hot partitions because the change feed
1:51:04 works at the partition level and your progress is per partition.
1:51:08 If you have a surge of writes that is generating a hot partition,
1:51:12 it may be difficult for your change feed to keep up.
1:51:15 Ensure that your leases are appropriately balanced across partitions,
1:51:19 across instances, so that if there is a hot partition,
1:51:23 it doesn't overload one of your instances
1:51:25 and result in higher lag across the board.
1:51:31 All right, let's take a look at a demo to see how the Change
1:51:35 Feed processor is actually configured and see
1:51:37 it in a real application application,
1:51:38 I'm going to pull up VS code with a simple app here.
1:51:42 I've got a social media app simulator which is really just two console apps.
1:51:48 My first console app is simulating users creating,
1:51:51 editing and deleting some of their posts.
1:51:54 And the program that we've got on screen here is
1:51:56 a second console app that has our Change Feed processor.
1:52:00 Now, because we know that we've got updates and deletes
1:52:02 and I want to react to all of them,
1:52:05 I'm going to use the Change Feed processor with all versions and deletes mode.
1:52:09 When I configure my Change Feed processor, I can choose a processor name.
1:52:14 This tells me the unique processor instance and will allow me
1:52:19 to coordinate multiple compute instances all to the same processor application.
1:52:23 This is really critical for least rebalancing
1:52:25 because if you're using different processor names,
1:52:28 the change feed will assume that it's four actual different,
1:52:33 change feed instances and different purposes.
1:52:37 So we've got the processor name which sort of ties all
1:52:39 of our instances together and we also have our instance name.
1:52:43 We'll take a look in the console app how the instance name allows
1:52:46 us to tie multiple different processes
1:52:48 of these together into that same processor.
1:52:52 I'm using the same lease container for all
1:52:54 of these, which is storing all those checkpoints.
1:52:57 And I'm also setting up a couple of notifications.
1:53:00 These lifecycle notifications are really important
1:53:03 for debugging the lease ownership in your application.
1:53:06 And you can get notifications for lease acquired and lease released,
1:53:11 as well as printing out some, error messages that may occur.
1:53:15 So this is the Change Feed processor itself.
1:53:18 But let's also take a look at the delegate.
1:53:20 Once we get a batch of changes,
1:53:23 this is the code that will execute and actually process those changes.
1:53:27 So first we'll simulate some extra processing with a short delay.
1:53:33 And of course, in your real application
1:53:34 this is whatever business logic you may have.
1:53:37 Then we can process our changes
1:53:40 slightly differently depending on the operation type.
1:53:42 The first thing we want to check for is any deletes.
1:53:45 The ID and partition key of deleted items
1:53:47 will come through in the metadata so we
1:53:50 can pull out those key pieces of information
1:53:52 to be used later for creates and replace operations.
1:53:56 We can directly get the ID and the user id,
1:53:59 which is our partition key from, the current aspect of this change.
1:54:06 Scrolling down a little bit further,
1:54:08 this change feed is writing to a notifications container
1:54:13 which then can power notifications for our social media app.
1:54:17 So now that we Looked at the code, let's see it running in action.
1:54:21 And I've got three separate terminals pulled up here.
1:54:24 The first thing I'm going to do is start running my social simulator.
1:54:27 This console app is going to be creating
1:54:30 items and posts from users in our application.
1:54:33 Then I'm going to run my notification processor.
1:54:36 Now notice I'm not submitting any arguments here,
1:54:39 it's just that basic, code that we talked about.
1:54:43 And it will have a default instance name which we'll see spin
1:54:47 up here in just a second as, the detailed stats come online.
1:54:52 Now in the meantime,
1:54:53 I'm going to start a stream of posts in my application and let's make sure
1:54:58 our processor is printing in verbose mode so
1:55:00 we can see all those creates coming in.
1:55:02 But we're using all versions and delete.
1:55:04 So let's simulate some delete events and some edit events and you
1:55:08 can see all of those are actually flowing through into our processor instance.
1:55:14 All right, let's turn it back into quiet mode so we don't have every single,
1:55:18 change, but just the batches of changes that are coming in.
1:55:21 And let's simulate a couple of viral bursts.
1:55:25 There's a lot of buzz on our social app and users are making
1:55:29 more posts than ever as we see these viral bursts start to flow through.
1:55:34 We see the lag is steadily increasing on our change feed processor instance.
1:55:40 So we can create a second instance to help us,
1:55:43 manage all of the changes that are coming through for all
1:55:46 of these partitions and balance it a little bit better across both of these.
1:55:51 So I'm going to print out the detailed stats for our second instance.
1:55:57 And you notice that this one is named Instance 2 as it starts up.
1:56:02 We start by acquiring Elise and we get an error on our initial prefacer.
1:56:07 Now this is really important because this is not actually a bad error.
1:56:12 This error is just telling us, hey, I lost a lease and it went to someone else.
1:56:16 This is the behavior that we actually
1:56:18 want change Feed processor is automatically rebalancing
1:56:21 leases for us so that we can
1:56:23 ensure they're evenly distributed across our processors.
1:56:26 If we print out the leases that are actually owned by each of these, we can see,
1:56:31 our first instance owns three of our leases and our second instance owns two.
1:56:36 This is great because this means that our processors
1:56:39 are working properly and our load is evenly distributed.
1:56:43 Now I'm going to go ahead and shut back down our second
1:56:46 instance and we should see these leases
1:56:49 redistribute back onto our initial processor.
1:56:52 We see those lease acquired messages now and when we print the leases
1:56:56 owned this time we see all five leases are back onto our main processor.
1:57:01 Let's go ahead and shut down this application
1:57:04 and we see how easily Change Feed allows us
1:57:08 to spin up new instances and dynamically react
1:57:11 to the volume of processes and changes coming into our container.
1:57:18 Flipping back into slides now that we've shown a basic Change Feed app,
1:57:22 let's take a look at some advanced patterns.
1:57:25 We can combine the event sourcing CQRs and materialized views
1:57:29 pattern and to create a really powerful application in Cosmos DB.
1:57:33 While our demo showed creates updates and deletes
1:57:37 inline to items with the event sourcing pattern,
1:57:41 every write, every create, edit,
1:57:43 delete actually becomes its own independent create event.
1:57:46 So it's a depend only event store.
1:57:49 We can write these commands into our event
1:57:52 store still stored in Cosmos DB as container and we can use Change Feed to read
1:57:57 this ordered log and populate some downstream systems.
1:58:01 These are also known as materialized views and it helps us craft a view
1:58:06 of our data that is better suited for the read patterns of my application.
1:58:13 In my app I've got a couple of read patterns that I want to support,
1:58:16 like show me my notifications.
1:58:18 This is that notifications container that we
1:58:20 were just populating in our example.
1:58:22 But let's say we also want to get posts by topic.
1:58:26 Or maybe we want to create a search index container.
1:58:29 We can use built in full text search with Cosmos DB
1:58:33 to create a second container that has full text policies on our content.
1:58:39 This allows our users to really easily search across those data patterns too.
1:58:45 With Materialized Views pattern.
1:58:48 Setting up all of these separate containers based
1:58:50 on the read pattern that we're trying to optimize really
1:58:53 gives us the best configuration for Cosmos to give
1:58:57 us these efficient reads without impacting our write path.
1:59:01 Our write path is still optimized for the writes.
1:59:05 Now the Materialized Views pattern is really common in many databases,
1:59:09 but I do want to highlight a Cosmos
1:59:11 DB feature that makes this pattern really easy.
1:59:14 We have a feature called Global Secondary Indexes which
1:59:17 are effectively managed materialized views directly into Cosmos DB.
1:59:22 So while I just showed you how you can
1:59:24 kind of build this yourself with the Change feed.
1:59:27 Reading the change feed of social events,
1:59:29 populating a new container with a new partition key,
1:59:32 you can add a global secondary index.
1:59:35 This will do the exact same thing for you,
1:59:37 but it will automatically maintain the secondary container without you
1:59:41 having to write the code to build and manage changepie yourself.
1:59:45 This is a really powerful container,
1:59:48 this is a really powerful feature that allows you
1:59:50 to build containers specialized to the read patterns that you have.
1:59:54 My container, I had five physical partitions in my source.
1:59:57 If I were to use my source container to serve this query by topic,
2:00:01 it would be really inefficient,
2:00:03 it would be slow because I'm checking every physical
2:00:05 partition and it also would cost a lot of RUs.
2:00:09 With GSI I'm able to automatically create this and I
2:00:13 don't have to manage the change feed myself.
2:00:16 So let's take a look at the Azure portal
2:00:18 and we can see how this is actually implemented.
2:00:22 I have my Change Feed application,
2:00:25 my Cosmos DB here that was powering that Change Feed app we
2:00:28 just looked at and you can see my source container is social events.
2:00:34 Now my source container is partitioned on user ID and we
2:00:37 see all of the posts that my users are making here.
2:00:41 But I added a GSI already.
2:00:43 Let's take a look at our GSI and we can see this has
2:00:47 a persistent copy of all of the items in our source container,
2:00:50 but now it's partitioned by topic
2:00:52 and this is automatically maintained by the Change feed.
2:00:56 We ensure any writes only go to social events.
2:00:59 We don't need to handle any dual write or error handling.
2:01:01 That's all given to us by the platform
2:01:03 and we're able to have this read optimized view.
2:01:07 Now let's take a look at a query.
2:01:10 I want to get the top 100 items where the topic is Cosmos Conf,
2:01:15 and this first query is against my social events container.
2:01:18 If I look at my query stats I can see this costs almost 27 RUs.
2:01:23 So that's quite a bit of RUs to get these topics.
2:01:27 But if I issue that exact same query against my global secondary index,
2:01:32 flip over to query stats.
2:01:34 I see this costs just about five and a half RUs.
2:01:36 So this gives us an roughly 80% savings.
2:01:40 All because we're using the GSI which
2:01:42 is partitioned and configured for our read pattern.
2:01:50 This demo really shows us how you can use Change Feed
2:01:55 and manage features to give you
2:01:57 the best configuration for your specific application.
2:02:01 All right, we covered a lot today, so the key takeaways are that Change Feed
2:02:06 really powers a lot of event driven applications.
2:02:09 There's many architecture patterns that help
2:02:11 you build these apps for your business.
2:02:13 Use case.
2:02:15 When you're using the Change Feed, consider the Change Feed mode that is best
2:02:19 suited for your application and also the consumption model.
2:02:22 Whether that's Change Feed, processor Azure functions or the pull model,
2:02:26 it's important to invest in monitoring and you can use
2:02:30 the estimated lag to ensure that your processors are Keeping up.
2:02:34 Keep resiliency best practices in mind in your handlers.
2:02:38 Ensure that they're idempotent as retries are happening and their resiliency
2:02:43 is all the best practices with circuit breaker retry timeouts, et cetera.
2:02:49 Advanced features like Global secondary Index really help
2:02:52 you simplify these patterns and let the platform
2:02:54 do the heavy lifting instead of you
2:02:56 needing to manage that Change Feed application yourself.
2:02:59 Now you can get started and learn even more about Change Feed at this link,
2:03:03 aka MsAzureCosmos DB ChangeBead.
2:03:07 Thank you so much and I hope you enjoy the rest of the conference.
2:03:13 Okay.
2:03:14 Hello everybody.
2:03:15 So, this is MultiCloudDB.
2:03:17 Write once, run anywhere.
2:03:19 I'm Theo van Kraay, I'm a PM in the Cosmos DB team.
2:03:23 I work on SDKs, connectors,
2:03:26 developer experience and lots of fun stuff like that.
2:03:29 But I'm going to be doing something very different this time.
2:03:32 Probably the first time we've done anything like this, on Cosmos DB Conf.
2:03:37 I'm going to be talking about something that we're calling MultiCloudDB.
2:03:41 This is a unified data access layer for best of breed cloud databases.
2:03:47 So for the first time ever we're talking about databases other than Cosmos DB.
2:03:52 But before I go into that, before I talk about what that does,
2:03:55 I want to set some context here.
2:03:58 So why teams still want cloud native databases?
2:04:03 Of course, databases like Azure, Cosmos DB, Amazon,
2:04:07 DynamoDB, Google Spanner, are still very highly sought after.
2:04:11 They offer high availability and elasticity built
2:04:14 from the ground up in those platforms.
2:04:16 These platforms are built to exploit cloud core,
2:04:22 properties in those environments.
2:04:24 They have strong operational guarantees and SLAs and manage scaling.
2:04:28 It's simply not possible to get this same kind of performance
2:04:32 if you're installing a traditional database software and putting that onto
2:04:37 VMs for obvious reasons and of course you have a fast
2:04:39 path to production without owning the database or the control planes.
2:04:44 Everything is very easy, it's very elastic,
2:04:46 it's very scalable, it's very highly available.
2:04:50 The benefits there are pretty obvious.
2:04:52 But teams tend to want managed database upside,
2:04:55 without the irreversible coupling that you tend to get.
2:04:59 There's a tension here.
2:05:01 The deeper that you adopt native SDKs of these types,
2:05:04 of databases and query models, the harder it becomes to move anywhere else.
2:05:13 So this becomes something that is concerning for customers,
2:05:17 if they want to exit the architecture or move into another cloud
2:05:22 and Portability concerns show up pretty much before any line of code is written.
2:05:27 And so sadly for us working on Cosmos
2:05:29 DB&M for our counterparts in Amazon and Google,
2:05:33 often these very good database services get excluded
2:05:37 right off the bat because there's no portability,
2:05:40 they only run in Azure or Amazon or Google etc.
2:05:45 And so the cloud native benefits are obviously significant,
2:05:48 everybody knows that, but so is architecture level lock in and risk.
2:05:52 And this is very real and we are obviously recognizing this.
2:05:56 So the customer reality is that customers do want
2:06:00 the benefits of cloud native databases as we've said,
2:06:02 but they also need a story for portability,
2:06:05 procurement and regional strategy and so on.
2:06:08 This is the gap that we're trying to fill with MultiCloudDB,
2:06:11 and it's designed to close.
2:06:13 So portability is no longer optional for teams
2:06:17 who are building these types of applications,
2:06:20 operating across regions, customers and different clouds and so on.
2:06:23 We see two different types of multi cloud
2:06:27 portability pressure if we can put it that way.
2:06:29 The one is cross cloud deployment.
2:06:31 So this is where you have the same solution delivered into multiple clouds,
2:06:35 typically white labeled products,
2:06:37 sovereign deployments or maybe partner managed environments,
2:06:40 where the application needs to behave consistently even when
2:06:44 the backing database changes by the customer or region, et cetera.
2:06:49 And then the other one maybe is more familiar is where you
2:06:53 want to deploy into one cloud but you want a credible exit strategy.
2:06:57 In other words you want to avoid cloud vendor locking at least.
2:07:02 And portability obviously matters there as well.
2:07:05 Even if the team never actually switches cloud providers.
2:07:11 In both cases the requirement is the same.
2:07:13 Application code doesn't want to have to be
2:07:16 rewritten if you're moving into a different cloud.
2:07:19 And again sadly for us working Cosmos DB this means a great database product
2:07:23 is often excluded right out of the bat because it only runs in Azure.
2:07:32 So there are some abstractions out there in the Java world you have things
2:07:37 like Spring Data and Hibernate of course
2:07:39 and the Net world you have Entity Framework,
2:07:41 Core and Python you have things like django etc.
2:07:46 That offer something close to this.
2:07:47 So you'll get something like a programming to an interface model.
2:07:51 In Spring Data you have repository abstractions and object mapping and so
2:07:55 on and you get something like a productive lowest common denominator.
2:08:00 The problem with this is that they're
2:08:04 not based strictly on portability contracts.
2:08:06 This is a side effect of developing an abstraction that is Meant to simplify
2:08:14 development on many different databases rather
2:08:17 than it being the core design goal.
2:08:20 So what you get over time is divergence still appearing.
2:08:24 You still get query features drifting by backend annotations and mappings code,
2:08:29 level changes that you need to apply even though
2:08:32 more or less things tend to be very similar.
2:08:35 So portability is sort of possible in part but it's not guaranteed by design.
2:08:39 And in practice this shows up when you
2:08:42 want to migrate from one cloud to another.
2:08:43 You still have in some cases almost as many changes if not quite as many.
2:08:49 And divergent still accumulates because as I
2:08:51 said this is not the core design goal.
2:08:55 It's a side effect of some principles to create an abstraction that you
2:08:59 get some level of portability but it's
2:09:01 not really completely portable in that sense.
2:09:06 So why MultiCloudDB and what we've built
2:09:08 here and what we are building is different.
2:09:11 We have a strict portability design goal from the ground up.
2:09:15 This is not a side effect capabilities explicitly
2:09:18 surface what is and what is not portable.
2:09:21 The promise that we're making is write once,
2:09:24 run anywhere at semantics MultiCloudDB
2:09:28 treats portability as a product guarantee,
2:09:31 not a best effort convenience that you get
2:09:34 as a side effect of some goals of abstraction and whatnot.
2:09:39 So again design goal strict portability same code
2:09:43 and query code should run unchanged across Cosmos DB,
2:09:47 DynamoDB and Google Spanner should be portable by default.
2:09:50 Transparent limits.
2:09:52 We will be providing some native escape hatches
2:09:54 but we don't expect customers to be using this.
2:09:58 The whole point of this abstraction is that you shouldn't be using those because
2:10:03 that will obviously break portability and that's
2:10:05 what you want if you're in this space.
2:10:08 So simple architecture view here you have a Java application
2:10:11 that's going to be connecting to multi CloudDB client and then
2:10:13 you have a service loaded discovery that's going to load
2:10:16 the appropriate module that is being driven by config files only.
2:10:21 And this is set very similar to what you get out of things like Spring,
2:10:24 and Hibernate and so on.
2:10:26 But again as we've said this is portability from the ground up and strict
2:10:30 guarantees around not only the feature set
2:10:34 but the behavior of those features as well.
2:10:37 And it follows from that we need a portable query dsl.
2:10:40 So you can see on the left there this is MultiCloudDB query dsl the same
2:10:45 regardless of the platform that's running on it
2:10:47 would be different on the right hand side,
2:10:49 syntax differences between those different databases and so on.
2:10:53 It also follows that we have a fairly conservative feature roadmap.
2:10:58 Bottom line is example a feature will not
2:11:03 be supported if all providers cannot support it.
2:11:06 So if it can't be supported in every
2:11:08 available provider then it's simply not shipped.
2:11:13 So you might be wondering at this point why
2:11:15 would we be crazy enough to do something like this?
2:11:20 And the answer is of course we've done this before.
2:11:22 Those of you who are Familiar with Cosmos DB,
2:11:24 Azure Cosmos DB was the first database platform
2:11:27 to build broad compatibility APIs at this level.
2:11:31 Those of you who've used MongoDB API, Cassandra API,
2:11:34 Table API, Gremlin API and so on, these were compatibility APIs,
2:11:40 APIs sitting on top of Cosmos DB,
2:11:42 providing the surface area of a different database along on the same platform.
2:11:47 This gives us years of real world experiences handling compatibility drift,
2:11:52 behavior gaps, edge case semantics, et cetera.
2:11:55 We've learned that how portability fails in practice
2:11:58 where feature gaps create friction and so on.
2:12:01 How to set clear contracts and manage the relevant trade offs et cetera.
2:12:07 Why this matters.
2:12:09 Well this is not our first rodeo as my American
2:12:12 colleagues would say we're applying hardened compatibility discipline
2:12:16 to multi cloud DB from day one and we
2:12:20 actually think this is the easier problem of the two.
2:12:24 Compatibility APIs are hard.
2:12:26 Fitting a database engine to a programmability surface
2:12:30 area that we don't control is well difficult.
2:12:34 Let's say surface area keeps expanding,
2:12:37 behavior contracts keep shifting, goalposts move continuously.
2:12:41 It's very difficult.
2:12:43 Now I'm not going to say that multi clouddb is easy
2:12:45 but certainly from our perspective given
2:12:47 our experience it's a more tractable problem.
2:12:51 We can constrain the programmability layer that we control across
2:12:56 a focused set of databases and therefore we define the contract,
2:13:02 the portability of the contract.
2:13:03 We scope that to known database targets and we
2:13:06 can enforce limits explicitly so we have more control.
2:13:09 Even though there are challenges, will be challenges around supportability.
2:13:14 It's certainly a more tractable problem than the one that we
2:13:17 have already solved in the past with Cosmos DB nonetheless.
2:13:23 So of course where this fits best,
2:13:27 of course if your most important things that you want to solve
2:13:31 that you want the benefits of cloud native databases, the elasticity,
2:13:35 the high availability guarantees,
2:13:38 the SLAs and so on the managed service aspect but you also
2:13:42 want Portability then this is going to be a great fit for you.
2:13:46 Of course there are trade offs to accept as there are everywhere.
2:13:49 Strict portability means some trade offs.
2:13:52 Of course you are going to get access to more granular features if you're
2:13:56 using the native SDKs for Cosmos DB and DynamoDB and Google Spanner and so on.
2:14:03 So bottom line you want to be choosing
2:14:06 this when portability and the benefits of a cloud
2:14:09 native database are really the higher order
2:14:11 bit for your architecture and for your strategy.
2:14:15 So then again ultimately the key takeaways,
2:14:18 managed cloud databases are still very,
2:14:20 very attractive to many customers for some very good high availability,
2:14:25 elasticity, performance maintenance, operational guarantees.
2:14:28 The list goes on.
2:14:29 The blocker isn't really the database capability, far from it.
2:14:34 It's of course the application level coupling to one provider's SDK.
2:14:38 It's lock in.
2:14:39 We don't like to say the words lock
2:14:41 in but that's really what we're talking about here.
2:14:43 And MultiCloudDB is designed to blow that open by having
2:14:50 a strict portability design from the ground up so teams can
2:14:54 keep an exit strategy and deploying cross cloud scenarios without giving
2:14:59 up the massive benefits of modern managed cloud databases like Azure,
2:15:04 Cosmos DB and our counterparts in Amazon and Google and so on.
2:15:10 So before I do a demo I just want
2:15:11 to mention we do have some documentation and you'll see there,
2:15:14 it says at the bottom there in small print.
2:15:16 The SDK is currently available as a public preview.
2:15:19 It is not yet fully ready for production use.
2:15:22 Expect breaking changes, incomplete features, limited support during this phase.
2:15:28 But we definitely will encourage you to try this out,
2:15:31 raise issues in GitHub, ask questions, make feature requests, et cetera.
2:15:36 We want to know how you want to use this, what this is solving for you.
2:15:39 We're definitely committed to it.
2:15:42 So yeah, check out the documentation.
2:15:43 There's a nice getting started guide, a developer guide as well.
2:15:47 It's fully comprehensive with everything that we have so far
2:15:50 and obviously we're going to be
2:15:51 developing on this, information about compatibility,
2:15:53 a couple of samples that you can get
2:15:55 your teeth into, an API reference and so on.
2:16:00 All right, so without further to do, let me set up my demo here.
2:16:05 Okay, so here we are.
2:16:06 In VS code I have an application called Risk Platform App.
2:16:11 It's a multi tenant app that is giving a dashboard
2:16:14 of different applications that are providing risk management services and so on.
2:16:21 And this is built on Multi Cloud DB and you
2:16:24 can See the multi cloud client is being used here,
2:16:28 and that is going to do a bunch of things and one
2:16:31 of the things it's going to do is load a config file.
2:16:34 And again I mentioned earlier,
2:16:35 the configs are the only thing that are different per provider.
2:16:39 So you've got one for Cosmos DB here
2:16:41 and one for DynamoDB and everything else is identical,
2:16:45 the code is the same regardless of which provider you're pointing to.
2:16:48 You don't need to change it.
2:16:49 You can treat it as basically the same database
2:16:52 and we guarantee that the behaviors and the functionality
2:16:55 that is available to you should be the same
2:16:58 and performance differences within some
2:17:00 acceptable limits of a statistical difference.
2:17:04 So let me just start up my Cosmos DB,
2:17:09 version here with the Cosmos DB properties that I have.
2:17:12 So let's go ahead and start this one up.
2:17:17 Okay, so that's started up.
2:17:18 So let me go into this app here and we can see I've just put
2:17:23 a little moniker there just based on the config
2:17:25 file to denote which database it's pointed to.
2:17:27 But everything else should be identical if we were going to run this in Amazon.
2:17:33 And so I can see here I've got different portfolios,
2:17:36 I click on these, I'm going to see
2:17:39 a different list of positions on those portfolios.
2:17:41 So just to prove it let me go back here and let me cancel
2:17:45 this app and let me start it up
2:17:49 again but use the Dynamodb properties files instead.
2:17:57 Okay, so that started up, now you can see there provider, DynamoDB.
2:18:01 Instead of just clicking on the link I'm just
2:18:03 gonna go straight back to my original link here
2:18:05 and I should able to refresh or hit that URL
2:18:08 and you see it's just now pointing to Dynamodb.
2:18:11 But everything else is exactly the same.
2:18:14 The data is the same and all the behaviors should be identical.
2:18:20 And if we were to look in the portal for Cosmos DB,
2:18:23 now of course this is where things may look different.
2:18:25 So you'll have databases and containers in Cosmos
2:18:28 DB for those of you familiar with Cosmos DB.
2:18:31 In Dynamo of course you have tables but you don't have databases.
2:18:35 So we prefix this with the database
2:18:37 name in order to replicate the resource hierarchy.
2:18:40 And so that's from the SDK perspective that's what it's abstracting.
2:18:43 But of course the data is ultimately identical
2:18:45 and most importantly of course how the SDK interacts
2:18:48 with the data and the results that come
2:18:50 back and the behaviors should be identical as well.
2:18:53 And ultimately that's what you will care about
2:18:55 if you want to maintain a single playbook,
2:18:57 a single set of application code for your solutions
2:19:00 that are being deployed into different cloud environments,
2:19:02 or if you want to make sure that you
2:19:05 have an easy migration and exit strategy from one,
2:19:08 cloud environment to the other.
2:19:12 So that's it.
2:19:13 That's everything for MultiCloudDB again, check out the documentation,
2:19:20 and the links that we gave here earlier on.
2:19:23 Thanks for watching.
2:19:30 Hey everybody.
2:19:30 My name is Andrew Ruffin.
2:19:31 I'm a senior Product Manager here at AMD and I
2:19:34 focus on launching and growing new products and services on Azure.
2:19:36 With AMD under the hood, you're here at Cosmos DB conference most likely because
2:19:41 you care about building systems that actually work at scale.
2:19:43 They're globally distributed, they're highly available,
2:19:46 and they're predictable under real workloads.
2:19:48 Azure Cosmos DB gives you all of that as a managed service,
2:19:51 and that abstraction is incredibly powerful, especially as a managed service.
2:19:55 This all comes by design.
2:19:57 They come from real engineering decisions, joint enablement,
2:20:00 and the deployment of actual infrastructure across
2:20:02 the global scale that only Azure can provide.
2:20:05 The hardware choice isn't your concern,
2:20:06 but the physics absolutely stick around and they matter.
2:20:10 And that's where AMD comes into the picture.
2:20:12 Today, Azure Cosmos DB runs on modern AMD EPYC processors across the globe.
2:20:17 We launched AMD EPYC in the cloud on Azure nearly 10 years ago,
2:20:21 and we've continuously executed and generated new products and services
2:20:25 that give the features that you need generation over generation.
2:20:29 You'll hear more details on how Azure
2:20:31 Cosmos DB takes advantage of AMD infrastructure,
2:20:33 including the latest 5th gen processors later on.
2:20:37 While you, as an Azure Cosmos DB user, don't need to spend any cycles
2:20:40 worrying about the hardware or the infrastructure.
2:20:42 You get the benefits in terms of cost
2:20:44 savings enabled by efficient and sustainable performance across Azure,
2:20:48 you'll find benefits when you're running on top of modern EPYC processors.
2:20:52 The newest V7 generation, for example, provides up to 35% more performance
2:20:56 in performance per dollar over the prior generation.
2:20:59 With these newest generations integrated into Azure Cosmos DB,
2:21:02 you get more performance per request unit
2:21:04 and a better experience to the many dynamic scaling
2:21:07 and elasticity features that the team continues to roll
2:21:09 out and you hear more about later today.
2:21:12 This is all enabled by the close partnership
2:21:13 we at AMD have with Microsoft as a whole,
2:21:15 and it creates a feedback loop between hardware and software development.
2:21:19 This results in optimized performance per dollar
2:21:21 per watt for the most advanced workloads,
2:21:23 like the ones you run every day on Azure Cosmos DB.
2:21:26 So thank you to the Azure Cosmos DB
2:21:28 team and all the organizers of the conference.
2:21:30 We appreciate the opportunity to help put
2:21:32 this all together and make this possible.
2:21:33 Enjoy the show.
2:21:36 Welcome back and a quick thank you to Andrew Ruffin and our partners
2:21:40 at AMD for being part of the Azure Cosmos DB conf.
2:21:43 Yeah, absolutely.
2:21:45 They were a key part of how we were able to extend this show to five hours,
2:21:50 more sessions, more speakers and more depth.
2:21:53 Quick reminder, you can explore the full agenda of all those talks at AKA Ms.
2:21:59 Azurecosmos dbconf.
2:22:00 All of the talks.
2:22:02 All of them.
2:22:03 All of them.
2:22:03 They're going to be live.
2:22:04 Tell them about it.
2:22:05 Yeah, there's going to be $0.21 sessions total and every
2:22:08 session will go live before the end of today's show.
2:22:11 You can find all of today's links up in the YouTube description, a lot of links.
2:22:16 So you're going to be looking past the comments.
2:22:19 So, we're going to talk a little bit about agents AI,
2:22:22 but before we do that we've got some comments
2:22:24 from you that have to do with agents in AI.
2:22:27 So let's take a look.
2:22:28 So, one of our viewers, a very, very engaged viewer, Chenola,
2:22:34 has said I asked the chat and I'd love you to answer those questions too,
2:22:39 what your AI agent development, environment may look like.
2:22:44 They said we are using GitHub Copilot at the enterprise level.
2:22:48 We've developed our agents to run widely with Copilot CLI,
2:22:53 using it for code test DevOps.
2:22:56 That's huge.
2:22:57 It's an end to end kind of experience.
2:22:59 And then Sandra says, I am a DB admin who usually uses the Azure
2:23:05 Portal for DB management and Bicep templates for deployment.
2:23:09 So they do have a copilot with PowerShell, for doing that automation.
2:23:13 So.
2:23:13 Hey, Patty, I hear you also got
2:23:15 another question about something related to AI coding.
2:23:18 Yeah, also from St.
2:23:20 Andrew.
2:23:21 First of all, thank you so much for the great comment on the keynotes.
2:23:24 Now the question is, about the agent skills.
2:23:27 Does it only work with the extension?
2:23:29 No, you can use NPX to install it.
2:23:34 Installation and usage instructions should be in the documentation.
2:23:40 Yeah, but what do we have, Jay?
2:23:42 Absolutely.
2:23:43 And up next, when you're building with AI agents,
2:23:46 token usage isn't just a billing detail, it's a design construction.
2:23:50 And so every extra turn and chunk of context adds up.
2:23:55 So you need an approach that stays efficient as your agent scales.
2:23:59 And I know I may get access to lots and lots of tokens,
2:24:03 but not everyone else does.
2:24:05 I know efficiency really is key and that's where memory becomes critical.
2:24:11 Farah Abdou will show you how an agent memory
2:24:13 layer on Azure Cosmos DB helps you reuse Cosmos context,
2:24:16 reduce tokens and keep response quality high while improving
2:24:20 latency and keeping costs as predictable as you scale.
2:24:23 Absolutely.
2:24:24 And I love this session.
2:24:26 Farah is so great.
2:24:27 This is cutting AI agent costs with Azure Cosmos DB, the agent memory fabric.
2:24:33 Here's Farah.
2:24:36 Hello everyone, my name is Farah Abdou.
2:24:39 I'm a lead machine learning engineer and AI researcher.
2:24:43 And today I will walk you through how
2:24:45 to cut AI agent costs with Azure Cosmos DB.
2:24:50 This is a war story about building an AI system
2:24:54 that costs way too much money and how I fixed it.
2:25:00 So for the agenda today you are going to talk about the problem.
2:25:04 Why is the multi agent AI based model the solution?
2:25:08 We have one database, three features,
2:25:11 then a live demo and the real production numbers.
2:25:17 So global AI spending in 2026 will be $2.5 trillion.
2:25:25 The agent market grows from 7.84 billion to 52.62 per year.
2:25:34 This is between 2025, 2030, a compound of annual growth,
2:25:40 rate of 46.3% and 75% of large enterprises will be running AI agents by 2026.
2:25:51 And inference demand will grow about a, thousand times by 2027.
2:25:57 And this will keep that host law is not an optional thing.
2:26:04 It's how we can stay in the business.
2:26:12 So this is the architecture that everyone is building right now.
2:26:16 Store data systems, a cache, a relational database,
2:26:19 a vector database and event pass.
2:26:22 We have four terms for this, four filler points,
2:26:25 we have dual coordination between them.
2:26:28 And this is exactly what I built first.
2:26:30 It worked, but it didn't work after this.
2:26:39 So what about the production?
2:26:43 So in the demo, first of all we have three agents,
2:26:48 we have 100 requests, it's about six donors.
2:26:53 Then we come to production.
2:26:54 In production we have now three agents,
2:26:57 still our agents, but now we have 10,000 requests.
2:27:03 This will cost us $18,000 a month.
2:27:06 Nothing changed about the design, we just got real traffic.
2:27:15 So we have five filler modes.
2:27:17 I have seen with my own eyes, number one is about the cost exclusion.
2:27:23 We have from 5 to $15 in a demo becomes 18,000 to $90,000 a month in production.
2:27:33 Number two is the reliability collapse.
2:27:36 So 95% per agent multiplied across three agents equals for example like 85.7%.
2:27:45 This is 14:30 every day at 10,000, huge number.
2:27:55 So number three is about the latency cascade.
2:27:59 For example agent one let's say take three seconds.
2:28:02 Agent B take about four seconds.
2:28:05 Agent C take about five seconds.
2:28:09 There are 12 seconds total plus 200 to 500 millisecond coordination between
2:28:15 the handoff and for the state corruption
2:28:20 like parel rights with no concurrency control.
2:28:24 It means silent data and 5 is the context degradation.
2:28:32 Agency gets only 60% of agents a original entity after 47 steps.
2:28:45 So what is the reliability mass that kills you?
2:28:49 One agent at 95%, 500 filler a day at 10,000 rubles.
2:28:59 Three agents at 85.7%.
2:29:03 1,430 filler per day.
2:29:08 We have five agents at 77% 2,300 Philip and they're gonna
2:29:16 say 40% or more of agent project will fail by 2027.
2:29:22 So look at this math.
2:29:23 I think huge numbers the little Pill now
2:29:30 that nobody plants are Winterbeat analyzed 100,000 production queries.
2:29:38 18% were exact duplicates.
2:29:42 47% were semantically similar.
2:29:45 Same meaning, different words.
2:29:48 Only 35% were genuine new questions.
2:29:51 When I run the same analysis on my own system,
2:29:55 about 65% of my LLM Spent was answering question I already answered.
2:30:05 So the breaking point is $47,000 a month.
2:30:11 40% of the sprint time debugging agent fillers.
2:30:15 Multi agent debugging takes three to five
2:30:19 times longer than the single agent debugging.
2:30:22 And only 11% of the organizations have AI agents in production at all.
2:30:29 We were in that 11% but it didn't feel like success.
2:30:34 Something now had to change and this is why we are here today.
2:30:41 We have one database, three capabilities.
2:30:44 The agent memory subgroup which is built entirely
2:30:49 ON Azure Cosmos DB for NoSQL this replaces four systems.
2:30:54 Cut the cost 73%, cut the latency 65% and let
2:31:00 me show you now how So why Azure Cosmos?
2:31:07 Because we have four properties that make
2:31:11 it uniquely suited for the multi agent footprint.
2:31:14 First, the native vector search using DiskANN sub 20 milliseconds
2:31:20 at 10 million vectors under 100 milliseconds at one period.
2:31:26 The second is a change feed like
2:31:29 the built in event streaming no message queue needs.
2:31:33 And the third is the ETag optimistic concurrency at the document
2:31:39 level no distributed op manager is here and the global
2:31:44 distribution so multi region 99.999% SLA single digit multi second point
2:31:54 triads all in one SDK one string connection all in one.
2:32:04 Now we know that a traditional cache matches exact strings 18% hit rate.
2:32:15 A semantic cache matches meaning using the vector similarity 65% hit age.
2:32:25 The extra 47% is the money you are leaving with a table evidence the query
2:32:32 on the screen does it in one SQL call vector distance with a threshold of 0.80.
2:32:43 If the incoming question scores at or above that returns a cache answer.
2:32:48 No LLM.
2:32:53 Now for the embedding pipeline we have four components.
2:32:56 An embedding model which will be a text
2:33:00 embedding.3 small by the Azure OpenAI 1536 dimensions.
2:33:07 Then we have a Cosmos DB container with a distance in vector index DTL
2:33:14 item from 32nd to 7 days partition key which will fill the query hash.
2:33:20 And we have the vector index configuration which where we set the indexing
2:33:26 search list size between 10 and um.500 to tune the accuracy versus build speed.
2:33:33 And then the lookup query which will be the select top
2:33:37 one with the vector distance greater than or equal to 0.80.
2:33:42 Below that call the LLM And store the new entry.
2:33:50 So now let's have a look how this works.
2:33:54 Like.
2:33:59 So we have here, this is the entire cache lookup one SQL query.
2:34:13 We have the vector distance built into the Cosmos DB.
2:34:19 No extra library, no separate vector database.
2:34:24 Then I have some agents to try on.
2:34:28 This is the first one with three questions and then we have a next
2:34:32 run with another three question to see the cache hit and the cache.
2:34:39 So if I run the demo, we have here three pilots on the screen.
2:34:47 Semantic cache change feed, optimistic concurrency, all in one database.
2:34:52 The first agent as you can see here, like for example it have a support agent
2:35:00 with what is the refund policy for electronics?
2:35:03 It have a cache miss.
2:35:05 It doesn't find this before.
2:35:07 So now we have the answer.
2:35:09 The second agent, the return agent.
2:35:12 How can I return an electronic product?
2:35:14 The third agent called the product itself or anti on laptops.
2:35:20 So now we have missed right now what if I just Change the question again?
2:35:29 So I will just comment this uncomment the second
2:35:34 one or the second run it has similar questions,
2:35:38 similar meanings but not exact same writing.
2:35:43 So here for example what is the electronics return undefined policy
2:35:47 Then join item I put Then the warranty comes with that.
2:35:52 So let's run this one more time.
2:36:03 So as you can see we have here a cache
2:36:06 hit similar to 0.89 which is more than that 0.80.
2:36:13 We have got it in the slide.
2:36:16 The second one is also a cache hit similar to 0.91.
2:36:21 The third is also cache hit with 0.92 so
2:36:25 it finds all of this info on the database.
2:36:28 It doesn't need to call the LLM again and cost me money.
2:36:33 So how much it saved you?
2:36:35 It saved me.
2:36:36 Now just 327.
2:36:39 Okay, let's go back to our slides.
2:36:46 So what you just saw was three agents, three different questions, zero impulse,
2:36:52 all answers compact in about one second I
2:36:58 say what is electronics return and refund policy?
2:37:01 And I found it matches the refund policy for electronics.
2:37:04 For example at 0.89 similarity a traditional
2:37:08 cache would have made a semantic cache.
2:37:12 Catch it every time.
2:37:16 So the production results and this is the result
2:37:20 from a real production deployment which is constant with published benchmarks.
2:37:28 Before $47,000 a month,
2:37:32 803.50millisecond average latency with 10% cash hit rate.
2:37:39 After that I have $12,700 a month.
2:37:45 These 100millisecond average three months,
2:37:49 300millisecond average latency and 65% cache hit rate and 73% cost reduction,
2:38:01 65% latency improvements, one architecture.
2:38:06 Change it all.
2:38:10 So what about the disk?
2:38:11 And at scale, let's see.
2:38:14 So under 20 milliseconds at 10 million vector was greater than 90% recall.
2:38:24 We have under 100 millisecond at 1 billion vectors.
2:38:30 This is still greater than 90% recall.
2:38:34 70iu per query let down a complex SQL query and when
2:38:40 the index grow 100 times the latency increases less than 2 times.
2:38:50 And so the change feed we have two ways,
2:38:53 the traditional way and the agent memory separate way.
2:38:56 The traditional is where the agent A will
2:38:59 be sent some message queue and Q delivers To Agent B 200 to 500 millisecond end
2:39:08 times handoff plus the cost of the queue service.
2:39:12 With the agent memory fabric that we have
2:39:15 just had the agent A will write Cosmos DB
2:39:19 change feed triggers agent pay in real time
2:39:23 and no queue to deploy monitor or to pay for.
2:39:29 So the database is the message us Here.
2:39:37 The three patterns I use in production
2:39:40 are first of all the fan out coordination.
2:39:44 So one agent will write many react and partition by the task ID.
2:39:55 The button 2 is that state machine translation where the status
2:40:02 moves pending to planning to execution to validating to complete.
2:40:09 So we have all the status like in front of us.
2:40:14 Change feed fires on every transition.
2:40:18 You get a full audit tree for free.
2:40:21 Button three is competing the consumers.
2:40:25 Multiple agent instances should work from one partition the least procedure
2:40:31 will rebalances automatically when you scale from 1 to 100 instances.
2:40:42 The problem was that agent A and C post read the same decode,
2:40:48 post, write back, last write, win and the first write is silently lost.
2:40:55 That's why I have here the mistake concurrency so we can prevent the corruption.
2:41:02 And the solution is the ETag where every document have a version stamp.
2:41:08 If a document changes between you read and you write,
2:41:13 Cosmos DB will reject the write.
2:41:15 No distributed lock manager needed, just a streamline retry rule.
2:41:26 What about the conflict resolution?
2:41:28 So a 409 conflict is not an error, it's an information.
2:41:34 And it has three patterns.
2:41:37 First one, the last writer wins.
2:41:40 Discard your changes, reread, retry, use when refresh data.
2:41:46 Is here is always more correct.
2:41:50 Next is the merge and retice.
2:41:52 So fetch is a fresh state merge.
2:41:55 Both changes retry with the new ETag.
2:41:58 Use when the agents update different fields of the same documents.
2:42:03 The pattern three is about partition and avoid.
2:42:08 So design the partition keys so agents never write at the same document.
2:42:14 The best conflict resolution is no conflicts at all.
2:42:22 So here's the full architecture we have.
2:42:27 Before we had like four systems,
2:42:31 four pills, four filler points, zero coordination.
2:42:36 After that we have one database, one perl,
2:42:40 one SDK, one dashboard, sync capabilities and only one fell.
2:42:51 So the lessons that I made that I learned the hard way are listed here.
2:42:57 These are the four things I got wrong.
2:43:00 First one is a partition key.
2:43:02 The partition key design done last, not first.
2:43:05 It caused a painful live migration.
2:43:08 The right key was/, task ID,
2:43:13 even distributed plus the change feed locality design itself.
2:43:20 Number two is that I killed the vector index after loading the data.
2:43:25 That causes a 90 minute window of hundred percent cache messes.
2:43:31 Build the index before loading the data.
2:43:34 Number three is the compliance team asked for an audit trail change.
2:43:41 The feed was already captured every state transition for free.
2:43:45 It saves reproduction incidents through the replay.
2:43:49 And number four is that start with the partition and avoid,
2:43:55 not the merge and retry.
2:43:58 The best conflict resolution is no conflict.
2:44:06 So for you now what to call Monday morning.
2:44:09 Four steps.
2:44:11 Create a Cosmos DB serverless account five minutes.
2:44:16 Then you'll create a container with a disk
2:44:19 in and vector index on/embedding 2 minutes.
2:44:26 You need then to store the LLM Responses with embeddings.
2:44:31 That is my existing code with plus let's say 10 other lines.
2:44:38 And before every LLM call, run the vector distance query first.
2:44:46 You can have semantic caching in production by the end of the Monday.
2:44:55 But why this pattern lens is because we can see
2:44:59 that McKinsey Global Institute published a Report in November 2020.
2:45:04 Five says that AI powered agents
2:45:08 and Reports could generate $2.9 trillion in U.S.
2:45:13 economic value per year.
2:45:17 So the winners will not be the teams that build the most agents,
2:45:22 there will be the teams that run them efficiently.
2:45:26 This is what the agent memory fabric gives you.
2:45:35 Thank you all the code is open source.
2:45:38 You will find this at this repo.
2:45:40 I'll be around.
2:45:42 If you have any question please open an issue on the repo and let me know.
2:45:50 Thank you.
2:45:54 Your favorite apps light up your world.
2:45:57 Azure Cosmos DB powers the spark behind their shine
2:46:00 with massively scalable data and built in AI capabilities.
2:46:06 Intelligence on the move.
2:46:07 Audi uses Azure Cosmos DB to fuel
2:46:10 AI assistants that deliver answers at highway speed.
2:46:15 Creativity that connects with Azure Cosmos DB.
2:46:19 Adobe helps brands craft personalized experiences with real time relevance.
2:46:25 Color reimagined built on Azure Cosmos DB.
2:46:29 Pantone's palette generator turns inspiration into expression.
2:46:34 Goal.
2:46:36 Premier League uses Azure Cosmos DB to keep fans close
2:46:39 to the pitch with insights drawn from match day model moments, talking cars.
2:46:45 The future of in car infotainment starts
2:46:48 with TomTom and Azure Cosmos DB answers uninterrupted.
2:46:54 Together.
2:46:54 Azure Cosmos DB and ChatGPT, the world's number one AI app help 800
2:46:59 million users keep work moving with near zero downtime.
2:47:03 From connected transportation to creative exploration and so much more.
2:47:08 Azure Cosmos DB powers the AI apps that move the world forward.
2:47:18 Hey folks, welcome to a session about Azure DocumentDB.
2:47:22 We have a really exciting session for you where I'm going to show
2:47:24 you first what and how things work under the hood for Azure DocumentDB.
2:47:29 What are the key functionalities and how you can modernize
2:47:33 your MongoDB compatible workloads that are on prem or on the cloud,
2:47:37 with just a few clicks.
2:47:39 So before we get into the slides and the content,
2:47:42 let's go over a brief demo and I'm going
2:47:44 to show you how everything works behind the scenes.
2:47:46 So over here I have a website where you place orders.
2:47:52 And our website is called Contoso Retail.
2:47:55 Think about it like a website like Walmart
2:47:56 or seven eleven where they have an online database,
2:48:00 where they sync up everything with their local instance and every
2:48:04 local store would have their own on prem DocumentDB server.
2:48:09 High level, to brief it up we have
2:48:12 a centralized database which connects all the local
2:48:16 stores that are in person with all
2:48:19 the local database for DocumentDB for these stores,
2:48:22 that are available out there today.
2:48:26 Over here we have the global database.
2:48:28 Let's also look at other databases that we have,
2:48:31 we have one store in Seattle and then we have a second store
2:48:34 in Chicago and it's all connected to our global database as, I mentioned before.
2:48:40 Now let's go through our Seattle store
2:48:42 and see what orders have been placed so far.
2:48:46 The last order that we had was a customer demo for like $249.
2:48:50 And the status still is still pending.
2:48:53 You can even add your own orders right now.
2:48:55 Let me briefly see how it works.
2:48:58 We place an order and the last order was just populated on the screen.
2:49:03 Everything is going to be synced up to a global database,
2:49:06 which is the Azure DocumentDB database.
2:49:10 If I hit refresh, the same order has
2:49:12 been popped up over here with the same numbers, and the same amount.
2:49:17 But let's say for some reason your local databases are not working today,
2:49:22 your WI fi, is off, or there's some kind of disaster that's happening.
2:49:28 We're going to try to replicate that in this demo as well.
2:49:31 I'm going to click on simulate Outage.
2:49:34 Everything that's connected to your local database
2:49:36 to your online database has been cut off.
2:49:39 There is no sync between the two.
2:49:42 If I go back to my Seattle database and add another order.
2:49:46 Let's see, let's add four items to the list and, place an order.
2:49:52 As you can see, we have an order Created
2:49:55 with order ID as last four digits as 5471.
2:50:01 If I hit refresh, it's right here, 5471.
2:50:04 But let me go back to my global database and see it has been synced up or not.
2:50:13 Let's give it a few more seconds.
2:50:16 Now we are connected to a global database.
2:50:18 Let me hit refresh.
2:50:21 As you can see, the last order ID is 5470.
2:50:26 Now let's take a look at what goes on behind
2:50:29 our database to make sure that we actually track,
2:50:34 this order in our local database.
2:50:37 I have, three sources over here.
2:50:39 First is my Azure DocumentDB instance, which lives on Azure.
2:50:43 Then I have a local DocumentDB instance which is, I've named that Chicago.
2:50:47 And then we have Seattle.
2:50:48 If I go to Orders and Documents and search for my last order,
2:50:55 which was 5471, you can see it has popped up.
2:50:59 It has all the items that I just placed an order for.
2:51:03 They are titled by SKU names.
2:51:05 It's still here.
2:51:06 But if I do the same thing, with my Azure database collection and if I go
2:51:12 to JSON view and search for 5471, nothing shows up.
2:51:18 We don't have anything in our global database just yet.
2:51:23 I'm going to cancel the outage and click to restore.
2:51:28 The disaster mode is off and the connection has been restored.
2:51:32 It knows all the changes that are in the pipeline
2:51:35 are going to be flush to our global database.
2:51:38 As you see it just mentioned,
2:51:39 auto sync 1 new order from Seattle store to global.
2:51:44 It's going to take a few minutes for it to pop
2:51:46 up on the recent orders and we can take a look.
2:51:50 The reason this is really crucial for you is because Walmart as a store,
2:51:53 at any store that are out there you have
2:51:56 local stores that are running on a 247 basis.
2:52:00 You cannot just rely on the cloud at every second.
2:52:03 Everything like we know some things could go wrong.
2:52:06 Customers if they want to buy something, we cannot be like hey,
2:52:09 the Internet is down, we cannot process a transaction.
2:52:13 So that's why having a local database is really important for these stores.
2:52:17 So as you can see it just flushed one new change to the global database.
2:52:21 Let me hit refresh again and you can see the order
2:52:25 id 5471 has been popped up on our global database.
2:52:29 And the amount everything is like searched and filled up over here.
2:52:34 So this is really crucial.
2:52:37 Since we are in truly open source database,
2:52:39 you can run DocumentDB on your local laptop computer wherever or even.
2:52:44 You can run it on aws, GCP or Azure.
2:52:48 It makes it so flexible for you to just
2:52:50 manage your workload based on what your needs are.
2:52:53 You're not tied into one vendor or one service at a time.
2:52:58 Also I want to show you that there is a bidirectional sync
2:53:01 between global and from local to global as well as global to local.
2:53:06 Let's say in this scenario the Contoso retail store,
2:53:09 they want to sync up everything once a day,
2:53:12 everything that's happened in 24 hours
2:53:14 of a local store to their headquarter database.
2:53:17 So they can do analytics,
2:53:19 do reviews and create dashboards on how many orders were placed.
2:53:23 We allow that as well.
2:53:24 Let's say like global database makes an update to the product sku,
2:53:30 or a product pricing has been changed.
2:53:33 Let's say the wireless headphone charge price goes from 249 to like 200.
2:53:38 They're running a sale for Thanksgiving or something like that.
2:53:41 So we can pull our data from a global database and it's going
2:53:44 to sync everything up with the local database that we have out there.
2:53:50 Right now we only have three databases stored.
2:53:52 We have one Chicago location, Seattle location and a global database.
2:53:56 But you can have multiple stores and have
2:53:58 a bidirectional sync between all these stores.
2:54:01 If something goes wrong in one of the local stores,
2:54:04 don't worry, your changes are not going to be gone forever.
2:54:07 They're just going to be queued up and they're going to be stored using change
2:54:11 streams later once the connection has been
2:54:13 restored and the disaster recovery has been fully.
2:54:17 You can also track what's going on for all these stores in our Azure service,
2:54:20 at the same time as well.
2:54:23 Now let's take a look at how our data is being
2:54:26 laid out behind the scenes using our own VS code extension.
2:54:29 DocumentDB for VS code extension let's go and take a look at how
2:54:33 our data is laid out for our online
2:54:36 database which is the Azure DocumentDB database.
2:54:39 Our local instances of DocumentDB which is
2:54:41 the Chicago store and the Seattle store.
2:54:45 Let's go through our orders and you can see we
2:54:48 have all the information necessary needed for our online headquarter database.
2:54:55 So if I click on the view for a specific document or a specific data point,
2:55:00 you can see it has all the information necessary
2:55:02 that an audience that a headquarter person would need.
2:55:06 Like where was the data synced from?
2:55:08 It said it was synced from the Seattle store at time April 16 at 7
2:55:14 30pm it also has all the information of what the status of that information is,
2:55:18 what was the total amount that was collected and what
2:55:21 was the items that were added to their sku.
2:55:25 Now let's go and take a look at the performance for our online database
2:55:30 and we want to make sure that the performance is optimal because let's say we
2:55:35 have an online store like Walmart where people want to search up what products
2:55:39 they are looking for and if it takes forever for them to load the website,
2:55:43 load the product they're looking for, it's
2:55:45 just going to be a bad user experience.
2:55:47 So now let's take a look at the orders and see how
2:55:51 long it takes to find a certain item that you're looking for.
2:55:55 So you can see we have multiple fields over here.
2:55:57 One of the fields is status.
2:55:59 So let's try to filter out based on the status of an order.
2:56:06 So if I just do status and shift and hit
2:56:11 run it's going to give me all the data points.
2:56:14 All the documents that have status equals shipped.
2:56:17 Now I want to take a look at the query performance for this exact same query.
2:56:22 As you can see it took like 3.5 milliseconds to examine over 15,000
2:56:28 documents to return only 2866 document the ratio for examine to return ratio
2:56:34 is 5 to 1 which is not good because you don't want your database
2:56:38 to do a full collection when
2:56:39 retrieving information for what they're looking for.
2:56:44 In this case it's only 15,000 documents but a store because Walmart there are
2:56:48 going to be millions and millions of millions
2:56:49 of documents or orders in the database.
2:56:53 If we scale it up to those levels it's just going
2:56:55 to be forever for them to get the information they are looking for.
2:56:58 So that's why we built something which is very useful using AI,
2:57:02 which is an AI powered performance insights.
2:57:06 Well how it works is it already has
2:57:08 all the information about your cluster, your database,
2:57:10 how your indexes are stored,
2:57:12 how your data is laid out and the performance of your database.
2:57:16 So it takes all that information and we
2:57:18 have a bunch of skills that MD built right
2:57:20 into it which gives you the recommendation of what
2:57:23 we think is going to be optimal for you.
2:57:26 Before we even give you the recommendation it
2:57:27 tells you what exactly is wrong with your query.
2:57:30 Over here you can see the performance summary said
2:57:32 the query performance is poor because you basically did
2:57:35 a full collection scan and which is inefficient for filtering
2:57:40 on the status field that we are looking for.
2:57:43 You don't want to do a full
2:57:44 collection scan when you're working with your database.
2:57:47 So and then it's going to give you a one
2:57:50 click recommendation solution where all it says is hey,
2:57:53 just create an index on the status field.
2:57:56 And it gives us, we have the button for it.
2:57:58 Just click on Create index and it's just going to create the index for you.
2:58:02 So as you can see it took like two
2:58:04 seconds but the index has been created on status field.
2:58:09 We are looking for.
2:58:10 So let me click over here, go to indexes and just hit refresh.
2:58:15 We can see our status index has been created.
2:58:17 Now let's run the exact same query to see
2:58:20 the performance difference with and without an index.
2:58:26 It's 1.74 milliseconds so we actually reduced it by exactly half that time.
2:58:33 We only examined 2,866 documents.
2:58:36 Return 2860 is the document that we are looking for.
2:58:41 This is how an index was actually optimized your performance.
2:58:47 You don't have to read any documentation or find the best indexing policies.
2:58:51 That's why we brought everything together using AI to make sure
2:58:54 your database performance is the most optimal you are looking for.
2:58:59 Not just this only works on an Azure DocumentDB service.
2:59:02 It works on any local DocumentDB service.
2:59:05 Or any Mongo compatible database that's running out there.
2:59:10 So let's do the same thing over here and let's do,
2:59:12 I'm going to open up a Seattle store and filter out by status
2:59:16 equals shift and run a fine query and it's going to give us,
2:59:24 since we don't have an index on, since
2:59:28 we don't have an index on the status field,
2:59:30 it's just going to examine all the documents that we
2:59:34 have to give us the response that you're looking for.
2:59:39 So AI performance also works on your local
2:59:42 instance as well as an Azure DocumentDB instance.
2:59:45 Now since we look at the demo, everything worked well.
2:59:48 Now let me give you a brief overview of what Azure DocumentDB is
2:59:51 or what is DocumentDB and why we even built it in the first place.
2:59:55 So you must be already familiar with MongoDB databases.
3:00:00 It's the most popular non relational database that is out there.
3:00:04 People allow the specifically people it allows to specifically
3:00:09 doc to specify documents in a JSON format.
3:00:12 And JSON fits naturally into many web developers code bases with the popularity
3:00:17 of JavaScript and it is very popular
3:00:20 to build solutions like web and mobile apps,
3:00:23 AI and rag applications and product catalog and personalization.
3:00:27 Personalization.
3:00:28 So what is DocumentDB?
3:00:30 It is an open source MongoDB compatible
3:00:33 database built for flexibility, scale and AI.
3:00:36 It is now part of the Linux foundation
3:00:38 and it has the industry support from Azure,
3:00:41 aws, gcp, Snowflake Cockroach Labs and so many more.
3:00:47 It also gives us ability to run your DocumentDB
3:00:50 instance on AWS GCP on Azure the same
3:00:53 time where you can replicate information across all
3:00:56 the clouds so you're not vendor lock in.
3:00:59 And also you can run it on prem.
3:01:02 As I showed in the demo few seconds ago,
3:01:05 the mission of DocumentDB is it's built on principles of transparency,
3:01:10 freedom and standardization visibility.
3:01:13 We want to ensure developers have
3:01:15 full visibility in the underlying architecture.
3:01:18 That's why with the MIT license users have complete freedom
3:01:22 to use the project as they please with no restrictions.
3:01:25 It is an open standard.
3:01:26 The goal is to create an Open Standard
3:01:28 for DocumentDB databases for a universally accepted implementation standard.
3:01:33 All right, so based on DocumentDB's popularity we thought
3:01:37 about let's also build an Azure first service built
3:01:40 on those open source DocumentDB and we call it
3:01:44 Azure DocumentDB With mongodb compatibility it is built innovative.
3:01:48 You can build innovative apps with truly open
3:01:51 source 99.03 MongoDB compatible document database in any environment.
3:01:55 You can also deploy in Azure for the best
3:01:57 Enterprise experience with hybrid and multi cloud capabilities.
3:02:02 It has AI driven AI built right into it.
3:02:05 It is a built in no cost vector search powered by DiskANN.
3:02:10 So you can combine your vector search queries with and your full
3:02:13 text queries with hybrid search right into the database.
3:02:17 Think about one database for all kinds of your search workloads.
3:02:20 You don't need to ETL your database out
3:02:23 of operational database to vector database like Pinecone or Vivid,
3:02:28 something like that.
3:02:29 Everything lives under one database so you can run all your queries together.
3:02:33 It is enterprise ready.
3:02:35 We have authentication with Entra ID.
3:02:37 You can get.
3:02:37 You get all the perks of being an Azure service.
3:02:40 We also give you up to 99.99995 availability SLA across the full service stack.
3:02:47 It also is cost effective.
3:02:48 It reduces the total cost of ownership
3:02:51 by scaling compute and storage independently.
3:02:54 And also 35 days of free backup and 24,7 support included in the price.
3:02:59 So there is no extra cost for licensing or support or even backup.
3:03:04 The price you see on the website is the price you pay every month.
3:03:07 It is also only based on your compute
3:03:09 and the storage that is required for your workload.
3:03:13 There are a lot of customers that are actually using Azure DocumentDB today.
3:03:17 We have AB and UBS, KPMG, UnitedHealthcare,
3:03:21 a lot of customers for different workloads.
3:03:25 Some of them like EY and KPMG are using DocumentDB
3:03:28 for their vector search workloads for the AI native applications.
3:03:33 One of the key differentiator for Azure DocumentDB is its
3:03:36 hybrid and multi cloud freedom like the demo I showed before.
3:03:40 You can run DocumentDB on any service that is out there,
3:03:44 local, aws, GCP and Azure at the same time.
3:03:47 And there's about unidirectional replication between all these services because
3:03:51 it runs on the same engine that powers them all.
3:03:54 And Also with Azure DocumentDB you can lower your monthly cost by up to 45%.
3:04:01 Also we give reserve instances.
3:04:03 If you commit for a year we give you up to 40% discount and if you
3:04:06 come in for three years we give up
3:04:08 to 60% discount on your MongoDB compatible database.
3:04:13 You can save up to 50% by switching to Azure DocumentDB
3:04:16 and it gives you the best perks that are out there.
3:04:19 We Give up to 32 terabytes per shard free point
3:04:23 in time backup and restore and no additional support contracts or licensing.
3:04:27 I'm going to emphasize this again.
3:04:29 The price you see on the website is the price you pay at the end.
3:04:33 There Are no surprises, no additional contracts,
3:04:36 licensing or any other fees that are associated with Azure DocumentDB.
3:04:39 But wait, there's more, there's more discounts.
3:04:43 We also give up to additional 20% customer incentive,
3:04:47 discounts on top of 1% and 3% RI.
3:04:51 If you're interested, feel free to reach out to us@DocumentDB@microsoft.com
3:04:54 and today we are introducing Premium SSD V2.
3:05:00 With Premium SSD V2 it gives you the maximum iops
3:05:04 and throughput for every disk selected at no additional fees.
3:05:07 So you can do 80,000 IOPS and around 1200 milliseconds per disk selected so
3:05:14 you can upgrade your speed but this the same price that you paid today.
3:05:19 We also have the LLM based Index Advisor
3:05:21 for DocumentDB which I demoed a few seconds ago.
3:05:24 So make sure you get the best optimum performance from your database,
3:05:28 wherever you're running it from.
3:05:30 And say goodbye to manual troubleshooting with just one click.
3:05:34 AI Optimization Then we also bringing advanced
3:05:36 full text search capabilities in Azure DocumentDB.
3:05:39 Think about Elasticsearch workloads that people usually run with the database.
3:05:44 We are bringing all those functionalities
3:05:45 built right in directly into Azure DocumentDB.
3:05:48 So no more paying for third party services for your search workloads.
3:05:52 Everything is built right into your database.
3:05:54 Some of the functionalities we're bringing in the next
3:05:57 month or so is going to be fuzzy.
3:05:59 Search proximity search,
3:06:01 multi language support like Chinese and Thai BM25 Ranking and rank Fusion.
3:06:05 So with Rank Fusion you can run
3:06:07 your hybrid search workloads with just one query.
3:06:10 So you combine your semantic search
3:06:12 and your full text search and use Rank Fusion
3:06:15 to find the best optimal solution best optimal
3:06:18 documents in the database for your exact query.
3:06:23 And in the future we're going to bring on multi
3:06:25 field indexing analyzers and tokenizers and tokenizers filters and so on.
3:06:30 And did I mention we have the best in class vector search.
3:06:33 We have performed significantly better than
3:06:36 any other MongoDB compatible database out there.
3:06:38 There are so many benchmarks that have
3:06:40 been reported about the performance with latency, lower latency and higher rps.
3:06:45 We also have Microsoft AI built in vector indexing which is the disk scan
3:06:51 and technology which supports richer vector search
3:06:54 capabilities so you get high accuracy and speed.
3:06:56 You can scale up to millions of vectors
3:06:58 embedding with high accuracy and fast speed.
3:07:01 We support up to 16,000 dimensions for context.
3:07:05 The largest embedding model that is out there
3:07:09 is 4096 dimensions but we support up to 16,000.
3:07:12 So let's say in the future, OpenAI, Anthropic or Gemini comes up with their own
3:07:17 model which has higher embedding dimensions.
3:07:19 We still support that in Azure DocumentDB.
3:07:21 No more thinking about what would happen.
3:07:24 So we are compatible with all the models
3:07:26 that are being released in the world today.
3:07:31 One other thing about Azure DocumentDB is you can do vertical
3:07:33 scaling as well as horizontal scaling along with the storage scaling.
3:07:38 So vertical scaling involves increasing the resources like vcores,
3:07:41 virtual cores and RAM of your existing nodes or shards.
3:07:44 Easily horizontal scaling means you are
3:07:47 distributing your data across multiple nodes,
3:07:50 enabling support for larger datasets and higher throughput.
3:07:53 You can also instantly independently increase
3:07:56 your disk size which without altering the compute
3:07:59 tier which gives you providing more flexibility
3:08:04 if your data grows in the future.
3:08:06 Let's go into the scale up process in much more detail.
3:08:10 Since Azure DocumentDB is built on scaling up process,
3:08:13 we support scaling compute which is vcores and RAM and storage iops this size
3:08:18 independently you can increase or decrease your computer
3:08:21 storage without service downtime or application changes.
3:08:25 We have larger disk up to 32 terabytes and 64 terabytes
3:08:29 per shard and we allow up to 10 shards at a time.
3:08:33 The architecture supports smaller compute
3:08:35 larger disks depending on your workload.
3:08:38 We also have horizontal scale out
3:08:40 when scaling vertically is no longer sufficient.
3:08:42 You want to scale horizontally by increasing the fuzzy physical shard count.
3:08:48 Azure DocumentDB manages logical physical shard mapping and the logical
3:08:52 shards are mapped into physical shards and requests are routed automatically.
3:08:56 Rebalances is handled behind the scenes.
3:08:59 We do everything for you and we go up to 10 shards at a time.
3:09:03 We are the only MongoDB compatible database with instant autoscale.
3:09:08 How it works is ensuring there's a peak demand.
3:09:10 We are always sure that your database is running at the same time.
3:09:14 Let's say there's a spike,
3:09:16 during a certain time of the month or a certain time of the year
3:09:19 where you see a huge spike from a user from your user using your website.
3:09:24 So instead of manually scaling up like the month
3:09:27 before to make sure that you meet the user demands,
3:09:30 we instantly auto scale for you.
3:09:32 And if we see the demand has slowed down,
3:09:35 we can manually lower the computer based on your workloads so your application
3:09:41 is still running without that unexpected need
3:09:44 that is going to crash the website.
3:09:46 We ensure that all our users have
3:09:48 instant autoscale so customers will choose this right
3:09:52 here for their workloads and over under
3:09:54 provisioning and reduce the escalations and budgeting.
3:09:58 Since we are a first party Azure service
3:10:00 you get all the authentication and access management,
3:10:02 data protection and network security built right in for no additional cost.
3:10:08 We have database security in Azure DocumentDB with in transit security,
3:10:12 Data Encryption and transit and at rest security.
3:10:15 So with a first party Azure service you
3:10:18 get all the Azure requirements that you're looking for.
3:10:20 And we also made sure migration is
3:10:23 seamless for whatever workflow you're looking at.
3:10:25 We offer three types of migration.
3:10:26 First we have the offline migration.
3:10:29 So think about it like a snapshot bulk copy from the source to the target.
3:10:35 If new data has been edited after we start an offline migration,
3:10:40 it's not going to be copied over.
3:10:41 But for that exact reason we have online migration.
3:10:44 But apart from the bulk start copy activities done in offline migration,
3:10:48 a change stream monitors all the addition,
3:10:51 updates and deletion that has happened.
3:10:53 Once the bulk copy start has been worked on, it remembers all that information.
3:10:58 Once that information has been migrated to Azure Document db,
3:11:01 it updates and deletes whatever the updates
3:11:03 have been done when the migration has started.
3:11:05 You can also cut over at any time when performing
3:11:08 an online migration to ensure it remains active until manually finalized.
3:11:14 Using cutover to prevent data loss and only proceed with cutovers
3:11:18 when the replication gap has been eliminated across all your collection.
3:11:24 Choose the best way you want to migrate your data
3:11:27 from your local or any other MongoDB compatible database to Azure DocumentDB.
3:11:32 We have an excellent tool for this where we have a VS
3:11:35 code extension where you just have to provide the connection string
3:11:39 from your source to the destination which is going to be
3:11:42 an Azure DocumentDB and you could
3:11:43 just migrate your data without any interruption.
3:11:47 Since we are MongoDB compatible database up to 99.03%
3:11:51 we will tell you if certain changes are required,
3:11:53 which I don't think is going to happen, but we give you an assessment,
3:11:57 a pre migration assessment which tells you what are at risk,
3:12:01 what needs to be changed before you even start your migration.
3:12:04 So you can have all the information you need before migrating
3:12:08 your huge workloads from on PREM or any other provider to Azure DocumentDB.
3:12:15 So since you have, let's say you are running through migration and you
3:12:18 have an issue migrating your data to the cloud to Azure DocumentDB.
3:12:22 We provide free migration with Cloud Accelerate Factory.
3:12:25 You can also reach out to the product team directly by just emailing
3:12:29 us@cbd migrationsupport@microsoft.com and one of the people
3:12:33 working on Azure DocumentDB will definitely reach
3:12:35 out and make sure your migration is seamless and there are no other issues
3:12:39 or data loss while you're migrating your data from On Prem to the cloud.
3:12:43 So we give you all the tools that are necessary to migrate
3:12:45 from On Prem art from any other Mongo compatible database to Azure DocumentDB.
3:12:51 If you have any questions, feel free to reach out.
3:12:53 And I think you should get started with Azure DocumentDB with a free
3:12:56 tier where we give up to 64 gigabytes of storage for completely free.
3:13:02 No credit card required to sign up.
3:13:04 Just try out and see how you like it.
3:13:06 Or you can even try our local instance, run a local DocumentDB server,
3:13:10 using a Docker container or even checking out all the samples that we
3:13:14 have built for Azure DocumentDB or DocumentDB
3:13:16 by checking out the website DocumentDB.io/samples.
3:13:20 The demo that I showed over here is also going to be available over there.
3:13:24 So if you want to take a look at the code,
3:13:25 how it worked, all the information is provided over there.
3:13:28 But feel free to reach out to us if you have any questions or drop a question
3:13:32 in the chat on the YouTube comments and we
3:13:34 will make sure we're able to answer it.
3:13:36 Thank you.
3:13:40 My name is Johnny Jalife and I'm CTO with Tao Works.
3:13:48 We typically use Cosmos DB in the solutions we build for our customers.
3:13:52 As the operational backbone for cloud native applications.
3:13:55 It works really well for customer facing platforms
3:13:57 that need to handle real time data like user context,
3:14:01 configuration and workflow state.
3:14:03 We also use it a lot in distributed and event driven
3:14:06 systems where data is coming in from fast and changing constantly.
3:14:10 And that's where Cosmos DB really stands out.
3:14:18 The main benefits we've seen are speed,
3:14:20 flexibility and less operational overhead.
3:14:23 Cosmos gives you low latency performance,
3:14:25 a data model that adapts easily as requirements evolve
3:14:28 and the ability to scale without having to redesign the system.
3:14:31 And it's fully managed, teams can stay focused on billing and shipping
3:14:35 instead of spending time running the database.
3:14:43 The main outcomes have been faster delivery,
3:14:45 consistent performance and the ability to keep expanding without friction.
3:14:49 We are able to ship features quicker,
3:14:51 keep performance strong as usage grows and add
3:14:53 new use cases without having to rework the data.
3:14:57 At the end of the day it helps
3:14:58 our customers customers move faster and scale with confidence.
3:15:05 Welcome back and welcome back.
3:15:08 And all I can say right now about what we've seen so far.
3:15:11 Amaze, amaze, amaze.
3:15:13 That's right Jay.
3:15:14 We've seen how developers are building
3:15:16 across clouds and designing flexible systems.
3:15:19 Yeah.
3:15:20 And you know what?
3:15:20 We've heard back from some of our great, great members of our community,
3:15:25 who have left some things in the chat, Chenola said,
3:15:29 looks like there is a great future for DocumentDB.
3:15:32 You know there is Patty.
3:15:34 I know that there is.
3:15:35 We're doing a lot of great things
3:15:38 with DocumentDB and especially with open source DocumentDB.
3:15:42 So we're always available and if you want to learn more about it,
3:15:46 just visit DocumentDB.io.
3:15:48 Also we have a great comment from behind.
3:15:52 Building real time dashboards is a breeze
3:15:55 using change feeds and it really, really is.
3:15:58 Absolutely.
3:15:59 And we heard so much great stuff from Justine today about
3:16:02 the change feed and how it's really helping people with real time applications.
3:16:06 But now it's time to go into some real,
3:16:09 real deep content things that we know are so important to having
3:16:14 a well created efficient and cost
3:16:19 especially considered when we're talking about indexing.
3:16:21 Yeah, I couldn't have said it better Jay.
3:16:24 This is really where performance comes together in just an application.
3:16:29 Yeah.
3:16:30 And if you want to revisit these sessions later,
3:16:32 you'll find everything in the full playlist.
3:16:35 It's going to be@ah, aka aka.ms/cosmosconf26playlist.
3:16:39 It's in the chat right now.
3:16:41 Go ahead and keep it bookmarked for later.
3:16:45 Yeah, I know for me personally we're so busy hosting,
3:16:49 that's definitely something that I will check out afterwards.
3:16:51 Absolutely.
3:16:52 And if you have a minute please, we'd appreciate your feedback.
3:16:56 So if you want to share your thoughts, go to aka Ms.
3:17:02 aka.ms/cosmosconf2026survey.
3:17:04 Yeah, we want to know so we can keep
3:17:06 doing this show and making it great for you.
3:17:09 So up next we've got a really great one.
3:17:11 It's called Querying and Indexing in Azure Cosmos DB.
3:17:14 The Complete Guide.
3:17:16 Yeah, here's Dr.
3:17:17 James Codella.
3:17:20 Hey everyone.
3:17:21 Welcome to Querying and Indexing in Azure Cosmos DB.
3:17:24 I'm James Codella, Product manager for Querying AI in Cosmos DB.
3:17:28 Let's get started.
3:17:30 So in this session we're going to imagine a multi tenant eretail platform where
3:17:34 we have customers that can shop across
3:17:36 a portfolio of different products and outdoor brands.
3:17:40 We're going to take a look at some common query workflows that such
3:17:44 a platform might use in order
3:17:45 to help their customers search for different products.
3:17:49 Now we're going to run the same queries against the same set of data,
3:17:52 but we're going to compare it in two different ways.
3:17:55 First is we have a container called
3:17:57 Products which has no indexing policy whatsoever.
3:18:01 And we're going to use some unoptimized queries to run These workflows.
3:18:05 Next we're going to run it the same
3:18:08 queries against a optimized collection that has
3:18:11 the best indexing policies and using the best
3:18:14 practices for creating queries in Cosmos DB.
3:18:17 And we're going to compare the performance and cost differences.
3:18:23 So what are we going to cover in this session?
3:18:26 We'll start briefly on partition keys, what they are and how to use them.
3:18:30 Then we'll dive into basic queries, queries that use a, composite index or more
3:18:35 complex patterns handling array data in queries,
3:18:39 geospatial data, computed properties,
3:18:43 and finally end with our advanced search capabilities including vector,
3:18:47 full text and hybrid search.
3:18:50 Plus we'll take a look at indexing metrics and query execution metrics
3:18:54 to help us make sure that we're preparing our queries are performing optimally.
3:19:00 So in all of these different scenarios we're going to see a few common patterns.
3:19:04 First is we'll see an example of a query
3:19:06 or two that is going to power the scenario.
3:19:09 Next we'll see the RU charge.
3:19:11 So how many request units that query execution cost.
3:19:14 We'll see the latency or the time it took
3:19:17 for the Cosmos DB container to execute that query.
3:19:20 And then we'll see the indexing metrics.
3:19:22 So this will show us which indexes were utilized by the query
3:19:26 and what potential indexes could have added benefit to the query if they were,
3:19:30 if they existed in the indexing policy.
3:19:33 So maybe better performance or better cost or both.
3:19:36 We'll also see query execution metrics which
3:19:38 show us exactly component wise where the query
3:19:41 took the most time and spent the most amount of work in the entire process.
3:19:47 Now, what does our data look like?
3:19:49 Well, imagine that we have a catalog of different products
3:19:52 for our E retail shop and we have different products as data items.
3:19:56 So you may have an id, a tenant or a store id.
3:19:59 So this is the Ridge Supply Store, a sku, a name for the product.
3:20:04 So in this case it's a camp stove and a bunch of other
3:20:07 metadata that you'd expect to see in a product catalog along with a description,
3:20:10 tags, some date timestamps as well as GeoJSON data that tells
3:20:16 us exactly where this product is located at that specific store.
3:20:20 As well as a embedding field that contains a vector
3:20:23 embedding that describes the description and name of the product.
3:20:26 And we'll use this later on for a vector similarity search.
3:20:33 So let's start with partitioning.
3:20:35 Partitioning Azure Cosmos DB determines how data is
3:20:39 distributed in routed Partition key provides a hint
3:20:42 to Cosmos DB to which logical partition
3:20:45 the data should be written to or read from.
3:20:49 So in our scenario we imagine that we have three
3:20:51 retailers or three shops and we're going to partition on that.
3:20:54 So most of our eretail queries can be scoped to a single
3:20:58 retailer unless the user wants to search across different retailers.
3:21:02 Now in general a good pattern to use is
3:21:04 queries should really include the partition keys whenever possible.
3:21:08 This allows your queries to be routed to the exact logical
3:21:11 partition or small set of logical partitions where your data lives.
3:21:15 This helps you keep RU costs and latencies
3:21:17 low and avoids potentially costly cross partition queries.
3:21:23 Okay, and we'll see an example of that in our first couple of queries.
3:21:26 So let's start with the basic query operations.
3:21:28 Right, so when I say basic queries we're going to use equality filters.
3:21:32 So quality filters, range predicates, single property sorts or order buys,
3:21:37 all these in combination to create a very basic common query pattern.
3:21:42 So for an example in our E retail shop maybe we would need
3:21:45 to filter on a category or do some simple sorts like sorting products by price.
3:21:51 That's useful really every single sort of product grid
3:21:54 or carousel that you'd see in an E retail shop.
3:21:57 And it's important to note as we'll see a query that returns
3:22:00 the correct results is not the same as a query that scales.
3:22:03 So we'll want to make sure that we're using best practices
3:22:05 and the right indexing policy to make
3:22:07 sure that our query can execute performantly.
3:22:12 Okay, so I'm going to switch over here to my terminal
3:22:17 and the first thing I'm going to do is I
3:22:21 have some sample runner here where I'm going to run some
3:22:24 sample queries against my data set of about 10,000 different products.
3:22:28 And so I'll go ahead and run this project and we'll get started.
3:22:33 So we're going to see two scenarios.
3:22:34 We're going to see a basic filter query and a basic filter plus order by query.
3:22:40 So here we can see the example of the basic filter.
3:22:44 So we're selecting a few of the different property paths
3:22:47 of our data items and we're adding where clause and two conditions.
3:22:51 Right?
3:22:52 So we're going to scope it to a particular
3:22:53 tenant or shop and then we're going to make
3:22:55 sure that the category name is targeted
3:22:58 to the category that we're really interested in looking for.
3:23:01 So maybe this user really wants to look at camping tents and backpacking tents.
3:23:06 So we're going to press enter and run the first query.
3:23:09 And so now we're going to spin up our job
3:23:10 and we're going to go execute a query from my laptop,
3:23:13 hitting the Cosmos DB backend from my Cosmos
3:23:16 DB resource deployed on the west coast
3:23:17 and come all the way back here to New York and show me the results.
3:23:21 And so here I can see the results for my query and I can see,
3:23:23 you know, a couple different products that came up.
3:23:25 And as I scroll up here, I'm going to see a few things also.
3:23:29 Right, so here are the index metrics that we had talked about.
3:23:32 So this first query is running on our container
3:23:35 products container that has no indexing policy whatsoever.
3:23:38 So the index metrics says that it wasn't able to utilize any index,
3:23:43 which makes sense because we don't have an indexing policy.
3:23:45 But it was looking for potential indexes.
3:23:47 So it was looking for a category name and it was looking for a tenant id.
3:23:51 So these are both things that we filtered on in the query itself.
3:23:54 And the impact score is high.
3:23:56 So if we add these to our indexing policy like we're really
3:23:59 going to get much better performance and better RU cost characteristics as well.
3:24:04 Scrolling up we can see the query metrics.
3:24:06 So this tells me how many documents were retrieved,
3:24:09 what the size was of all those documents and other
3:24:12 metadata that helps me understand the total query expense,
3:24:14 execution time or where the query spent a lot of time in processing.
3:24:19 And then I also print out some summary metrics
3:24:21 like how many rows were returned in this query.
3:24:23 So 136 match the query filters and then the RU charge.
3:24:27 So 343, almost 344 RU.
3:24:30 So this is not a cheap query.
3:24:33 Okay, so let's go and run this on the optimized container now.
3:24:36 So we'll now run this on the container that has an indexing policy,
3:24:39 that will make this query more performant.
3:24:42 And we can see the same results come back.
3:24:45 And this time in the index metrics we see that it
3:24:47 was able to utilize the indexes that are now in the policy.
3:24:51 And you know there are no indexes that were
3:24:53 not in the policy that it was looking for.
3:24:54 So that's great.
3:24:55 So we're effectively using all the indexes that we
3:24:57 can and we can see that you know, RU charge is significantly less.
3:25:01 Actually let's do a side by side comparison.
3:25:03 So we can see that without the index it was about 240
3:25:06 milliseconds of execution time for this query
3:25:09 with the index about 38 milliseconds.
3:25:12 So more than a 6x improvement on latency and then more than
3:25:17 a 7x increase or sorry 7x improvements on RU charges for this query.
3:25:22 So really demonstrating the power of having proper index even just with a simple
3:25:27 filter with just a couple of query where clauses in the query.
3:25:33 Okay so now we're going to run a basic filter and an order by clause.
3:25:37 So we're going to go ahead and run this again using our.
3:25:40 NET SDK running sort of a default query.
3:25:42 And we get back like this error actually
3:25:45 wasn't able to operate because it was looking for an index to do the order
3:25:49 by on the container itself and we didn't have that there.
3:25:52 So let's go ahead and do the same query on the optimize path.
3:25:57 So an optimized collection where we actually do have the indexing policy.
3:26:01 It was able to leverage the indexes and actually it's telling us that I forgot
3:26:06 to add a composite index here but we'll see an example of that maybe next.
3:26:10 So we can see that it was able to execute.
3:26:12 It's actually relatively performant comes back
3:26:15 27 milliseconds of query execution time.
3:26:17 That's a pretty fast query.
3:26:21 All right, so now let's go back to our slides and we'll go to the next scenario.
3:26:26 So we saw basic query operations with filtering and order by clause.
3:26:30 Next what we'll do is look at composite indexes.
3:26:33 Composite indexes are really useful when I have
3:26:37 where clauses or filters that cover multiple different
3:26:40 properties and or an order by with multiple
3:26:44 different properties in the order by as well.
3:26:46 Right.
3:26:47 So adding complexity to my queries really makes
3:26:49 composite indexes a really useful pattern for me.
3:26:52 Right.
3:26:52 So maybe if I'm searching through my product catalog for you know,
3:26:56 the latest backpacking tents.
3:26:57 So I want to search for some keywords and I also want
3:26:59 to have an order by clause in there for the most recent items.
3:27:03 Or I want some you know price order by along with my query filters.
3:27:07 There's really great use cases for composite
3:27:10 indexes and you know the composite index needs
3:27:14 to match the order and sort direction in which
3:27:16 you're going to leverage it in the query.
3:27:18 So that, just keep that in mind.
3:27:19 And of course you have to add composite indexes.
3:27:21 It's not index by default in Cosmos db
3:27:24 Okay so let's clear our screen here and what we're going to do is we're going
3:27:32 to go run the demo scenario with composite indexes.
3:27:39 So let's spin this up and if you're curious,
3:27:41 this tool that I'm using in the terminal, it's going to be available on GitHub.
3:27:45 There's actually a link and a QR code at the end of this, this presentation.
3:27:49 So you know, don't need to bother taking notes or anything like that.
3:27:52 You have access to all the code at the end of this talk.
3:27:55 Okay, so we're going to test a, couple different scenarios here.
3:27:58 We're going to look at composite,
3:27:59 equality and range filters and then we're going to look
3:28:01 at composite filtering and composite sorting in the same query.
3:28:05 So in this example here we're going
3:28:07 to select some properties and we want to make
3:28:09 sure that we're scoping to a particular shop
3:28:11 and scoping to a particular category for those products.
3:28:15 And we want the price, you know, to be greater than you know, something.
3:28:20 Right?
3:28:21 So that something's going to be 100 so 100 US dollars, right?
3:28:25 And so let's go ahead and execute that query.
3:28:27 And again with all these scenarios,
3:28:28 we're first executing on the unoptimized collection with the collection without
3:28:32 any indexing policy and we can see that the results come back,
3:28:36 maybe it took some time for us to run, you know,
3:28:39 we can scroll up and see the metrics and it took
3:28:41 about 350 RUs to execute that query, which is not cheap.
3:28:44 And we'll go and run this on the optimized container and it's going to execute
3:28:48 much much faster and we're going to skip
3:28:50 directly to the side by side comparison.
3:28:52 Look, there's no, there's no contest here, right?
3:28:56 13x more than 13x improvement in server latency, so query execution latency,
3:29:01 more than 7, almost 8x improvement in RU charges.
3:29:05 It's no brainer that you want to use composite indexes,
3:29:07 for these sort of complex query filter scenarios.
3:29:10 Now when we compare in the next scenario using a composite filter and sorting,
3:29:17 we'll have sort of the same issue that we saw last time.
3:29:19 Right.
3:29:20 For composite indexes I really excuse me, for composite order bys,
3:29:24 I really need a composite index to be able to execute that effectively.
3:29:27 So we're not able to execute
3:29:29 this on our container that has no indexing policy whatsoever.
3:29:32 But when we run this on the container that does have the indexing policy,
3:29:35 not just the standard indexing policy but also our composite indexes,
3:29:39 we're able to realize that it's still incredibly performant, right?
3:29:44 16 seconds.
3:29:45 It took 16 seconds for this query to run on over 10,000 data items.
3:29:49 And it cost about 47 RUs.
3:29:51 So much, much cheaper than we would expect, for any of these operations,
3:29:56 for not using sort of an optimal
3:29:59 way and best practice of leveraging composite indexes.
3:30:04 Okay, so let's talk about the next area.
3:30:07 So we covered basic queries and composite indexes.
3:30:10 Now let's jump into arrays.
3:30:12 So Cosmos DB can index the content in an array.
3:30:17 And then you can use query functions like array contains or you can
3:30:21 reference certain positions of the arrays or you can even do these joins
3:30:25 like intra document joins and you can use exist subqueries and all sorts
3:30:29 of great functions to resolve from the index
3:30:32 instead of scanning every single document, for the contents of those arrays.
3:30:36 Right.
3:30:37 So for example, maybe my products have different tags like a waterproof
3:30:41 tag or a tag on what seasons that product can be used for.
3:30:46 And these tags can power, you know, filter navigation.
3:30:49 So imagine that a customer is searching
3:30:50 for products they want to filter by other countries,
3:30:52 categories or their characteristics in this product.
3:30:56 Important thing to know is that Cosmos DB is schema free.
3:31:00 Right?
3:31:00 It's essentially JSON documents.
3:31:02 So instead of joining data across different
3:31:03 entities like you would in a relational database,
3:31:06 joins occur within a single item.
3:31:07 So there's scope to that item.
3:31:09 You don't do cross item or cross container joins.
3:31:13 It's really meant to sort of bring up and interact
3:31:16 with the objects or items that are contained within an array.
3:31:20 Okay, so let's see an example of this.
3:31:23 So we will switch back here to our terminal
3:31:28 and then we're going to run the demo scenario for arrays.
3:31:34 Let's go ahead and load this up.
3:31:36 And so we're going to test two different scenarios.
3:31:38 So we'll run a scenario where we're leveraging the array contains properties,
3:31:42 we're searching an array to see if it
3:31:43 contains some specific value or object and then
3:31:46 we're going to see an example of join and how that can be used as well.
3:31:49 Great.
3:31:51 So in this example, maybe we're filtering on the tenant and we want
3:31:54 to make sure that the tags array contains a specific tag like waterproof.
3:32:00 So we're going to go ahead and run that without any index
3:32:02 whatsoever on our products collection and we'll see the results that come up.
3:32:08 And we see the results and of course you know,
3:32:11 we see that it was looking for indexes
3:32:12 that it couldn't find in the indexing policy
3:32:14 because there is no Indexing policy and we can see it took about 411 RUs to run.
3:32:20 So we'll go ahead and run this on the container with the index
3:32:24 defined and we'll see that the same results come up here.
3:32:29 And we'll see that this is actually a much, much cheaper and efficient query.
3:32:34 Right?
3:32:34 So we did use the indexes that we have defined and it cost you know,
3:32:40 significantly less ru.
3:32:41 So let's do a side by side comparison.
3:32:43 So over you know, 1.4x improvement in latency
3:32:46 and over 3x improvement in RU chart.
3:32:48 Right.
3:32:49 And these performance gains that you get from adding the index
3:32:52 only grow when the size of the data grows, right?
3:32:55 If I went, if I had instead of 10,000 plus documents,
3:32:57 if I had 100,000 or a million or 100 million documents, right,
3:33:01 you can imagine that the improvements that I'm seeing here would actually be
3:33:05 many fold more than what I'm getting just from like the small scenario.
3:33:11 Next let's see an example of join.
3:33:14 So we have select, have a couple properties like
3:33:15 our name and description and then also the tag, right?
3:33:19 So we're getting the tag, we're getting tags from the C tag.
3:33:23 So C tags, tags is the array of different tags for that product.
3:33:28 So we're scoping again to the tenant ID and we're looking for a particular tag.
3:33:31 So this time we're going to again look
3:33:33 for waterproof as a particular tag using a join syntax.
3:33:37 So again we're going to run without
3:33:38 the index and then once this executes we're going
3:33:42 to just go right ahead and run it
3:33:44 on the optimized collection that does have indexes that were,
3:33:47 you know that, that we chose specifically where that Actually
3:33:51 to be honest this is the standard indexing policy, right?
3:33:54 By default Cosmos DB indexes all the properties
3:33:56 by default for range and inequality inequality filters.
3:34:00 And so I actually don't need to add anything specific for array operations here.
3:34:05 The one in the queries that I'm showing you,
3:34:07 this is part of the default indexing policy.
3:34:09 And so we can see that without the index versus with the index.
3:34:12 So with the index you get over a 5x
3:34:14 improvement on latency and over 3x improvement on RU charge.
3:34:18 So again just demonstrating the power of having an index policy for my arrays.
3:34:29 And now what we're going to do is
3:34:30 we'll move on next to talk about computed properties.
3:34:33 So computed properties let you define derived
3:34:35 fields from query expressions on the container Properties,
3:34:38 So they're not materialized, at read, time or query time.
3:34:43 They're not computed at query time,
3:34:44 but they're materialized and indexed at break time.
3:34:47 Right.
3:34:47 And the properties aren't persisted in the items themselves,
3:34:50 which is really makes it really nice for efficient storage.
3:34:52 Right.
3:34:53 So imagine that in our product catalog, imagine that we're storing,
3:34:56 you know, the different category path in a sort of compound string.
3:35:00 Like maybe our string is camping and tents and backpack tents.
3:35:04 And what we want to do is maybe we want to filter on some,
3:35:07 some, you know, subset of these.
3:35:08 Right.
3:35:09 Some, some, some subset of, of, of these categories that are stored in a string.
3:35:13 So without a computer property,
3:35:14 maybe I'd write a query that separates these out at runtime in the query itself.
3:35:19 And that could lead to, you know, when I'm doing a lot of queries,
3:35:22 I'm just doing that operation over and over and over again in the query,
3:35:25 which is not the most efficient way.
3:35:27 So in a computer property, I can do this, I can write it once and then
3:35:30 I don't have to worry about it at query time.
3:35:32 And it's important to note computer properties.
3:35:35 Like you have to define these yourselves and they
3:35:36 need to be explicitly added to the indexing policy.
3:35:39 But once you do this, then you can reference them
3:35:41 just as you would any other property in, your document.
3:35:44 Except they're not property of your document, they're a computer property.
3:35:48 So let's see what that looks like.
3:35:52 So we'll go back to our terminal and we'll clear the screen,
3:35:58 we'll go ahead and run this computed properties workflow.
3:36:01 Okay, so a bunch of stuff's popping up here.
3:36:03 So we're going to look at a derived field lookup.
3:36:07 So, here's the query that we're going to run without computer properties.
3:36:10 We want to project a couple different properties of the products.
3:36:14 And then of course, we're going to scope to our, tenant or shop.
3:36:18 And then we have, in addition,
3:36:19 we have this additional part of the where clause where we're going
3:36:21 to look at the category name and we're going to take a substring,
3:36:26 and take the first, element of it.
3:36:28 So, right, we're going to trim, this as well.
3:36:31 And so what we're doing is we're basically getting
3:36:33 the first category that's listed in that string, right?
3:36:36 So we're going to call this the primary category.
3:36:38 So maybe the first category of this product is the most important,
3:36:42 the main category of this product.
3:36:43 Unfortunately, all the products categories are Maybe you know,
3:36:47 because of my data model or the way that my application is designed,
3:36:50 it's stored in a way that is in a string rather than a array.
3:36:55 And so maybe I want to search for a particular,
3:36:57 just what this primary category might be.
3:37:00 So without computer properties,
3:37:01 this is what my query would look like with computer properties.
3:37:05 I can actually just reference a computed property that I create
3:37:08 and I can reference it like any other property in my document.
3:37:11 C cprimarycategory.
3:37:14 And what I can see here is that when
3:37:15 I define a computer property from my collection,
3:37:17 I can give it a name and then basically I
3:37:20 can give it a query that defines a computer property.
3:37:24 Right?
3:37:24 So the logic that we had here that was previously
3:37:27 in our unoptimized query is now in a computed property.
3:37:31 So now this can compute it and store it
3:37:34 in the index and makes it really efficient for search and retrieval.
3:37:37 And I don't have to rewrite this in every single query.
3:37:40 I don't have to rerun this logic every single time, during query execution.
3:37:44 So it should make our queries more performant.
3:37:46 So let's go ahead and run this.
3:37:47 So we'll run without the computer properties.
3:37:50 We'll go ahead and execute again on our product container.
3:37:56 And again so we're going to actually,
3:37:58 so in this example we're actually going to show
3:38:00 the difference between a computer property and no computer property.
3:38:03 So we're still going to leverage the index in both of these scenarios.
3:38:06 But one time we'll be using,
3:38:08 in this first one we're not using computer properties.
3:38:11 And then, Right, and so we can see the characteristics of this.
3:38:15 Right, so we can see there's 378 RUs.
3:38:18 And then in this optimized version we're
3:38:20 going to be using the computed property itself.
3:38:22 So again the first query did not use computer property,
3:38:26 but it still used the indexing policies.
3:38:29 Now in this second run we're going to use the computer properties and the index
3:38:35 computer property as well and we'll look at the side by side comparison.
3:38:39 Okay, so the execution time of the query was 88 milliseconds.
3:38:44 88.77 milliseconds with computed property.
3:38:46 So over 3x performance improvement with computed property in terms of latency,
3:38:51 1.45x improvement from RU charge.
3:38:54 And again as your data set scales this benefit only increases.
3:38:58 So really, really interesting application of computer properties here.
3:39:02 Right.
3:39:02 It's a really great feature that if you're
3:39:04 doing this complex logic over and over again,
3:39:06 in your queries you can really simplify your workflow,
3:39:09 and get better performance out, of your queries by using the computer property.
3:39:17 All right, so what do we have next?
3:39:19 Geospatial.
3:39:20 Great.
3:39:20 Geospatial data enables you to do all sorts of cool filtering
3:39:24 or bounding or sorting by physical distance by having locations like longitude,
3:39:29 latitude locations, on Earth.
3:39:31 This allows you to use all sorts of complex
3:39:34 geospatial functions that we have in the cosmos.
3:39:36 DB Query syntax.
3:39:37 So in this example we're going to use the stdistance function.
3:39:40 So a query common example that you find in e retailers, right?
3:39:44 You are looking for products and you want to make sure that you maybe you can
3:39:48 find a store nearby that has this product in stock that you want to go pick up.
3:39:52 Right?
3:39:53 And again, you know, spatial data here needs to be valid GeoJSON.
3:39:57 And of course you need to add it to your indexing policy.
3:40:00 By default, custom, DB doesn't index geospatial data,
3:40:03 but it's really simple just to add it to your indexing policy wherever
3:40:06 you're going to store whatever property you're going to store your GeoJSON data.
3:40:10 Okay, now let's see this in practice in our demo code.
3:40:14 So first we're going to clear a terminal.
3:40:17 So we'll go ahead and do that.
3:40:22 Then we're going to go and run the geospatial scenario.
3:40:26 So what are we going to take a look at in this scenario?
3:40:29 Well, we're going to look at a couple of things.
3:40:32 So first we're going to do this distance filter.
3:40:34 So maybe in this scenario we want to make sure that the product
3:40:39 that the customer is looking for is in stock in a store that's close to them.
3:40:43 Right?
3:40:43 So we are going to have a couple things in our query, right?
3:40:48 So we're going to project the id, the name, the description and the distance,
3:40:52 or the geospatial distance as this distance meters, alias.
3:40:57 And then we're going to have a where clause on the tenant id.
3:41:00 So we're going to look for a specific, tenant or a specific shop.
3:41:04 And then we're going to, we want
3:41:05 to make sure that it's within this maximum distance.
3:41:08 So we want to make sure it's within, you know, say 50, kilometers, of our user.
3:41:14 And then we're in order by ST distance.
3:41:16 Right.
3:41:16 So what this is going to do is going
3:41:18 to rank the results in order by their geospatial distance,
3:41:23 from the user to the stores.
3:41:25 So let's go ahead and run this query
3:41:27 and we're going to run it again first without
3:41:29 the index and we can see some of the results
3:41:32 that come back and some of the characteristics.
3:41:34 But let's just skip this and we'll go ahead and run it
3:41:37 with the index and then we'll look at the side by side comparison.
3:41:41 So side by side comparison shows something really interesting.
3:41:43 Right?
3:41:45 Without the index we see about 89 millisecond of execution time on the server.
3:41:49 With the index 55.88 milliseconds.
3:41:52 So it's a 1.61x improvement with the index.
3:41:56 And then of course the RU charge as well,
3:41:58 we get a 2.34x improvement on RU charges.
3:42:02 Right.
3:42:03 And again this only increases, right?
3:42:05 The improvement level only increases as the scale of my scenario.
3:42:09 So the more data that I have, especially if I have really high throughput
3:42:13 account or account with lots of different partitions.
3:42:15 Right.
3:42:16 So we get a lot of benefit by using
3:42:18 the index and also by specifying the partition key,
3:42:23 in that where clause as well.
3:42:24 So now that we talked quickly about geospatial,
3:42:26 let's look at some of our search capabilities.
3:42:29 So vector, full text and hybrid search.
3:42:31 Let's start with vector search.
3:42:32 So vector search uses vector embedding and our vector
3:42:36 distance function combined with a vector index.
3:42:38 So in Cosmos DB we have multiple vector indices including disk ann,
3:42:41 which allows you to do semantic search with very
3:42:44 high accuracy and low latency really at any scale,
3:42:48 even up to billions of vectors.
3:42:50 So imagine that a shopper wants to search for something,
3:42:54 you know, a product and uses some, you know, text phrase, right?
3:42:58 Like lightweight tent for rainy hikes.
3:43:00 So the vector search will allow us to find semantically related products,
3:43:03 not just products that match these keywords or phrases and support a note.
3:43:09 So vector search is really good to add
3:43:11 a vector embedding policy and vector index.
3:43:13 So you could add like DiskANN index.
3:43:15 And if you want to do multi tenant search, like for example here,
3:43:18 if I wanted to search specifically
3:43:19 on a particular tenant or our particular shop,
3:43:23 I could shard our vector index to that, you know,
3:43:26 by the different shops or tenant IDs.
3:43:29 And so I could isolate my vector search just
3:43:31 to those shops and you can get really performant,
3:43:34 you know, tenant isolated search that way.
3:43:37 So let's see an example of vector search and practice.
3:43:40 Okay, so we'll clear screen again.
3:43:47 And we'll do a run of the Vector scenario.
3:43:52 So here our index.
3:43:54 Excuse me, here our query is very straightforward.
3:43:56 We're selecting the top five most similar items and we're going to project
3:44:00 the vector similarity score and then we'll also order by the vector distance.
3:44:04 Right?
3:44:05 So we'll order by the vector similarity score from most similar
3:44:08 to least some more and we'll run it with and without the index.
3:44:19 So we ran it with and without the index.
3:44:22 We'll skip right to the comparison and we can see the server time.
3:44:25 So almost 2024, almost 25x improvement in server side latency using
3:44:31 the vector index compared to not using a vector vector index.
3:44:34 And then the RU charge is also much much lower.
3:44:37 Right?
3:44:37 Not even 1615 RUs on vector search on over 10,000 documents.
3:44:44 So really, really performant using that to scan in vector index.
3:44:49 Going back to talk about full text search real quick.
3:44:52 So full text search uses different linguistic
3:44:55 analyzers so tokenization and stemming and functions like
3:44:58 full text contains and full text score
3:45:01 to match and rank documents by keyword relevance.
3:45:04 Right.
3:45:04 So I want to make sure that certain
3:45:07 products contain a certain phrase or keywords.
3:45:09 I can use full text contains if I want to search for Items using this BM25
3:45:13 scoring which is a measure of frequency of whether
3:45:17 these terms are in the documents or not.
3:45:20 I can do an order by rank full text score, clause as well.
3:45:24 So we'll see examples of that.
3:45:26 Right.
3:45:26 So an example here is when a customer wants to do a search
3:45:28 and you want to search by those text keywords or by those text phrases.
3:45:33 This is really powerful tool for enabling that.
3:45:35 Okay, so we'll run this scenario as well.
3:45:37 So we'll clear a terminal again and go ahead and run the full text demo.
3:45:45 So here's what we're going to test the full
3:45:47 text contains and full text scoring using BM25.
3:45:50 So we'll do a simple query here where we have we want
3:45:52 to make sure multiple terms are contained
3:45:54 within the products that we're looking for.
3:45:56 We want to make sure that it contains waterproof and conditions and windproof.
3:46:01 And so we're to run this without any index whatsoever.
3:46:04 And then we'll run this with a full text index
3:46:07 defined on that collection and we can see that you know,
3:46:11 it's 2.3x more performant and uses far fewer RUs.
3:46:15 And again the difference here the benefit
3:46:17 only increases as your data size scales.
3:46:20 Next we'll run the full text scoring.
3:46:22 Right.
3:46:23 So what we're going to do is we're going to search
3:46:24 for these items and we're going to order by their BM25 score.
3:46:28 So how relevant the products are by how often these keywords
3:46:32 appear in the descriptions or the names, of these items.
3:46:36 So again, running without the index versus with the index,
3:46:40 a 30x improvement in server latency using the index
3:46:44 and a 116x improvement on RU charge by using the index.
3:46:48 Right.
3:46:48 So, this really demonstrates that, you know,
3:46:52 we're able to really effectively use the index
3:46:54 once we do these advanced search or ranking,
3:46:57 methods for lexical or full text search.
3:47:00 The index is incredibly powerful for doing these very efficiently at scale.
3:47:07 And finally we have hybrid search.
3:47:09 Hybrid search combines multiple different measures like vector similarity
3:47:13 and full text scores into a single ranked result.
3:47:17 So for example, if you want your shoppers to be
3:47:20 able to make sure that maybe your searches contain certain keywords,
3:47:24 but you also want to capture the semantic
3:47:26 similarity from vector search to find related products.
3:47:29 Hybrid search could be a really great approach.
3:47:31 And again, you have to define both
3:47:32 the full text and vector index to utilize this.
3:47:35 And then the RRF function merges, the two scoring methods together.
3:47:43 Okay, so we'll clear the screen and run our final demo scenario here.
3:47:52 Okay, so we're going to see an example of the hybrid search.
3:47:55 So we're going to select the top five items
3:47:58 from our product catalog and we'll also project the vector similarity score.
3:48:02 This time we're going to order by rank.
3:48:04 We're going to use the RF function to fuse together
3:48:07 results from full text score and from vector similarity score.
3:48:11 And we're going to do this without using an index and with using an index.
3:48:14 So we'll go ahead and run this first without using the index.
3:48:21 And then we'll run this again while using the index.
3:48:25 And side by side comparison we see,
3:48:28 the server time 1 point 1 3x so faster and then ru charge is 2.8,
3:48:34 almost 3x, cheaper to run.
3:48:37 Right.
3:48:37 And again, these improvements, you know,
3:48:40 the server time looks relatively modest,
3:48:41 but you have to keep in mind that this is just on 10,000 data items.
3:48:44 As I keep mentioning, as the size of your data set grows,
3:48:48 benefit of using the index also increases.
3:48:50 The benefit of scoping to a particular partition also increases.
3:48:53 Right.
3:48:54 Especially if you're using any sort of search
3:48:56 like vector similarity search or the hybrid search.
3:48:58 Right.
3:48:59 Our vector indexing algorithms,
3:49:00 allow you to search in a way that scales sublinearly.
3:49:04 So as your data set grows, maybe your dataset grows 200 times in size,
3:49:08 your vector similarity search won't even increase by 2x in latency.
3:49:13 It's that efficient.
3:49:14 It really scales sublinearly with the size of your data.
3:49:17 So keep that in mind that as your scenario grows,
3:49:21 as your E retail shop becomes more and more popular and you
3:49:24 have more and more products and deal with more and more customers,
3:49:26 you're going to really retire the performance benefits
3:49:28 of using proper indexing techniques and query techniques.
3:49:35 Okay, so that's all we have today.
3:49:38 There are some references here to get you started.
3:49:40 All the material from this session, including the slides and the code examples
3:49:43 that I ran along with the sample data, are available at aka Ms.
3:49:48 Cosmos DB conf 26 query.
3:49:51 If you're interested in learning more about anything we talked about here,
3:49:54 whether it's queries, vector searches, full text search, hybrid search,
3:49:57 there's also some great learn documentation to get started.
3:50:00 And we also have a QR code that if you want
3:50:01 to scan that with your phone or scan it with your browser,
3:50:04 and it'll take you immediately to the GitHub repository
3:50:06 where we have all this content available for you.
3:50:09 All right, that's it.
3:50:11 Hope you enjoyed the content.
3:50:13 Thanks a lot folks.
3:50:14 See you next time.
3:50:18 Hi everyone, I'm Jess Ramos, or you might know me as Jess Ramos.
3:50:21 Data Online.
3:50:23 I'm the founder of Big Data Energy and I create data
3:50:26 and AI content for over half a million followers on social media.
3:50:30 Right now, everyone's talking about how good AI is.
3:50:33 You know, which model is faster, smarter, cheaper.
3:50:36 But for me personally, I can't stop thinking about the data.
3:50:40 We've all heard the saying garbage in, garbage out.
3:50:44 And this has never been more true than it is right now,
3:50:47 because you can have the most powerful model in the world
3:50:51 and still get completely unreliable outputs if your data layer is not solid.
3:50:57 AI doesn't run on neat structured tables.
3:51:00 It runs on vectors, documents in real time,
3:51:04 streaming data that's messy by nature and massive by scale.
3:51:09 And that's exactly where NoSQL shines.
3:51:12 We're talking AI search, real time personalization,
3:51:15 rag pipelines and agentic apps that need to read and write data fast.
3:51:21 And this is the stuff that actually
3:51:22 breaks your infrastructure if it's not built right.
3:51:26 Cosmos DB is built for exactly that.
3:51:28 Global distribution, single digit millisecond latency,
3:51:31 and both operational NoSQL and vector data.
3:51:35 And I'm so excited to keep learning more about this@ Cosmos DB conf.
3:51:39 Thanks so much for coming and let me know on LinkedIn what you learned today.
3:51:42 Bye hey, welcome back.
3:51:47 We are so glad you're here with us.
3:51:49 We've had a great start.
3:51:51 Well, we got about an hour to go.
3:51:55 Wow.
3:51:55 Yeah, thank you so much for sticking around and we have so
3:51:58 much more to show you as well as just the on demand sessions.
3:52:02 And you know what, we've seen so many things,
3:52:06 so many demos to help us move along,
3:52:08 we've got another amazing live session with none other than Andrew Lu.
3:52:13 You might have seen him earlier today.
3:52:14 Yeah.
3:52:15 Andrew is going to help us learn even more
3:52:17 about how Azure Cosmos DB works under the hood.
3:52:20 Andrew, welcome back to the stage.
3:52:22 Cool.
3:52:22 Hey, nice to be here.
3:52:24 Hey everyone, we're going to do a bit of an experiment.
3:52:28 We're going to do a whiteboarding session and we're going to go right into this.
3:52:32 So have you ever wondered like what happens when you go to the portal,
3:52:36 you hit Create Cosmos DB and just what happens behind the hood?
3:52:40 That's the intention of this session.
3:52:43 So let's start with our SDK and taking that perspective
3:52:49 from this point of the view of the world.
3:52:54 So the SDK is going to be embedded either in an application or even when we
3:52:58 build things like the Azure Portal we view
3:52:59 that as a client from the databases perspective.
3:53:02 So think of this as the app or the portal.
3:53:05 This is wrapping our SDK.
3:53:07 We actually have multiple flavors of the SDK beyond just programming languages.
3:53:10 We have a management SDK as well as a data plan SDK.
3:53:14 And what ends up happening is when we do from a client server,
3:53:17 perspective behind the scenes what this is going to go and talk to is a control
3:53:21 plane and a data plane and we're going
3:53:23 to think of these two things as very separate.
3:53:26 The control plane is going to be wrapped over arm.
3:53:33 This stands for Azure Resource Manager.
3:53:35 So whenever you make REST calls over into a DNS that looks something
3:53:39 like management.uhazure.com what happens behind the scenes
3:53:44 is you can think of arm, one mental model is like a big reverse proxy
3:53:49 and that reverse proxy when it goes hey, architecture, okay,
3:53:51 you want to create a Cosmos DB or add a region
3:53:54 or failover to a different region or scale your RUs.
3:53:57 What that's going to do is it's going to go
3:53:59 route those requests into what's called a resource provider.
3:54:02 So Cosmos DB has a resource provider
3:54:04 and of course every other Azure service out there also
3:54:06 has its own resource provider Arm's job is
3:54:10 to go and route to the correct Resource provider.
3:54:13 Now on the data plane this is going to be a different DNS entry.
3:54:16 This is going to be Cosmos Azure.
3:54:20 And the data plan is responsible for what you think of a database actually does.
3:54:24 Right?
3:54:25 This is going to be writing your data,
3:54:27 indexing it and then making it available for fast efficient queries.
3:54:32 So when we go and look at that path,
3:54:40 All of these requests will first flow into our front end server.
3:54:44 In Cosmos DB we like to call this our gateway.
3:54:48 The gateway's job is to do request routing because what happens is
3:54:51 we have a distributed backend behind the scenes in the back end.
3:54:59 What this will map to is effectively big clusters that we run behind the scenes.
3:55:04 They are multi tenant federations or clusters.
3:55:09 We like to use those words interchangeably within those clusters.
3:55:15 This is going to be the deployment model.
3:55:20 In Cosmos DB so if we go look
3:55:22 from a physical to a logical mapping on the logical
3:55:25 layer you're making this thing called a database
3:55:28 account down to a database down to a container.
3:55:33 I might use different words
3:55:34 like container or collection interchangeably container.
3:55:37 We use that as the API agnostic way of saying that.
3:55:40 And then depending on some of the Cosmos DB APIs,
3:55:42 you might have seen the word collection especially for any
3:55:45 of the document oriented data models within that we store our data.
3:55:50 Everything within this container is what is being
3:55:53 powered by the backend on the data plane.
3:55:56 Behind the scenes what this container is, it's not a single server.
3:56:01 What we're doing is we're stitching together replica sets.
3:56:05 So think of I have a little replica set powered by service fabric,
3:56:12 where I have a primary and a few different secondaries.
3:56:15 Each of these is a database, replica a database replica is going to be the core
3:56:22 building block for what's powering the software behind the scenes.
3:56:26 This is going to have the query engine, it's going to have the indexing layer,
3:56:29 it's going to have the, the storage and beyond.
3:56:32 Just doing the bread and butter of what a database does.
3:56:35 It's also going to have other things like admission control
3:56:37 and that's where we're going to have things like our resource governance.
3:56:41 So think of this what we're doing is we take this software,
3:56:45 we stitch it together into a replica set and on that replica
3:56:50 set this is what's going to form a physical partition.
3:56:56 And what happens is when you scale a container
3:56:59 we're going to create one or more physical partitions.
3:57:02 Each physical partition today is able to Support up
3:57:06 to 10 thousand request units per second worth of bandwidth.
3:57:09 It's also able to support up to 50gb worth of storage.
3:57:13 So when we talk about infinite scale, like millions of RUs,
3:57:16 what we're really doing is we're just taking
3:57:18 that container and we're sharding it aggressively where we're creating
3:57:21 tons of these, these different physical partitions behind
3:57:23 the scenes now these get placed onto these big federations.
3:57:27 So back to from a federation perspective,
3:57:29 think of this as in the data center I have a bunch of racks.
3:57:33 Each of these racks are going to have some
3:57:36 shared infrastructure like a top of rack network switch.
3:57:39 And we can also think of those in terms of helping us
3:57:42 compose what will become a set of fault domains or update domains.
3:57:47 So when we set up out these federations,
3:57:48 that go and host all these different database replicas,
3:57:52 what we're doing is on these different nodes,
3:57:55 think of these nodes as basically spinning
3:57:58 up these replicas and having them on standby.
3:58:00 And so when you go through the control plan saying hey,
3:58:02 I need to go scale my Cosmos db give me some more ru.
3:58:05 Behind the scenes what that resource provider is
3:58:07 doing with our control plan is saying hey,
3:58:10 this workload needs to scale and you some more partitions.
3:58:12 Do we have some bindable partitions?
3:58:14 Let's go take a look at this federation.
3:58:16 We might have a replica stitched together
3:58:19 across different fault domains and update domains.
3:58:22 So let's say I have a couple of different nodes here.
3:58:28 We're going to stitch together a replica from each one of these into a bindable
3:58:32 partition and we're going to give that partition over to that container.
3:58:35 Now these federations don't think of as an account
3:58:39 as having to live on a single federation.
3:58:41 We span across many federations, for infinite scale.
3:58:44 And why we do that is we decouple the physical and logical
3:58:48 layer and this allows us to always stay current on our hardware.
3:58:51 Speaking of the hardware, this is a nice shout out to our sponsor amd.
3:58:56 We use especially in the newer generations of our federations,
3:59:00 a massive amount of AMD CPUs.
3:59:02 Why we love the AMD CPUs is because they're incredibly power efficient.
3:59:06 They give really good watts, for the unit of compute,
3:59:09 that we give as well as that helps us drive the cost down.
3:59:12 And as we drive the cost down this is what we're also able to go
3:59:16 and Turn that into cost savings for our users
3:59:19 like you in the form of new features.
3:59:22 So things like the dynamic auto scale,
3:59:23 the RU pooling with the new Cosmos DB fleets the Cosmos DB serverless,
3:59:28 those Effectively what we're doing is,
3:59:31 is we're using elasticity to floor the TCO of operating a database.
3:59:36 Every workload is going to be bursty, any transactional workload is like that.
3:59:41 What we're doing is we're passing all of the savings we
3:59:43 get from the AMD into nice elastic features for our users.
3:59:48 Now on the federation,
3:59:49 as we go and stitch together these replicas we can jump into our first I
3:59:55 think bigger meteor topic and that is going
3:59:57 to be how we look at the partitioning.
4:00:08 So Cosmos DB uses a technique called consistent hashing.
4:00:11 It's a fairly ubiquitous concept for any kind of sharded database.
4:00:17 It tends to be the industry standard and I'll
4:00:19 kind of walk through that in terms a little bit
4:00:21 just to help understand it and compare it to what
4:00:25 maybe folks will think of as like a naive hash.
4:00:27 So in a naive hash think of it as I'm going to go partition my data
4:00:31 and behind the scenes that container now
4:00:33 has a ton of these physical partitions, right?
4:00:36 Like a container has let's say 50,000 RUs.
4:00:40 We know each one can support 10k each.
4:00:43 So this is going to go and create
4:00:46 five of these physical partitions behind the scenes.
4:00:48 Means it's going to go then route let's say some data onto this node,
4:00:53 some data over into this or this physical partition this physical partition.
4:00:57 How does that know?
4:00:58 And to deterministically route if I go and query
4:01:01 for let's say Alice's data and a user based system,
4:01:05 I need to be able to actually hit the partition hosting Alice's data.
4:01:08 Otherwise when I do that query I'm not going to find anything.
4:01:12 So what happens is we ask Hugh for, for a partition key.
4:01:17 So let's put in some values here just to give you an idea of what is happening.
4:01:25 So we have a logical partition key.
4:01:27 This is the partition key that you're supplying
4:01:29 and that's going to go hashed to some numeric value.
4:01:33 Behind the scenes we can use different hashing algorithms.
4:01:37 In this case Cosmos DB uses this algorithm
4:01:40 called Murmur hash because it's very compute efficient.
4:01:45 Let's look what happens if you do a naive hash like in a naive approach.
4:01:49 Let's say if I only had two partitions.
4:01:51 The challenge you're going to run into in a naive hash is if I,
4:01:54 let's say go and let's say mod two.
4:01:57 Well what I'm going to see is a distribution that looks like 1 2, 1 2, 1 2.
4:02:01 Now what happens if I need to grow
4:02:04 that on an elastic workload I'm scaling that over to three partitions
4:02:08 and right away you're going to go see what
4:02:11 the problem is and why a consistent hash is so valuable.
4:02:15 So when I do a mod 3 I'm going to get a distribution
4:02:17 that looks a little bit more like 123-123- Notice what's happening here.
4:02:24 If I am mapping data, let's say Alice,
4:02:27 Bob, Carroll, let's see, put Danny, Eve, et cetera.
4:02:32 Notice this challenge where if Danny Mapped to partition 2
4:02:35 when we mod 2 but now maps to partition 1,
4:02:39 we're doing effectively a huge shuffle.
4:02:41 A lot of data would have to go from the second partition into the first,
4:02:45 from the first into the second,
4:02:48 from the first into the third, the second to the third.
4:02:50 That's not very efficient.
4:02:51 So this is why a consistent hashing scheme is a much
4:02:55 better way of doing this than rather than a naive hash.
4:02:58 But this will give you the intuition in terms of how this works out.
4:03:02 So a consistent hash you can think of as a range
4:03:05 partitioning on top of a hashed based scheme.
4:03:11 So if I go and think about the hash values
4:03:13 where this is the min value on a number line,
4:03:17 and I do this as a circle since everything is going to go map onto this ring,
4:03:22 if I look at min value on where it hashes to 000 to ffff.
4:03:27 If I have two partitions on day one then
4:03:31 the way of scaling this is actually very easy on mapping.
4:03:36 What we can do is we can just go chop the number line, right?
4:03:40 And so from zero to approximately half and then
4:03:43 from half to let's say plus one over to max value,
4:03:46 I can go map this into physical partition one
4:03:49 and this can go map to physical partition two.
4:03:51 Now what's neat here is in a consistent hash,
4:03:53 rather than doing any of this shuffling,
4:03:55 what happens if the first partition is filling up
4:03:57 or what happens if the second partition is filling up?
4:04:00 Well then we can go and dynamically allocate more partitions by splitting them.
4:04:04 So what happens is let's say I need to do a split
4:04:07 where I need to go and map Now a third partition.
4:04:10 So we can go from zero to let's say quarter
4:04:14 and then quarter to quarter plus one over to half.
4:04:19 Well notice that I can now relocate, let's say, or not even relocate any data.
4:04:23 I can just go in on the hash values, directly.
4:04:27 If Alice was here, Bob was here and Carol was here,
4:04:32 rather than having to shuffle things around, I can just say hey,
4:04:35 Alice and Bob used to be co located on partition one,
4:04:38 but partition one has now turned turn into partition one,
4:04:41 let's say prime and partition one double prime.
4:04:44 All we've done is split that data.
4:04:46 What's awesome about this is Cosmos DB can do
4:04:48 this as an online operation where we'll handle the splitting or merging
4:04:51 of the data on allocating the partitions and then dynamically cut
4:04:56 over whenever you need to go and perform one of these operations.
4:05:00 So what Cosmos DB does is really the runtime aspects of partitioning,
4:05:05 so that you have auto partitioning.
4:05:07 Really as a user, what you're left with is the design time aspect,
4:05:10 which is choosing a good partition key.
4:05:12 And this is why it's so important to choose a good partition key.
4:05:16 Because ultimately if you're partitioning, let's say by a user id,
4:05:18 then what we're doing is we're using that user
4:05:20 ID to go deterministically route which piece of data,
4:05:24 let's say Alice's piece of data into partition one prime,
4:05:27 Bob's data into partition one double prime and Carol into partition two.
4:05:33 Now this is one aspect of scale out.
4:05:35 This is the partitioning aspect.
4:05:36 The other thing I'd like to cover is the replication side.
4:05:40 What replication is really therefore is a, little bit different.
4:05:43 Partitioning is what gets you ability to scale your request traffic.
4:05:47 In particular it does really well at scaling your write traffic.
4:05:50 It can also scale your read traffic.
4:05:52 If I have millions of QPS,
4:05:55 well I cannot have an infinitely number of CPU on a single box on a single vm.
4:06:00 There's still the physical aspect of it.
4:06:02 We just put you on more computers and so what we're doing is we can power many,
4:06:07 many CPUs as long as we divide that data set and partition it up.
4:06:11 Now what replication is doing a little bit different is copying that data.
4:06:14 Copying that data gives us two different things.
4:06:16 One that's similar and one that's very different.
4:06:18 The similar thing is going to be it's one way of scaling your read throughput.
4:06:22 The other thing that it can do is it gives you a lot of Redundancy.
4:06:25 And this is how Cosmos DB gives you five nines of availability.
4:06:30 And we can think of this in multiple grains.
4:06:33 So let me go and get some board real estate.
4:06:41 On the replication side.
4:06:42 When we zoom in on any of these partitions,
4:06:45 remember what we have is a replica set.
4:06:50 What we're doing is we're using a concept of called
4:06:53 quorums or quorum replication to go and provide that availability.
4:06:56 Now how we do those replica sets is going to be
4:06:58 a little bit different in terms of within a region versus across region.
4:07:02 When you do across regions, what we're effectively doing is we're spawning
4:07:05 another partition in each of the different regions.
4:07:08 And then what happens is we can have multi write regions.
4:07:12 What happens is one of these secondaries will get designated as a forwarder.
4:07:17 We might have three and we might have four.
4:07:19 This is a bit black boxed.
4:07:21 But what we do is we dynamically size a quorum across a replica set.
4:07:25 And this can be in let's say region two,
4:07:27 this can be in region one, within a region.
4:07:33 We rely on surface fabric heavily to do the within a region, replication.
4:07:40 And what's beautiful about this is it handles
4:07:43 a lot of the things like leader replication.
4:07:45 For us, let's say a primary is going down,
4:07:47 whether it's willingly, intentionally.
4:07:50 Let's say we're deploying a database software update.
4:07:52 Have you noticed Cosmos DB doesn't really have maintenance periods.
4:07:55 It's just always online.
4:07:56 You've never seen us say hey,
4:07:57 it's time for a Windows update or it's time for an application update.
4:08:01 That's because behind the scenes what
4:08:03 we're doing is we're just dynamically changing.
4:08:05 Okay, let me take down a replica that's
4:08:07 on a different update domain or a different fault domain,
4:08:09 and let me designate and promote one of these secondaries
4:08:12 over to a primary and then shift that traffic over.
4:08:16 What we're doing is we're live doing rolling updates.
4:08:19 And that way when this replica is brought back online,
4:08:22 what we can do is we're just dynamically writing to a rolling set of replicas,
4:08:28 behind that replica set.
4:08:29 Now what a quorum, is a little bit
4:08:32 different from your traditional primary versus secondary replication.
4:08:37 In a traditional database like let's say a relational database,
4:08:41 what you're typically working with is
4:08:44 either synchronous or asynchronous replication.
4:08:48 Those come with pretty steep trade offs.
4:08:50 When you do synchronous replication.
4:08:52 If your goal was high availability, if the primary goes down, that's fine,
4:08:56 you can go and fail over to a secondary,
4:08:59 but actually goes and has some risks as well.
4:09:04 What if the instead of the first failing, it's the second one that went down?
4:09:09 Right?
4:09:09 Well, if you're synchronously replicating,
4:09:11 well even if the primary is healthy now I have
4:09:13 to go and cut that secondary in order to restore,
4:09:16 work back onto bring the right availability back.
4:09:21 So what synchronous replication is doing is it's giving,
4:09:25 giving you really strong consistency.
4:09:27 And in terms of disaster recovery terms,
4:09:30 what you're really prioritizing is a low recovery point objective, right?
4:09:35 Like if I go from the primary to the secondary,
4:09:37 since everything was synchronously replicated,
4:09:39 if I rotate the primary, I know it's in the secondary.
4:09:42 So I have a recovery point objective of effectively zero.
4:09:45 Now where synchronous replication is really
4:09:47 bad at is the recovery time objective.
4:09:49 If you are having a coordinate effect failover.
4:09:51 Failovers take time.
4:09:53 So this is where you're going to be less happy, as opposed to more happy.
4:09:59 And then asynchronous is the exact opposite, right?
4:10:01 In an asynchronous model,
4:10:03 what you're going to see is your recovery time objectives.
4:10:06 Well, hey, if the secondary fails, that's fine.
4:10:09 I can just keep serving out of the primary.
4:10:11 I'll just let the secondary lag behind.
4:10:15 But how far is it lacking behind, Behind.
4:10:17 If I do a cutover to the secondary, is the data that I just wrote even there?
4:10:23 So what you're going to see is the recovery point objective is going to be,
4:10:26 well it's an eventually consistent model.
4:10:27 So you're going to have a lack of consistency,
4:10:30 but you're trading up the recovery time objective.
4:10:33 Now where quorum replication,
4:10:35 innovates on this is rather than thinking of the world as hey,
4:10:39 let me just have one additional redundant copy or N
4:10:41 number of redundant where I have a primary secondary model,
4:10:44 what I can do is I can always write to a quorum of replica.
4:10:48 So imagine that if I have four replicas here,
4:10:51 a majority quorum would be something like writing
4:10:54 to at least three out of four of those replicas.
4:10:58 That means that if any replica were to fail, in this example, let's say,
4:11:03 this fourth replica bombed out,
4:11:05 and I was able to write to the primary and to these secondaries,
4:11:09 or even if that was primary dies,
4:11:12 we go and promote One of the other replicas to be the new
4:11:15 primary and we get three out of four on the new topology.
4:11:18 Any node can fail and we're able to maintain that availability.
4:11:24 Now on the read side this is what is really interesting is if
4:11:27 I go and circle any two replicas I can do a minority quorum.
4:11:33 Well because I'm allowing fault tolerance for, for any
4:11:36 replica to go down on the write path.
4:11:38 On the read path as long as I read
4:11:40 at least two different replicas I can still guarantee strong consistency.
4:11:45 Even though this is in a fault tolerant
4:11:46 system where I allow a replica to go down,
4:11:49 I'm always going to have replica overlap.
4:11:52 If replica D goes down well A and B will of course both have the data.
4:11:56 B and C would both have the data C and D.
4:11:59 The D would miss, but the C would still
4:12:03 because it's further forward in time it's always going
4:12:06 to win that vote that you're able to get
4:12:09 strongly consistent reads in the midst of a downed replica.
4:12:14 Now that's what we're doing within a region and then across regions.
4:12:20 What we do is we designate one of these secondary
4:12:23 replicas to go and replicate over to another region.
4:12:29 There's some slight caveats to that.
4:12:31 Depending on whether you're using strong consistency versus
4:12:33 any of the other Cosmos DB consistency levels,
4:12:35 if you're using any of the weaker consistency levels you can think
4:12:39 of it as we're doing an asynchronous replication over to the secondary region.
4:12:43 And this is what gives you well higher availability,
4:12:47 faster RTOs at the trade off of of consistency
4:12:56 or if you're using global strong consistency.
4:12:58 What we're doing is you can think
4:13:00 of it as synchronously writing to multiple regions.
4:13:03 Now there's a bunch of more things that go into that.
4:13:05 For example we can do dynamic quorums where if we have a third region
4:13:08 we can vote two out of three regions to give you even strong consistency.
4:13:12 That allows a region to go down.
4:13:15 And what's really cool is Cosmos DB allows it's not a single primary
4:13:19 system where able to put primaries in each of these two different regions.
4:13:24 On the SDK side you never have to really wait for a failover
4:13:28 if you're any of the weaker consistency levels using multi region writes.
4:13:32 If I'm writing to region one and region one goes down that's fine.
4:13:36 The SDK is actually smart enough when you instantiate the SDK,
4:13:39 yeah, that connection policy, you can set what the preferred location is.
4:13:43 What it's doing is it's going to do a live reading redirection
4:13:46 to the secondary region without even waiting for a server side failover.
4:13:50 This is how you can collapse that recovery time objective down to zero,
4:13:54 even in a strong consistency manner,
4:13:58 we're able to go and minimize the RTO where we now
4:14:01 have the ability of doing fast failovers in a strongly consistent setup.
4:14:06 We are very happy about.
4:14:08 Recently we had launched per partition automatic failover.
4:14:12 What it does is rather than looking at hey, if one partition's unhealthy,
4:14:15 another partition is healthy and trying to have
4:14:18 a live ops team go hey, should I fail over?
4:14:20 Should I not fail over?
4:14:21 That tends to be the bottleneck for any failover.
4:14:24 We don't have to go and do any
4:14:25 decision making in a strongly consistent system with ppaf.
4:14:29 What we can do is we can also
4:14:30 then trigger that one partition to independently trigger
4:14:33 a failover in this big distributed database without
4:14:36 having to impact all of the other healthy partitions.
4:14:39 So this is how we're able to go and try
4:14:40 to give you just softer trade offs rather than thinking
4:14:43 of the world as extremes in terms of eventual or strongly
4:14:46 consistent systems with like hey do I want RTO or rpo?
4:14:50 You can get both by using quorum topologies for replication,
4:14:57 and then doing a lot of magic around redirections.
4:15:01 Now a third thing I want to go and show you that I think is
4:15:04 good for every developer to understand in Cosmos
4:15:07 DB is the resource governance aspect, right?
4:15:10 Like when you go and create a Cosmos DB container you're going to see like hey,
4:15:12 what is this thing, request units.
4:15:14 This is what allows us to go and do paper request models around serverless.
4:15:19 It allows us to give you guaranteed throughput rather than just best effort.
4:15:24 Hey, we have some nodes with some cores,
4:15:26 we're able to give you predictable performance.
4:15:28 So what happens is each of those replicas, when I look at the replica,
4:15:32 where behind the scenes you can think of it as we have a resource
4:15:36 governor and you can think of that as a token bucket algorithm, right?
4:15:40 If you haven't seen token bucket algorithm before,
4:15:43 think of it as like of a bucket
4:15:44 of tokens and the request units are those tokens.
4:15:48 When you go and do a database read,
4:15:51 a one KB read is actually the baseline charge.
4:15:53 That's one request unit.
4:15:55 What we'll do is we'll take a token out of the bucket.
4:15:58 So I'll subtract one request unit and then
4:16:01 on the resource governor in order to go and say hey,
4:16:03 is the system overloaded, is it not overloaded?
4:16:05 We want to give predictable latency so let's
4:16:08 try to avoid anything like overly saturated cpu.
4:16:13 What we'll do is we'll govern that based off
4:16:14 of the relative CPU memory and IO footprint of that request.
4:16:18 So each second what we're doing is that we're
4:16:20 then replenishing those tokens back in on a timer interval.
4:16:25 As long as I have greater than 0 RUs, we're able to keep serving requests.
4:16:31 Now one innovation that we had that's a bit new is
4:16:36 the RU pooling functionality that we did with the Cosmos DB Fleets.
4:16:42 So what happens is we have one
4:16:44 of these token bucket resource governors on each individual replica.
4:16:48 So on that, that partition, so I have a partition,
4:16:52 the RUs right now first is a shared nothing counter.
4:16:55 We have this on each individual replica.
4:17:00 This is now a partition and then we go and divide that evenly across partitions.
4:17:06 What happens if I have a hotter partition,
4:17:08 and then I have let's say a colder partition.
4:17:11 Can we go and donate budget for from a quarter
4:17:15 partition over to a hot partition in a responsible manner.
4:17:19 So that's what the RU pooling is all about.
4:17:21 We're able to pool across different accounts up to we
4:17:24 have also a node level governor but on the local
4:17:28 governor what's happening is historically if we just use
4:17:32 a shared nothing counter we have a few different trade offs.
4:17:35 Right?
4:17:35 The benefit of having a shared nothing counter is let's say
4:17:38 if one node goes down it's not a global rate limiter.
4:17:42 I can go and continue to serve requests and the system is highly resilient.
4:17:46 But the trade off here is well what
4:17:49 happens if I'm overly spread thin across many replicas,
4:17:53 many partitions is all of these now also
4:17:57 report into we call a global rate limiter,
4:18:00 or we internally like to call it a controller service.
4:18:03 And so think of it as a heartbeat where it's not
4:18:05 on the critical path where each of the local replicas is
4:18:09 going and saying hey here's my utilization on how many RUs
4:18:12 I am consuming and here's the leftover budget that I have.
4:18:16 All of that leftover budget is then sent over to the global
4:18:19 rate limiter where the global rate limiter on the response
4:18:22 can also go heartbeat back saying hey you Just use
4:18:25 all of these RUs and you have some leftover budget.
4:18:28 Meanwhile this other partition is running hot.
4:18:31 But because we have leftover budget somewhere else.
4:18:34 Let's say if this is the one that's
4:18:39 running blazing hot and this one's got leftover budget,
4:18:42 what we're able to do is the global rate limiter can say,
4:18:46 hey, let's relax some limits.
4:18:48 I can give and donate some RUs back over
4:18:50 to the other partition and help you pull that.
4:18:53 So these are some of the really cool
4:18:54 new capabilities where we can give you leftover budget,
4:18:57 give you higher burst ability and do so
4:18:59 in a responsible manner that still preserves a shared nothing architecture.
4:19:03 On the front door.
4:19:05 Now, this is a lot to take in.
4:19:07 Let me go and zoom out.
4:19:11 We talked a little bit on partitioning,
4:19:12 on replication and on request units on what
4:19:15 happens behind the scenes in Cosmos DB.
4:19:17 Let us know what you think about this format if this works out for you.
4:19:20 There's a lot of other topics we can do in the future like indexing,
4:19:23 there's query engine, et cetera.
4:19:24 Let us know in the comments on what you
4:19:26 think and please do fill out our form evaluations.
4:19:29 Otherwise, thank you so much.
4:19:32 Thank you so much.
4:19:33 Thank you, Andrew.
4:19:34 You know what's crazy is that that was
4:19:37 such a great behind the scenes look and you
4:19:40 gave me something like this, I think maybe the first day I started on the team.
4:19:45 So it's always great to kind of see this again.
4:19:47 It's a refresher and keeps me really up to speed on what we're doing.
4:19:52 Yeah, I mean I figure we keep doing this as an onboarding thing internally,
4:19:56 but honestly it helps our developers so much
4:19:58 when they know what's happening behind the scenes too.
4:20:00 On rationalizing, hey, why is the system behaving the way it is?
4:20:04 And that's the intent of also disclosing this out broadly here.
4:20:07 Fantastic.
4:20:08 Well thanks Andrew.
4:20:10 We appreciate it.
4:20:11 And we now know a little bit more about how Azure Cosmos DB operates, at scale.
4:20:17 And you know what, we had so much information shared with us and I
4:20:21 want to go ahead and share some things that we've had from our community.
4:20:25 Nicholas Chang, he's a really,
4:20:27 really wonderful member of the global AI community,
4:20:30 working to help get everybody.
4:20:32 You might have seen him on one of these AI tours that we've got going right now,
4:20:36 but he actually said that if Damian Brady,
4:20:38 who's a buddy of mine who's over at GitHub in their DevRel.
4:20:43 He said it was good to go.
4:20:44 I think you should go.
4:20:45 Patty.
4:20:46 What else do we have as far as programming language is a concern?
4:20:48 Yeah.
4:20:49 We also released a poll about what programming language used the most
4:20:52 and it was great to hear that 45% of udevelopers use.
4:20:58 Net.
4:20:59 I personally am a Python developer,
4:21:01 so I'm glad that I'm joined by 32% of the audience.
4:21:04 And we're followed by Java, JavaScript, TypeScript and Java.
4:21:08 So really exciting stuff.
4:21:09 Thank you for filling out our poll.
4:21:11 Absolutely.
4:21:11 We really, really appreciate it.
4:21:13 Yeah.
4:21:13 But now we're moving into one
4:21:15 of the most critical topics which is data modeling.
4:21:18 Yeah.
4:21:18 The decisions that you make early can define your application success
4:21:21 and it's a great reason to keep learning and sharpening your fundamentals.
4:21:25 I know one of the things that I always hear from Andrew,
4:21:28 who we just heard from is making these decisions early
4:21:32 are going to help you save yourself from headaches later.
4:21:36 No, 100% preparation really is key and if
4:21:39 you want to lock in those fundamentals.
4:21:41 The Cloud Skills Challenge, again it's a great follow on.
4:21:44 So data modeling is one of those topics and again it runs through May 8th.
4:21:49 Yeah.
4:21:49 You complete the modules and you can enter
4:21:51 to win one of 500 free DP-420 exam certification vouchers.
4:21:56 So start here at aka Ms.
4:21:59 Cosmos DBConf challenge.
4:22:02 So next we have an awesome talk
4:22:04 from one of my favorite community champions, Hassan Sovereign.
4:22:07 Yeah, and we're going to be talking about data modeling decisions,
4:22:10 Jay and that they will make or break your Cosmos DB application.
4:22:14 Let's go.
4:22:18 Hi everyone.
4:22:19 Hi Azure Cosmos DB Community, I am Hassan Savran.
4:22:23 I truly enjoy meeting many of you at test conference and I cannot
4:22:27 wait to make new connections and talk about Cosmos DB at upcoming events.
4:22:33 Today I'm going to talk about the data modeling with you and I'm
4:22:37 going to try to make it interactive as much as I can.
4:22:40 So I have no slides for you.
4:22:42 Today we are going to work on the VS code, as much as we can.
4:22:46 And you know, we are going to actually look at the Stack Overflow data model
4:22:50 and try to move it to Cosmos DB and make it scalable and very affordable.
4:22:56 Currently it's in the relational database.
4:22:59 So we are going to take it from there
4:23:01 and move it to Cosmos DB by using different VSCode, extensions and tools.
4:23:06 So I just want to kind of show you how
4:23:08 I guess I handle when I get a project like this.
4:23:13 So I guess before, we Move to the VS code.
4:23:17 Let's look at what we are dealing with.
4:23:18 Right.
4:23:19 So we have limited time so we cannot
4:23:21 kind of look at all the tech overflow tables.
4:23:24 So we are going to actually focus on the busiest
4:23:27 part of the application which is we display the questions and we are going to go
4:23:32 and click on a question and get all the answers.
4:23:36 Right.
4:23:37 Sounds simple but as you can see there's all things that are going on here.
4:23:41 So we have like questions, we have answers,
4:23:44 we have comments, we have badges, we have votes.
4:23:48 So we have potentially really a large application page to display here.
4:23:54 So this is what we are going to tackle today.
4:23:57 So let me go back to my VS code and let's
4:24:00 start from the relational database where the data is actually currently at.
4:24:05 So these are the tables we are going to tackle today.
4:24:08 So we have the post table here.
4:24:11 As you can see there's a bunch of kind of like a columns here.
4:24:16 We have each post can have multiple comments and multiple votes.
4:24:21 We have the user's information and users can have number of batches.
4:24:25 So in a relational database you can easily join this table
4:24:28 together and really deliver whatever the query is asking for.
4:24:33 But we are going to go to the Cosmos DB, which we don't have any more join.
4:24:37 So what are we going to do?
4:24:39 Well before we jump what we are going to do, let's talk about what we should,
4:24:43 we shouldn't just take this as it is and move it to Cosmos.
4:24:49 This is not going to work because we don't have joins in the database side.
4:24:53 So you have to kind of go and join each of these, you know,
4:24:57 data in your front end or in the middle layer which is not going to be easy.
4:25:02 So let me actually show you if we will actually
4:25:06 take this as it is and put in Cosmos db, what's going to happen?
4:25:10 So I have a bunch of the different databases here.
4:25:13 So what the application we are seeing here, this is the one that I created.
4:25:17 It's available in VSCOPE extension, it's just for the Azure Cosmos DB community.
4:25:22 It's free, you can download it.
4:25:24 So what I'm doing here is I'm just picking my database.
4:25:28 As you can see all these tables that I just show you, they are here available.
4:25:32 So we are going to go and start with the probably the posts table.
4:25:37 So first we have to go get the questions
4:25:38 and answers and the post table has both of them.
4:25:41 So I'm just going to go and actually go pick this one before I click to execute.
4:25:48 Actually there's A good function here.
4:25:50 I can.
4:25:51 So I'm just going to run my query analyzer here so
4:25:54 I can track how much everything is costing to me later.
4:25:58 So the first we are going to go get the questions and answers.
4:26:01 As you can see it looks like we have three.
4:26:03 I'm guessing I have one question and two answers here.
4:26:07 As you can see it cost me three requests.
4:26:10 Okay, not bad.
4:26:11 But we are not done yet because we have many other
4:26:14 stuff that we actually have to go and get like comments.
4:26:17 So to get the comments we kind of need to use
4:26:19 the in here because we are dealing with three objects.
4:26:22 So we have to pass their post ID and get this you know, information.
4:26:27 So let's go and get that too.
4:26:30 That cost me 337.
4:26:31 Okay, not bad.
4:26:33 But now we are going to continue.
4:26:35 You might need to get the votes or maybe try
4:26:38 to count the votes and display them on the page.
4:26:41 So we need this information too.
4:26:43 Another three then as you saw on the, you know,
4:26:48 the page there are some user information.
4:26:50 So I need to go and get the users,
4:26:52 for these questions and answers I'm going to go and get
4:26:56 those another three that I have to get, you know, paid for.
4:27:00 Then I'm going to go and get the badges.
4:27:02 I may need to maybe count the badges of each of these users,
4:27:05 display emojis or you know, whatever depending on what they have.
4:27:09 Plan to get that information.
4:27:12 Well, what's happening here?
4:27:13 Looks like I have 277 badges and it actually cost me 16 request units.
4:27:19 If I'm going to go back to my query analyzer, this is really what's happening.
4:27:23 Just to display one page.
4:27:25 I need to sum it up.
4:27:27 All these request units.
4:27:29 I couldn't join any of this data in database level.
4:27:32 So that's going to be my problem to solve in front end or in the middle.
4:27:36 So that's why you don't want to get a relational
4:27:40 data model and move to Cosmos DB as a base.
4:27:43 It's not going to be you know, good for feature for today.
4:27:48 As we can see it's kind of as the data is going to get larger,
4:27:51 if the traffic is going to get larger,
4:27:52 it's going to get, you know, cost you more and more.
4:27:56 So what do we do?
4:27:57 Well, what do we do is it's pretty
4:27:59 clear we need to reduce the number of containers.
4:28:04 So for that let's actually look at this diagram here.
4:28:07 That's where we are right now.
4:28:09 We want to try to write and read data, but currently we have five containers.
4:28:16 We only look at the read, we didn't even look at the right.
4:28:19 So if you actually.
4:28:19 I want to look at the right, I can quickly do that here.
4:28:23 So let's actually try an update.
4:28:25 So I'm going to just update a post.
4:28:27 I'm just increasing those numbers here by using the patch operation.
4:28:31 Let's see how much this is going to cost me.
4:28:33 15 requests and that's only one container.
4:28:37 So you know, things can get really expensive really quick with this.
4:28:42 So as I said before, we want to, you know,
4:28:45 reduce these numbers, the number of the containers here.
4:28:49 So what are we going to do here is.
4:28:52 Well, we have the post ID is the partition
4:28:55 key in each of these, contexts as you can see.
4:29:01 Now the good, news is we have the post ID for, you know,
4:29:05 getting shared by these three containers and the user
4:29:08 ID is getting shared by this too.
4:29:10 Since they share the same partition key,
4:29:12 why don't we just put them in the same container,
4:29:15 just like the version two suggests?
4:29:18 Right.
4:29:19 So we have a post.
4:29:19 Our partition key is post ID and each comments
4:29:22 and was is going to be under the same post document.
4:29:27 Sounds good on paper.
4:29:28 We can try it.
4:29:30 Also that applies to the users.
4:29:32 So rather than we have five containers, we end up with two containers.
4:29:36 So let's try that.
4:29:37 Let's see how this is going to work.
4:29:40 So I'm going to go and change my data model to two here.
4:29:44 Go back to results.
4:29:46 And as you can see, I have only two containers now.
4:29:49 So let's start with the post.
4:29:51 Well, for that I'm gonna, run or execute the same exact, query that I just ran.
4:30:00 So let's see what's gonna happen here.
4:30:01 Execute this one.
4:30:04 It cost me three request units, but this time, I got all the information.
4:30:09 I have the questions, I have the answers,
4:30:11 I have the comments, and I have the votes.
4:30:13 And if you look at the document size 5.71, it's really not that bad.
4:30:18 The good part is if I actually go scroll down,
4:30:21 you're gonna see everything is already mapped.
4:30:23 So I know that these words are for this question,
4:30:26 these comments are for this question.
4:30:28 So everything is really kind of ready to render, in the front end.
4:30:32 So I kind of like that.
4:30:34 The next one, this is, I think the answer.
4:30:37 There's no what's and comments.
4:30:38 Great.
4:30:40 The third one, let's see what's going on here.
4:30:42 Oh, that one kind of gives me some red alerts here.
4:30:46 Why?
4:30:47 Well, look at the number of comments we have.
4:30:49 21.
4:30:51 So which tells me this, the Relationship between the post
4:30:54 and comment is not one to few, it's one to many.
4:30:57 Which means if comments are going to grow and we don't have a limit for it
4:31:01 then that means that we are going to have an issue in feature for sure.
4:31:06 This is great for today but you know
4:31:08 you really need to worry about the feature too.
4:31:10 When you start to have more data, more traffic comes in your database.
4:31:15 So I know this cost great in all three but I don't think that this idea is
4:31:21 going to actually work for this data model
4:31:25 because in feature you know we will have issues.
4:31:28 Well let's look at the users table too I guess that was the only post.
4:31:31 So this is my user table.
4:31:33 Let's execute it.
4:31:35 Let's see what's happening here again.
4:31:37 299 I have looks like one user here but look at this badges.
4:31:42 I have 54 badges which kind of tells me
4:31:46 that maybe users can get the same batch multiple times.
4:31:50 Then this is not going to be easy to kind
4:31:52 of keep them under the one object like that.
4:31:55 So this is a one to many relationship too and really just doesn't
4:32:00 make sense to me especially for feature to put them in the same place.
4:32:06 The mainly because it's going to actually increase our data
4:32:09 model and there's going to be no limit for it.
4:32:12 So I'm going to actually go back to my diagram
4:32:15 and try to maybe talk about the sizes.
4:32:19 The size is important.
4:32:21 As you can see the data model size, you have a limit for it.
4:32:24 Two megabyte is the limit.
4:32:25 It cannot be larger than that.
4:32:27 And your data model, as, as long as your data model gets getting large and large
4:32:32 that means that your cost of the query is getting you know,
4:32:36 expensive and expensive.
4:32:37 Same with the right because well Cosmos DB needs to have you know,
4:32:41 handle many more properties and it's going to cost you more.
4:32:46 So whenever you create a data model you want to be you know,
4:32:49 like give some kind of estimate how big this data model is going to be.
4:32:53 So if your data model is going to continue growing
4:32:55 and growing probably you're going to hit that two megabyte
4:32:58 limit and your solution is going to get more expensive
4:33:01 and more expensive every year as data grows and traffic grows.
4:33:05 So because of that, yeah it is a good
4:33:09 kind of it give me a good result right now.
4:33:12 But for feature, I don't think that I can
4:33:14 end all this under the two documents or two containers.
4:33:20 But what do we do?
4:33:22 Well we know one thing we are going to roll back,
4:33:25 but we really cannot roll back to that because it's not really cheap either.
4:33:30 The good news is, well we are still,
4:33:32 you know, I guess sharing the same partition keys here.
4:33:37 And I'm not sure if you notice one thing here,
4:33:39 there is actually a nightmare here waiting to happen too.
4:33:43 So if you look at our select statement,
4:33:45 it says Post ID is this number or Parent ID is this number.
4:33:50 So the first one is giving you the question,
4:33:52 the second one is giving you the answers because
4:33:54 both of them are actually in the same post.
4:33:57 Table.
4:33:58 The problem is the post ID is the partition key, Parent ID is not.
4:34:03 So as soon as, right now we have all data in one physical partition.
4:34:08 As soon as you start to distribute this data,
4:34:11 this query is going to become a cross partition
4:34:14 query and it's going to get ready expensive really quickly.
4:34:18 So we need to do something about that too.
4:34:21 So what do we do from here?
4:34:23 Well we can try the document type data modeling.
4:34:29 Document type data modeling is working a bit different.
4:34:35 So we still have two containers here, Post documents and user documents.
4:34:42 But rather than saving only one entity type in the, you know, one container,
4:34:46 we want to save different types and we want
4:34:49 to control that with a new property named document type.
4:34:53 So I can easily figure out which post, this question,
4:34:56 which post is answered by looking the post type id.
4:34:59 So I just, you know,
4:35:01 change that and add question and answer document types in this post.
4:35:05 So I can easily, for example when you click the question,
4:35:09 tell me your question id.
4:35:10 And by knowing the question id I should be able to get all the questions,
4:35:14 all the answers, all the comments and all whatever really from this container.
4:35:19 So this might actually work very well for the situation.
4:35:23 Same with the users.
4:35:24 You know, you can just get the user information
4:35:26 or you can get the badge or you can get all.
4:35:29 And the other part that I really like about this one, this is open so you,
4:35:32 if you have different kind of, you know, document types later you can continue
4:35:36 adding here rather than creating new containers.
4:35:40 So let's see how this is going to work.
4:35:42 I'm going to go back to my Cosmos DB NoSQL studio,
4:35:45 change my data model to data model 4
4:35:49 and let's start with the post document first.
4:35:54 Now as you can see I don't have the post ID anymore.
4:35:58 My partition key is question id and if I want
4:36:01 to get only the question for this, for example question,
4:36:04 I can run this and just execute this one
4:36:10 here and this will give me only the question.
4:36:15 2.94 not bad.
4:36:17 I can get the only the answer here too.
4:36:20 I'll execute this one.
4:36:22 296 as you can see I have only the answers
4:36:25 because there are two answers here or I can get everything.
4:36:30 I don't care.
4:36:30 It's a question answer comment at the end.
4:36:32 I need to render this in some way in my front end.
4:36:36 So I can just do that and look at it.
4:36:39 This is beautiful.
4:36:40 You know what is beautiful?
4:36:41 This is beautiful here.
4:36:43 The cost is 3.61 and I was able
4:36:46 to get everything 29 items with that much request.
4:36:51 So this is really very scalable and very affordable solution.
4:36:56 I am really happy about this one.
4:36:58 So as you can see I have some answers here, I have some comments here.
4:37:02 And you can easily look at the document type and you
4:37:05 know render that in the correct places in the front end.
4:37:08 Or if you want to kind of use a factory
4:37:10 design and kind of put them in your, you know,
4:37:13 the C, sharp object or anything like that, you can do both.
4:37:16 So this is, this is working great.
4:37:17 I'm really happy about this one.
4:37:21 Now next is usually developers will kind of stop here.
4:37:28 Why?
4:37:28 Because probably that was the requirement.
4:37:31 Create a scalable affordable data model.
4:37:35 Which we did.
4:37:36 Right.
4:37:37 Well I would suggest you not to stop there because there's
4:37:41 all kind of other things that you can still work on.
4:37:44 And those are what about the feature?
4:37:47 What is going to happen in feature?
4:37:49 How can I make this data model better so it can
4:37:52 tackle the problems is going to come in feature for the system.
4:37:57 So for that there are a couple of options you can actually look at.
4:38:01 The first thing I can look at is probably the tenants, right?
4:38:05 So yeah, we have on the stack overflow of questions
4:38:08 today but as you know stack overflow got much bigger.
4:38:12 There's a stack exchange.
4:38:13 You can actually take this, you know,
4:38:15 questions and answers and apply to different domains and they
4:38:18 can use talk about the totally different questions and answers.
4:38:22 So therefore it will be a different ui but in the back end it will be the same.
4:38:27 So we can really easily make that ready if
4:38:30 we actually add the tenant ID to our data model.
4:38:36 But we can do that.
4:38:37 And if you're going to do that then that means
4:38:39 we can actually use the hierarchy of partition key.
4:38:41 Currently we are using only regular partition key which is the question id.
4:38:45 But if I'm going to do something like that, what is that going to look like?
4:38:48 Well all I need to do is I need to go and create
4:38:52 a tenant ID in here and whenever I'm asking the question here.
4:38:56 You know I'm going to Add the tenant ID1 and question ID is this.
4:38:59 So really that's what my select statement is going to change into.
4:39:03 And when I run that issue, it should run just fine.
4:39:07 Let me go data model 5 first.
4:39:12 As you can see I have the data so it's still three or five.
4:39:16 And if you actually go to Cosmos DB info tab here and look at the oral.
4:39:21 So look at the partitioning here.
4:39:23 My partition key is the tenant ID and question id.
4:39:26 So whenever I actually I ask this question id let's say we have two tenants.
4:39:30 One is about the C sharp.
4:39:32 Other one is maybe the farmers who's you
4:39:34 know asking a question how they're your chickens are, you know, can do better.
4:39:39 So in this case if you're going to tenant
4:39:41 ID 1 we are only search the C sharp question.
4:39:44 We are not going to search for you know, farmer's question in that case.
4:39:48 So you can easily you know change where they actually
4:39:53 are looking by you know changing the tenant ID here.
4:39:58 So I like this one.
4:40:00 This will be actually make the system ready or multi tenant system.
4:40:05 What else we can do?
4:40:07 Next one is find questions.
4:40:09 Now you're going to say well this is great
4:40:11 but how are we going to find these questions?
4:40:13 Right?
4:40:14 So the question is already in the page, you click on it.
4:40:17 But users needs to find these questions in some way.
4:40:20 So there are two things I will suggest for that.
4:40:22 In older days we used to have like categories,
4:40:25 we used to have tags user supposed to go and click the category,
4:40:29 find the products, what they need to find.
4:40:32 But well the things start change very quickly, right?
4:40:35 So right now everybody likes the chat, everybody likes to write,
4:40:38 you know what they need and they are supposed
4:40:40 to go and find whatever they are looking for.
4:40:43 So for that case what we can do here is
4:40:46 Cosmos DB has two options that you can easily set up.
4:40:49 The first one is full text search.
4:40:51 In the full text search all you have to do here
4:40:54 is you have to go and look at your full text policies.
4:40:58 You just need to define which language the data is in.
4:41:02 In post body that's where my data is for that description,
4:41:06 about the questions or the answers.
4:41:08 And I just create a full text index 1 with the post body.
4:41:12 After that, after re indexing is done,
4:41:15 all I have to do here is I can use the full text contains.
4:41:19 That's one of the functions for full text.
4:41:22 But this is one of the easiest ones.
4:41:24 So I'm looking for the questions.
4:41:27 I'M looking the post body and I'm looking
4:41:29 for if it has any SQL Server in the text.
4:41:32 So as you can see the post body is a pretty large text.
4:41:36 So if I'm going to go run this one
4:41:38 and if the user goes and writes SQL Server and search,
4:41:42 this is what's going to happen.
4:41:43 This is going to find top 10.
4:41:45 And as you can see the SQL server is in the post body out there.
4:41:49 I'm just giving the title here.
4:41:50 And the best part is look at the cost.
4:41:52 It really didn't cost us that much to make a search like that.
4:41:58 Next one is the vector search.
4:42:00 So the vector search in that one
4:42:04 user is really not looking specific words anymore.
4:42:06 It's looking maybe like looking
4:42:08 for a relational database operation problems for example.
4:42:12 So you have to Go and find MySQL, PostgreSQL and all kind of stuff.
4:42:15 So if that is the case,
4:42:17 the first thing you need to do for that is you just need to create no vectors.
4:42:21 And as you can see I created policies here.
4:42:24 You just need to give some information about the model that you pick.
4:42:28 And then this is my indexes,
4:42:32 interestingly there's other actual options here which I don't show it,
4:42:35 but you can actually see in the Azure portal
4:42:37 that is the shard ID for the vector index.
4:42:40 So in that case we have tenant one and two, right.
4:42:43 So you can easily create those indexes
4:42:46 and the indexes will distribute it by the tenant id.
4:42:49 You can actually write that.
4:42:50 So whenever you make the vector search,
4:42:52 it can search only that tenant ID vectors.
4:42:55 So it will make it much cheaper that way.
4:42:58 And DSCan is actually supporting that.
4:43:00 So that's why I kind of like the D scan indexes.
4:43:03 Anyway, before we go, you know more details, I go more complex.
4:43:06 Let's see.
4:43:08 Also this application is going to make the vector search much easier
4:43:12 because before you make the vector search you need to generate the vector,
4:43:16 sorry vector itself.
4:43:17 So as you can see, if you have a vector policy,
4:43:19 I have this vector search option here.
4:43:22 You just need to go and this is going
4:43:24 to go and discover all the AI policies you have.
4:43:28 And I have this embedding model for example.
4:43:30 And this is the data I want
4:43:32 to vectorize or search for relational database performance issues.
4:43:36 You click vectorize it, it calls and creates the embedding for you.
4:43:41 You can actually easily see it here.
4:43:43 So this is the text and this is the.
4:43:45 Well all that numbers is the vector I'm searching for.
4:43:49 So I can easily use this embedding parameter Before I
4:43:53 click on running that, I hope you notice one thing here.
4:43:56 Look at the document type.
4:43:58 It is post vector.
4:43:59 So I am not searching the whole database.
4:44:03 I'm only searching for document type.
4:44:04 So that actually makes my search much lighter.
4:44:07 And I'm passing the tenant ID too.
4:44:09 So I won't kind of go and search for farmer stuff I guess in this case.
4:44:14 So click on execute and this is going to cost a little
4:44:18 bit more because this is much more complex than looking for the word.
4:44:21 As you can see it find MySQL it find
4:44:25 WSQLDesigner insert problems dynamic SQL store procedures in Oracle.
4:44:30 It's much more complex than the other one that I just showed.
4:44:35 But it's available and you can easily use that too.
4:44:40 All right, the other three things that I
4:44:42 can just talk quickly is schema version.
4:44:46 You should have a property for the schema versioning out there
4:44:49 and every time the schema changes you can just increase that.
4:44:52 Then you can use that property to, if the business workflow changes
4:44:57 from schema version to other schema
4:44:59 version you can actually handle that retention.
4:45:02 You should really look into that because
4:45:04 Cosmos Signature have only the hot data.
4:45:07 If you don't care about that data anymore,
4:45:10 you can just take it out and put it somewhere much, you know, cheaper.
4:45:14 So that's going to make your data size smaller and that's going
4:45:16 to make your storage size much smaller and things going to get go,
4:45:20 you know, work much better for that.
4:45:22 You can use a TTL function of Cosmos DB for example.
4:45:26 If you don't care about the data after two years,
4:45:29 well you can easily make the TTL two years and Cosmos DB
4:45:33 automatically will delete your data after two years for free for you.
4:45:37 So that will be a great thing to ask to all the requirements,
4:45:42 whoever give you the requirement, what is your retention policies for that?
4:45:45 The last one is the analyzing data.
4:45:47 We create a great operational database,
4:45:50 data model but at the end somebody needs to analyze
4:45:54 this data so you will see how the business is doing.
4:45:58 So for that Cosmos DB has the mirroring data to fabric
4:46:03 and you can easily turn that on and data analyst
4:46:06 or BI developers can easily kind of go and look
4:46:09 at the data and analyze the data and create reports from there.
4:46:13 So those are the things that you should really
4:46:15 kind of consider and think about in the data model.
4:46:19 Well that's all I have for you today.
4:46:21 I hope everybody you know learns something new today.
4:46:24 If you have any questions,
4:46:26 you know how you can find me on the LinkedIn and Twitter or you know,
4:46:30 and other social media social platforms.
4:46:33 And I'll be more than happy to help you from those areas.
4:46:37 Thank you everyone.
4:46:38 Thank you for coming to my session.
4:46:46 At the intersection of science, artistry and design, Pantone gives designers,
4:46:51 brands and manufacturers a shared language to turn creative ideas into reality.
4:46:58 I'm Sky Kelly and I'm the president of Pantone.
4:47:01 Going beyond color standards, Pantone is about the business of creativity
4:47:05 and the trust and precision behind it.
4:47:08 Pantone harnesses decades of research,
4:47:11 being expertise and deep ties with the design community
4:47:14 to understand and predict what's new and what's next in design.
4:47:20 We really bring life to trends that shape culture.
4:47:24 When Pantone and Microsoft teamed up,
4:47:26 we built a system that translates those insights into actionable inspiration.
4:47:31 With the help of Azure AI Foundry, Azure Cosmos DB and GitHub Copilot,
4:47:36 we built AI agents that launch a new way
4:47:39 to engage with our platform and instantly unlock our expertise.