How video compression works - VLC lead developer explains | Lex Fridman Podcast

How video compression works - VLC lead developer explains | Lex Fridman Podcast

Lex Clips

0:02 So, the thing that we're talking about

0:05 is everything around video codecs, video encoding,

0:08 video decoding, video streaming,

0:10 video player client that I'm wearing on my head.

0:13 The entire ecosystem enabling free media.

0:16 We'll talk about FFmpeg, we'll talk about VideoLAN, VLC,

0:20 and all the other incredible video technology

0:23 that is used probably by billions of people.

0:26 So, JB you're the lead developer behind the legendary VLC player.

0:33 Kieran, amongst many other things,

0:34 you're the lead developer behind the legendary FFmpeg handle on Twitter.

0:39 And both of you have spicy opinions, I would say.

0:43 So, today we want to talk about FFmpeg and VLC.

0:48 For context for people who are not aware,

0:51 and I'm sure basically everybody listening to this have

0:54 used these two technologies probably regularly without knowing it.

1:01 So, FFmpeg underlies basically most video on the internet, including YouTube,

1:05 Netflix, Chrome, Firefox, of course VLC, and countless other video platforms.

1:11 It is estimated that over 90% of video

1:14 processing workflows online and offline involve FFmpeg.

1:18 VLC has been downloaded at least 6.5 billion times.

1:25 But likely that number, cuz it's impossible to really count the number,

1:29 is much higher than that.

1:31 Virtually any operating system supports virtually any media format.

1:39 The limitation being it can't open pancakes.

1:41 So, can we just lay out some of the basics

1:45 so to help people understand what's involved in all of this?

1:49 So, when we press play on a video player, like VLC, what happens?

1:54 What How does it go from the the file or the stream

1:59 to the pixels on the screen and the sound on the speaker?

2:02 What are the big stages to be aware of?

2:04 So, there are several stages, right?

2:06 The first stage is to get from an address, right?

2:09 Which is the type of URL to give you a bite of streams, right?

2:14 So, this would be, for example, HTTP, file, DVD, right?

2:18 You give the path to the media and give you a stream of data.

2:23 The stream needs to be cut up by what's known as the container,

2:25 the demultiplexer or demux.

2:28 Um we'll try and keep the jargon light throughout this, but um

2:30 it needs to go and start demarcating video and audio frames.

2:33 So, it just gets data from the operating system blocks at a time

2:36 and needs to start cutting these frames up into compressed data.

2:40 It then needs to start doing simple parsing of the video frames,

2:44 mainly to figure out whether that codec is

2:47 GPU decodable or needs to fall back to software.

2:50 We're very sort of used to assuming the GPU will play all of these things,

2:54 there'll be hardware acceleration.

2:56 I think it's up to 45% of files are not GPU decodable.

2:59 So, these need to be probed, they need to be detected.

3:02 There can be variants of a given codec, some of which are decodable on the GPU,

3:07 different vendors of GPU might have different capabilities.

3:10 So, those need to be detected.

3:12 So, if if it's GPU capable, you pass it through to the GPU black box.

3:16 So, now if there's a software fallback,

3:18 that means in the beginning is is to first do de-entropy coding,

3:22 so removing the mathematical coding of the bitstream.

3:25 So, this uses capabilities such as Huffman coding or arithmetic

3:29 coding to actually decompress the mathematical layer of the bitstream.

3:34 We then need to start reading the syntax elements for intra prediction.

3:37 So, intra prediction are like still images of the video, so your I-frames.

3:43 So, this works and operates in the spatial domain.

3:45 So, you do your intra prediction spatial domain, you you have a residual because

3:49 your prediction isn't quite matching that of reality.

3:52 So, you've made a prediction, but then there's a little bit left,

3:55 and that's what's known as the residual.

3:57 This is stored in the frequency domain,

3:59 and these are quantized to decompand their space.

4:02 We then need to do the inverse transform to bring

4:04 them back to the spatial domain and apply these residuals.

4:09 So, a lot of the process of the decoding is this thing is compressed.

4:13 Yes.

4:13 Yes.

4:13 And you have to predict the highest quality thing that's supposed to go there.

4:18 I-frame is the best representation you have spatially.

4:22 And then you And then there's a lot

4:24 of temporal compression that can happen depending on the codec.

4:27 And then you're predicting.

4:29 You're predicting what the reality that was captured in this rawest form.

4:33 Yeah, because what people don't realize is that the compression

4:36 on video and audio is one in the times, right?

4:40 Like, people don't realize how compressed we we do, right?

4:44 For audio, you move You compress by when you go from normal audio to MP3,

4:49 you compress by 10 times, right?

4:50 When when you move to video, you need one in the time, 200 times, right?

4:54 So, you need to remove all the details but that you

4:58 don't care about because all the compressions that we do,

5:01 and that's very important,

5:02 people forget about that, is to be viewed by humans, right?

5:05 So, all the codecs either for audio mimic basically how your ear works, right?

5:10 And and a lot of things about like the the the response on the ear,

5:14 and same for for your eyes, right?

5:16 And and so, for example, on video, we don't work on RGB, right?

5:20 Everyone expects to work in RGB.

5:22 We don't, right?

5:23 We move to YUV, which is basically one is luminance,

5:27 brightness, and the other are colors.

5:29 And this matches your eyes, where inside your eyes,

5:31 you have the cones and the buttons, right?

5:33 Where some of them look on brightness

5:35 and more on And the other on colors, right?

5:36 So, we need to compress a lot.

5:39 And so, we need to degrade, but in order to degrade,

5:42 we need to match the human perception.

5:44 And this is why it's so difficult.

5:46 And then we need to use the maximum power,

5:49 mathematical power, very complex technologies.

5:52 We move to the frequency domain as Kieran said.

5:54 We do a lot of dequantizing and in order to get the best compression,

6:00 but it still looks good.

6:02 You're trying to compress in order to maximize

6:05 the highest quality thing for human perception.

6:08 That is correct.

6:09 And that is correct.

6:09 And this is very important, right?

6:11 Compression is not like a zip, right?

6:13 A zip, you have data in, you get data out, right?

6:16 And you try with all the the zip compression to arrive with the limit.

6:21 Here, we are degrading the signal, right?

6:23 And so we need to degrade both the audio

6:25 and the video signal in the best way possible.

6:28 And we can do that, but it involves first

6:31 a lot of theoretical knowledge about how it works,

6:35 the eye works, but it a lot of mathematical change,

6:38 a lot of mathematical tricks, right?

6:40 For example, when you move to RGB and you you go to YUV, for example,

6:45 what we do very often is that we scale

6:48 down the resolution of the color compared to the brightness.

6:51 And most of the time and just this without compression,

6:54 it divides the size by two.

6:57 But most people don't see it, right?

6:59 Um and so on and so on, right?

7:01 And then you go to very complex mathematical change.

7:06 So of course Fourier transform, which de facto are not Fourier transform,

7:09 they are like discrete cosine transform, but that's the same idea.

7:13 So frequency domain, we split the video by blocks, right?

7:17 So that's why when it's wrongly decoded, you see those blocks and badly encoded,

7:21 you see those blocks and so on to arrive

7:24 to compression states that are insanely high, right?

7:28 And each generation of the codec is like 30% less Mhm.

7:32 for the same quality, right?

7:33 And this requires amount of power of computational power that are huge.

7:39 No, but you should you should elaborate.

7:40 It's 30% better, but an order of magnitude,

7:43 perhaps perhaps even two orders of magnitude more compression power.

7:47 That That's the big difference.

7:49 What do you mean by compression power?

7:50 So, CPU power to achieve that level of compression.

7:52 Oh, yeah.

7:53 So, you have to be able to leverage

7:54 the CPU and sometimes GPU like you mentioned.

7:57 And then we should mention that a lot

7:59 of this programming uh is done at the lowest possible stack.

8:04 Whether it's C and of course as as the legendary

8:08 Twitter handle um reemphasizes over and over a lot of assembly.

8:13 So, what happens is globally is that you have an address, right?

8:15 Which gives you uh with the operating system a stream of bytes,

8:19 a stream of data, right?

8:20 And this is the first step.

8:21 And the second step arise with demaxing where you're going to separate audio,

8:25 video, subtitle in type of different tracks.

8:28 And then on each of those tracks you're going to decompress them, decode them,

8:32 either audio with an audio codec,

8:34 video to video codec, and subtitle to subtitle codec.

8:37 Um and once you've decompressed those type of things, you have raw images,

8:41 raw and then you're going to talk to your uh

8:44 graphic card in your screen and display that.

8:46 And same for the audio, you're going to talk to your audio card,

8:49 which then is going to go um in analog to to your audio speakers.

8:53 And everything we've just said in the past couple of minutes,

8:57 every sentence is someone's lifetime's work.

8:58 There are books about every sentence.

9:00 So, the level of complexity in many cases is is inordinate.

9:04 You know, it's it's it's every sentence has thousands

9:08 of people working on this in in industry as a whole.

9:12 Books written about it.

9:13 So, there's a lot of detail, there's a lot of subtleties,

9:17 there's a lot of both academic and practical realities, um both of which matter.

9:23 Uh we're mentioning codecs, but I don't think you mentioned uh containers.

9:27 So, what what's the actual containers for some of the stuff we're

9:32 talking about so people are familiar with the MP4, uh MOV, MKV.

9:39 So, anyway, what what are containers versus uh the thing that goes inside?

9:44 So, the container is what we call also the muxer, right?

9:47 When I say demuxing, it means decontainerizing, right?

9:49 So, I actually if you look mux multiplexer and demultiplexer, right?

9:56 Mux and demux are those and same a codec is actually coder decoder, right?

10:01 Um, and um, so containers are these collection of multiple tracks, right?

10:07 So, it's a what normal people call the file format,

10:10 but it's a bit more um, subtle than that.

10:13 But, the most known one, of course, is MP4,

10:16 but uh, when I started it was AVI, right?

10:18 AVI was the the video format from from uh, Microsoft and uh,

10:23 move MOV, which became MP4, was a format from Apple.

10:27 Um, in the open source community, uh,

10:29 one of the person that is still active on VideoLAN is called uh,

10:31 Steve Lhomme and started this Matroska format,

10:34 which is like a bit more complex and and and more feature uh, proof.

10:38 Um, and um, there are so many others.

10:41 So, I mean, there's a it's a pretty common thing and maybe it'll

10:45 even happen in this conversation that people

10:46 confuse container and the codec, right?

10:50 So, they confuse MP4 and H.264, for example.

10:53 Is that a horrible violation?

10:55 No, it's not because technically the name of H.264 is

10:59 MPEG-4 Part 10 because MPEG-4 is actually a meta specification,

11:05 which has several things in it, right?

11:08 There is the Part 2.

11:10 Uh, so, there is like audio codecs, right?

11:12 AAC is the factor is MP4 audio something.

11:15 There is uh, actually several video codecs, right?

11:18 Inside the MPEG-4 specification, one of them is MPEG-4 Part 10,

11:22 called also AVC, called also H.264, right?

11:26 So, it's completely the fault of the industry

11:29 to to to to make things difficult to understand.

11:32 So, that's very difficult so that people then don't understand why sometimes you

11:36 talk about MPEG-4 Part 10 where you mean H.264 and why it's not MP4.

11:41 So, you can technically shove in all kinds

11:44 of different codecs inside containers and horribly so.

11:47 But, broadly speaking though,

11:49 MP4 is understood to generally be H.264 plus AAC audio.

11:54 99% of the time, that's that and that.

11:58 The rest are de minimis, they're smaller effects,

12:00 you know, edge effects really compared to that.

12:02 So, it's not the end of the world

12:03 that there are people who do get annoyed by that.

12:07 But, also in reality, something like VLC,

12:08 just to point out, the file may say dot mp4,

12:12 but it may be something completely different

12:13 and that's one of the challenges both FFmpeg and VLC have is the real world is

12:17 a completely different place to a three-letter file format.

12:20 And this is very important to say, right?

12:22 Like, for example, in VLC and in FFmpeg, we discard the file format, right?

12:27 We We look into the file to understand what's

12:30 in it because so many people, like, they say,

12:33 "Oh, it's a video, it must be MP4." But, technically,

12:35 it's an MOV or maybe it's a MKV, right?

12:38 So, we analyze in real time everything that we

12:42 have and we don't trust the the the format.

12:46 So, what information does the fact that it's dot mp4 give you?

12:49 It helps, right?

12:50 It gives you a hint, right?

12:51 Just like, "Oh, it It's finished by dot mp4.

12:54 I will start first by opening probing it with the MP4 container demuxer to say,

13:01 "Well, it should be that." But, I don't trust it and if I'm lost,

13:04 I say, "Okay, maybe I'm going to try to." So,

13:07 it bumps the priority of the module.

13:10 So, how do you get to, uh, just to take a bit of a tangent there, you know,

13:14 the dumb thing is if you try the MP4,

13:18 but it turns out a different codec than you would have expected,

13:22 uh, most players just break there.

13:25 Yes.

13:25 Yes.

13:26 So, how do you not break?

13:27 Cuz just uh philosophically, I'm sure there's a bunch of stumbling blocks along

13:31 the way where you it's easy to just break and stop, freak out, that's it.

13:36 How does VLC not?

13:37 This is why VLC is popular.

13:40 Um but the reason is because actually VLC was is just a client of a streaming

13:46 solution called VideoLAN from from from very long time ago, from the late '90s.

13:51 And when you playing video which are on UDP,

13:55 right, in network, they might be damaged, right?

13:58 So, you don't trust your inputs.

13:59 And this is very important in today's

14:00 security is that you don't trust your inputs.

14:02 So, everything in VLC is prepared to um work with broken files.

14:09 Mhm.

14:09 And it's a philosophical idea from the beginning,

14:13 and everything is engineered into that.

14:16 And and it's a culture, right?

14:17 And so, for example, and VLC became very popular on that because

14:21 a long time ago when people were uh pirating content,

14:24 um which they do a lot less today.

14:26 And none of us ever have.

14:28 No, of course not.

14:29 Um the metadata to play some files like AVI is at at the end of the file, right?

14:35 And when you're downloading, you don't have that, right?

14:37 So, VLC was just like, "Hey, this file is broken,

14:40 but I'm still going to try to interpret it." And this was very useful.

14:44 We hinted at the awesomeness of the various different stages.

14:48 We hinted at the awesomeness of codecs,

14:51 the depth and the richness and the complexity of everything involved there.

14:54 What Let's Let's try to define what is a video codec.

14:58 What What's involved there?

15:00 What What does it mean to compress something?

15:01 You already started to hint at it, but can can we elaborate a little bit more?

15:05 So, there's a huge amount of redundancy in any video,

15:08 uh both spatial and temporal.

15:10 And the point of any video codec is to remove this redundant data,

15:14 use mathematical properties as part of this reduction process.

15:17 So, more often than not using several

15:19 orders of magnitude more compute to compress because

15:22 that's more costly versus both costly both

15:24 financially and in CPU resources versus the decompression.

15:28 So, it's asymmetric in that respect.

15:31 Often the case because compression is done once,

15:33 but there could be lots of viewers of another file.

15:36 So, to take that information and compress it by 100x, 200x,

15:41 removing redundant information and using

15:43 mathematical properties to make that small,

15:45 but also have properties such as error resilience.

15:48 So, as as JB suggested VLC in the beginning

15:52 was was used to play UDP network feeds.

15:54 And UDP network feeds lose packets.

15:56 And so, some of the design goals of a codec is also to be recoverable.

16:01 You You need to actually be able to join the stream.

16:03 It's not necessarily a file.

16:04 You need to join get on the decoding process and start decoding.

16:09 And and to give it more image to to to to people who are not familiar, right?

16:14 Like when you're going to see any type of movie, right?

16:17 You're going to see the camera is going to pan, right?

16:20 And and travel.

16:21 And you realize that, for example,

16:23 all the background is the same from for like a minute, right?

16:26 Or 30 seconds, right?

16:27 So, you can reuse the cloud that you see on the background.

16:31 You can reuse that from a frame to another, right?

16:34 And so, it's gets the more the more memory you have,

16:39 the more power, the more comparisons you can make, right?

16:42 And so, the more compressed you can be.

16:44 And most of the modern codecs are basically doing that.

16:47 So, just to make it even more explicit.

16:50 So, what is video?

16:52 Video is a bunch of pixels of an RGB of three

16:58 values and you have a grid of pixels and you have,

17:01 let's say, 24 or 30 or 60 of frames a second.

17:07 And you just have all these pixels repeating

17:10 and showing different stuff 30 times a second.

17:13 And so, the question, the philosophical,

17:16 the technical question is how can I compress all

17:19 of that, store all of that at 100 x?

17:24 Or 1,000 x, right?

17:25 1,000 x.

17:26 The target is 1,000 x, right?

17:27 And the goal is when you say redundancy, what is redundant meaning stuff at best

17:36 that humans wouldn't notice if it was missing.

17:38 So, for example, you have a picture of a cloud, right?

17:41 And from the next frame there's still going to be the same cloud.

17:44 So, it's redundant.

17:45 You could just put it once and not do it, right?

17:47 Or you have a a black background behind me, for example,

17:50 the black is the same on the whole picture, right?

17:52 So, you can say, "Well, you know, in this picture,

17:55 take the pixels that you have on the top left and the one on the top right.

17:59 I'm not going to give the value.

18:00 I'm just going to tell you it's the same

18:01 as the top left." And then you can say for frame one,

18:05 um reuse something from the previous frame or the previous

18:08 previous frame and so on and so on, right?

18:10 So, you could basically it's unlimited,

18:14 but then it's limited in terms of memory or in terms of compute power because,

18:19 for example, if you need to compare pixels

18:21 on 200 frames in the past on 4K resolutions, it's a huge amount of compute.

18:29 And then when you're showing it, you have to do the decompress of all of that.

18:33 So, you is it the codec that has the encoding

18:37 and the decoding is a is a coupled process that you're developing?

18:41 Exactly, right?

18:41 And those are two different um uh tradeoffs, right?

18:45 Are you going to compress more,

18:47 uh but then it might be more difficult to to to decode?

18:51 Um are you going to comp to make it a codec

18:54 that is more complex to encode and easier to decode?

18:57 Are you going to make a codec that is

18:58 easier to encode because you need to be fast,

19:00 but then the the client side, the player is going to spend more time?

19:04 That's why you have so many different type of codecs.

19:06 Is that it's not always easy.

19:09 A- A- And to make it even more complex,

19:11 modern like AV1, AV2, or VVC are actually not codec's.

19:16 They are a collection of tools, right?

19:18 They are multiple tools,

19:19 multiple codec's in the same codec to depending on the image,

19:23 get the more compression.

19:24 So, just to elaborate, codec's like AV1,

19:28 VVC have a much wide have a wide audience.

19:31 It could be a screen share content, it could be video, it could be animation.

19:36 All of these require different coding tools.

19:40 So, what happens these days is a collection of tools are put in and called AV1,

19:46 called AV2, called VVC to allow for different use cases.

19:49 So, you may be on Zoom and sharing your PowerPoint,

19:53 and then you need to show the audience a video.

19:55 That codec needs to start changing its tool set

19:59 depending on the content to compress in a different way.

20:02 And like you said, there's a a bunch of incredible engineers behind each

20:06 part of that, each part of the tools that make up AV1, for example.

20:09 Sure.

Study with Looplines Download Captions Watch on YouTube