How video compression works - VLC lead developer explains | Lex Fridman Podcast
Lex Clips
0:02 So, the thing that we're talking about
0:05 is everything around video codecs, video encoding,
0:08 video decoding, video streaming,
0:10 video player client that I'm wearing on my head.
0:13 The entire ecosystem enabling free media.
0:16 We'll talk about FFmpeg, we'll talk about VideoLAN, VLC,
0:20 and all the other incredible video technology
0:23 that is used probably by billions of people.
0:26 So, JB you're the lead developer behind the legendary VLC player.
0:33 Kieran, amongst many other things,
0:34 you're the lead developer behind the legendary FFmpeg handle on Twitter.
0:39 And both of you have spicy opinions, I would say.
0:43 So, today we want to talk about FFmpeg and VLC.
0:48 For context for people who are not aware,
0:51 and I'm sure basically everybody listening to this have
0:54 used these two technologies probably regularly without knowing it.
1:01 So, FFmpeg underlies basically most video on the internet, including YouTube,
1:05 Netflix, Chrome, Firefox, of course VLC, and countless other video platforms.
1:11 It is estimated that over 90% of video
1:14 processing workflows online and offline involve FFmpeg.
1:18 VLC has been downloaded at least 6.5 billion times.
1:25 But likely that number, cuz it's impossible to really count the number,
1:29 is much higher than that.
1:31 Virtually any operating system supports virtually any media format.
1:39 The limitation being it can't open pancakes.
1:41 So, can we just lay out some of the basics
1:45 so to help people understand what's involved in all of this?
1:49 So, when we press play on a video player, like VLC, what happens?
1:54 What How does it go from the the file or the stream
1:59 to the pixels on the screen and the sound on the speaker?
2:02 What are the big stages to be aware of?
2:04 So, there are several stages, right?
2:06 The first stage is to get from an address, right?
2:09 Which is the type of URL to give you a bite of streams, right?
2:14 So, this would be, for example, HTTP, file, DVD, right?
2:18 You give the path to the media and give you a stream of data.
2:23 The stream needs to be cut up by what's known as the container,
2:25 the demultiplexer or demux.
2:28 Um we'll try and keep the jargon light throughout this, but um
2:30 it needs to go and start demarcating video and audio frames.
2:33 So, it just gets data from the operating system blocks at a time
2:36 and needs to start cutting these frames up into compressed data.
2:40 It then needs to start doing simple parsing of the video frames,
2:44 mainly to figure out whether that codec is
2:47 GPU decodable or needs to fall back to software.
2:50 We're very sort of used to assuming the GPU will play all of these things,
2:54 there'll be hardware acceleration.
2:56 I think it's up to 45% of files are not GPU decodable.
2:59 So, these need to be probed, they need to be detected.
3:02 There can be variants of a given codec, some of which are decodable on the GPU,
3:07 different vendors of GPU might have different capabilities.
3:10 So, those need to be detected.
3:12 So, if if it's GPU capable, you pass it through to the GPU black box.
3:16 So, now if there's a software fallback,
3:18 that means in the beginning is is to first do de-entropy coding,
3:22 so removing the mathematical coding of the bitstream.
3:25 So, this uses capabilities such as Huffman coding or arithmetic
3:29 coding to actually decompress the mathematical layer of the bitstream.
3:34 We then need to start reading the syntax elements for intra prediction.
3:37 So, intra prediction are like still images of the video, so your I-frames.
3:43 So, this works and operates in the spatial domain.
3:45 So, you do your intra prediction spatial domain, you you have a residual because
3:49 your prediction isn't quite matching that of reality.
3:52 So, you've made a prediction, but then there's a little bit left,
3:55 and that's what's known as the residual.
3:57 This is stored in the frequency domain,
3:59 and these are quantized to decompand their space.
4:02 We then need to do the inverse transform to bring
4:04 them back to the spatial domain and apply these residuals.
4:09 So, a lot of the process of the decoding is this thing is compressed.
4:13 Yes.
4:13 Yes.
4:13 And you have to predict the highest quality thing that's supposed to go there.
4:18 I-frame is the best representation you have spatially.
4:22 And then you And then there's a lot
4:24 of temporal compression that can happen depending on the codec.
4:27 And then you're predicting.
4:29 You're predicting what the reality that was captured in this rawest form.
4:33 Yeah, because what people don't realize is that the compression
4:36 on video and audio is one in the times, right?
4:40 Like, people don't realize how compressed we we do, right?
4:44 For audio, you move You compress by when you go from normal audio to MP3,
4:49 you compress by 10 times, right?
4:50 When when you move to video, you need one in the time, 200 times, right?
4:54 So, you need to remove all the details but that you
4:58 don't care about because all the compressions that we do,
5:01 and that's very important,
5:02 people forget about that, is to be viewed by humans, right?
5:05 So, all the codecs either for audio mimic basically how your ear works, right?
5:10 And and a lot of things about like the the the response on the ear,
5:14 and same for for your eyes, right?
5:16 And and so, for example, on video, we don't work on RGB, right?
5:20 Everyone expects to work in RGB.
5:22 We don't, right?
5:23 We move to YUV, which is basically one is luminance,
5:27 brightness, and the other are colors.
5:29 And this matches your eyes, where inside your eyes,
5:31 you have the cones and the buttons, right?
5:33 Where some of them look on brightness
5:35 and more on And the other on colors, right?
5:36 So, we need to compress a lot.
5:39 And so, we need to degrade, but in order to degrade,
5:42 we need to match the human perception.
5:44 And this is why it's so difficult.
5:46 And then we need to use the maximum power,
5:49 mathematical power, very complex technologies.
5:52 We move to the frequency domain as Kieran said.
5:54 We do a lot of dequantizing and in order to get the best compression,
6:00 but it still looks good.
6:02 You're trying to compress in order to maximize
6:05 the highest quality thing for human perception.
6:08 That is correct.
6:09 And that is correct.
6:09 And this is very important, right?
6:11 Compression is not like a zip, right?
6:13 A zip, you have data in, you get data out, right?
6:16 And you try with all the the zip compression to arrive with the limit.
6:21 Here, we are degrading the signal, right?
6:23 And so we need to degrade both the audio
6:25 and the video signal in the best way possible.
6:28 And we can do that, but it involves first
6:31 a lot of theoretical knowledge about how it works,
6:35 the eye works, but it a lot of mathematical change,
6:38 a lot of mathematical tricks, right?
6:40 For example, when you move to RGB and you you go to YUV, for example,
6:45 what we do very often is that we scale
6:48 down the resolution of the color compared to the brightness.
6:51 And most of the time and just this without compression,
6:54 it divides the size by two.
6:57 But most people don't see it, right?
6:59 Um and so on and so on, right?
7:01 And then you go to very complex mathematical change.
7:06 So of course Fourier transform, which de facto are not Fourier transform,
7:09 they are like discrete cosine transform, but that's the same idea.
7:13 So frequency domain, we split the video by blocks, right?
7:17 So that's why when it's wrongly decoded, you see those blocks and badly encoded,
7:21 you see those blocks and so on to arrive
7:24 to compression states that are insanely high, right?
7:28 And each generation of the codec is like 30% less Mhm.
7:32 for the same quality, right?
7:33 And this requires amount of power of computational power that are huge.
7:39 No, but you should you should elaborate.
7:40 It's 30% better, but an order of magnitude,
7:43 perhaps perhaps even two orders of magnitude more compression power.
7:47 That That's the big difference.
7:49 What do you mean by compression power?
7:50 So, CPU power to achieve that level of compression.
7:52 Oh, yeah.
7:53 So, you have to be able to leverage
7:54 the CPU and sometimes GPU like you mentioned.
7:57 And then we should mention that a lot
7:59 of this programming uh is done at the lowest possible stack.
8:04 Whether it's C and of course as as the legendary
8:08 Twitter handle um reemphasizes over and over a lot of assembly.
8:13 So, what happens is globally is that you have an address, right?
8:15 Which gives you uh with the operating system a stream of bytes,
8:19 a stream of data, right?
8:20 And this is the first step.
8:21 And the second step arise with demaxing where you're going to separate audio,
8:25 video, subtitle in type of different tracks.
8:28 And then on each of those tracks you're going to decompress them, decode them,
8:32 either audio with an audio codec,
8:34 video to video codec, and subtitle to subtitle codec.
8:37 Um and once you've decompressed those type of things, you have raw images,
8:41 raw and then you're going to talk to your uh
8:44 graphic card in your screen and display that.
8:46 And same for the audio, you're going to talk to your audio card,
8:49 which then is going to go um in analog to to your audio speakers.
8:53 And everything we've just said in the past couple of minutes,
8:57 every sentence is someone's lifetime's work.
8:58 There are books about every sentence.
9:00 So, the level of complexity in many cases is is inordinate.
9:04 You know, it's it's it's every sentence has thousands
9:08 of people working on this in in industry as a whole.
9:12 Books written about it.
9:13 So, there's a lot of detail, there's a lot of subtleties,
9:17 there's a lot of both academic and practical realities, um both of which matter.
9:23 Uh we're mentioning codecs, but I don't think you mentioned uh containers.
9:27 So, what what's the actual containers for some of the stuff we're
9:32 talking about so people are familiar with the MP4, uh MOV, MKV.
9:39 So, anyway, what what are containers versus uh the thing that goes inside?
9:44 So, the container is what we call also the muxer, right?
9:47 When I say demuxing, it means decontainerizing, right?
9:49 So, I actually if you look mux multiplexer and demultiplexer, right?
9:56 Mux and demux are those and same a codec is actually coder decoder, right?
10:01 Um, and um, so containers are these collection of multiple tracks, right?
10:07 So, it's a what normal people call the file format,
10:10 but it's a bit more um, subtle than that.
10:13 But, the most known one, of course, is MP4,
10:16 but uh, when I started it was AVI, right?
10:18 AVI was the the video format from from uh, Microsoft and uh,
10:23 move MOV, which became MP4, was a format from Apple.
10:27 Um, in the open source community, uh,
10:29 one of the person that is still active on VideoLAN is called uh,
10:31 Steve Lhomme and started this Matroska format,
10:34 which is like a bit more complex and and and more feature uh, proof.
10:38 Um, and um, there are so many others.
10:41 So, I mean, there's a it's a pretty common thing and maybe it'll
10:45 even happen in this conversation that people
10:46 confuse container and the codec, right?
10:50 So, they confuse MP4 and H.264, for example.
10:53 Is that a horrible violation?
10:55 No, it's not because technically the name of H.264 is
10:59 MPEG-4 Part 10 because MPEG-4 is actually a meta specification,
11:05 which has several things in it, right?
11:08 There is the Part 2.
11:10 Uh, so, there is like audio codecs, right?
11:12 AAC is the factor is MP4 audio something.
11:15 There is uh, actually several video codecs, right?
11:18 Inside the MPEG-4 specification, one of them is MPEG-4 Part 10,
11:22 called also AVC, called also H.264, right?
11:26 So, it's completely the fault of the industry
11:29 to to to to make things difficult to understand.
11:32 So, that's very difficult so that people then don't understand why sometimes you
11:36 talk about MPEG-4 Part 10 where you mean H.264 and why it's not MP4.
11:41 So, you can technically shove in all kinds
11:44 of different codecs inside containers and horribly so.
11:47 But, broadly speaking though,
11:49 MP4 is understood to generally be H.264 plus AAC audio.
11:54 99% of the time, that's that and that.
11:58 The rest are de minimis, they're smaller effects,
12:00 you know, edge effects really compared to that.
12:02 So, it's not the end of the world
12:03 that there are people who do get annoyed by that.
12:07 But, also in reality, something like VLC,
12:08 just to point out, the file may say dot mp4,
12:12 but it may be something completely different
12:13 and that's one of the challenges both FFmpeg and VLC have is the real world is
12:17 a completely different place to a three-letter file format.
12:20 And this is very important to say, right?
12:22 Like, for example, in VLC and in FFmpeg, we discard the file format, right?
12:27 We We look into the file to understand what's
12:30 in it because so many people, like, they say,
12:33 "Oh, it's a video, it must be MP4." But, technically,
12:35 it's an MOV or maybe it's a MKV, right?
12:38 So, we analyze in real time everything that we
12:42 have and we don't trust the the the format.
12:46 So, what information does the fact that it's dot mp4 give you?
12:49 It helps, right?
12:50 It gives you a hint, right?
12:51 Just like, "Oh, it It's finished by dot mp4.
12:54 I will start first by opening probing it with the MP4 container demuxer to say,
13:01 "Well, it should be that." But, I don't trust it and if I'm lost,
13:04 I say, "Okay, maybe I'm going to try to." So,
13:07 it bumps the priority of the module.
13:10 So, how do you get to, uh, just to take a bit of a tangent there, you know,
13:14 the dumb thing is if you try the MP4,
13:18 but it turns out a different codec than you would have expected,
13:22 uh, most players just break there.
13:25 Yes.
13:25 Yes.
13:26 So, how do you not break?
13:27 Cuz just uh philosophically, I'm sure there's a bunch of stumbling blocks along
13:31 the way where you it's easy to just break and stop, freak out, that's it.
13:36 How does VLC not?
13:37 This is why VLC is popular.
13:40 Um but the reason is because actually VLC was is just a client of a streaming
13:46 solution called VideoLAN from from from very long time ago, from the late '90s.
13:51 And when you playing video which are on UDP,
13:55 right, in network, they might be damaged, right?
13:58 So, you don't trust your inputs.
13:59 And this is very important in today's
14:00 security is that you don't trust your inputs.
14:02 So, everything in VLC is prepared to um work with broken files.
14:09 Mhm.
14:09 And it's a philosophical idea from the beginning,
14:13 and everything is engineered into that.
14:16 And and it's a culture, right?
14:17 And so, for example, and VLC became very popular on that because
14:21 a long time ago when people were uh pirating content,
14:24 um which they do a lot less today.
14:26 And none of us ever have.
14:28 No, of course not.
14:29 Um the metadata to play some files like AVI is at at the end of the file, right?
14:35 And when you're downloading, you don't have that, right?
14:37 So, VLC was just like, "Hey, this file is broken,
14:40 but I'm still going to try to interpret it." And this was very useful.
14:44 We hinted at the awesomeness of the various different stages.
14:48 We hinted at the awesomeness of codecs,
14:51 the depth and the richness and the complexity of everything involved there.
14:54 What Let's Let's try to define what is a video codec.
14:58 What What's involved there?
15:00 What What does it mean to compress something?
15:01 You already started to hint at it, but can can we elaborate a little bit more?
15:05 So, there's a huge amount of redundancy in any video,
15:08 uh both spatial and temporal.
15:10 And the point of any video codec is to remove this redundant data,
15:14 use mathematical properties as part of this reduction process.
15:17 So, more often than not using several
15:19 orders of magnitude more compute to compress because
15:22 that's more costly versus both costly both
15:24 financially and in CPU resources versus the decompression.
15:28 So, it's asymmetric in that respect.
15:31 Often the case because compression is done once,
15:33 but there could be lots of viewers of another file.
15:36 So, to take that information and compress it by 100x, 200x,
15:41 removing redundant information and using
15:43 mathematical properties to make that small,
15:45 but also have properties such as error resilience.
15:48 So, as as JB suggested VLC in the beginning
15:52 was was used to play UDP network feeds.
15:54 And UDP network feeds lose packets.
15:56 And so, some of the design goals of a codec is also to be recoverable.
16:01 You You need to actually be able to join the stream.
16:03 It's not necessarily a file.
16:04 You need to join get on the decoding process and start decoding.
16:09 And and to give it more image to to to to people who are not familiar, right?
16:14 Like when you're going to see any type of movie, right?
16:17 You're going to see the camera is going to pan, right?
16:20 And and travel.
16:21 And you realize that, for example,
16:23 all the background is the same from for like a minute, right?
16:26 Or 30 seconds, right?
16:27 So, you can reuse the cloud that you see on the background.
16:31 You can reuse that from a frame to another, right?
16:34 And so, it's gets the more the more memory you have,
16:39 the more power, the more comparisons you can make, right?
16:42 And so, the more compressed you can be.
16:44 And most of the modern codecs are basically doing that.
16:47 So, just to make it even more explicit.
16:50 So, what is video?
16:52 Video is a bunch of pixels of an RGB of three
16:58 values and you have a grid of pixels and you have,
17:01 let's say, 24 or 30 or 60 of frames a second.
17:07 And you just have all these pixels repeating
17:10 and showing different stuff 30 times a second.
17:13 And so, the question, the philosophical,
17:16 the technical question is how can I compress all
17:19 of that, store all of that at 100 x?
17:24 Or 1,000 x, right?
17:25 1,000 x.
17:26 The target is 1,000 x, right?
17:27 And the goal is when you say redundancy, what is redundant meaning stuff at best
17:36 that humans wouldn't notice if it was missing.
17:38 So, for example, you have a picture of a cloud, right?
17:41 And from the next frame there's still going to be the same cloud.
17:44 So, it's redundant.
17:45 You could just put it once and not do it, right?
17:47 Or you have a a black background behind me, for example,
17:50 the black is the same on the whole picture, right?
17:52 So, you can say, "Well, you know, in this picture,
17:55 take the pixels that you have on the top left and the one on the top right.
17:59 I'm not going to give the value.
18:00 I'm just going to tell you it's the same
18:01 as the top left." And then you can say for frame one,
18:05 um reuse something from the previous frame or the previous
18:08 previous frame and so on and so on, right?
18:10 So, you could basically it's unlimited,
18:14 but then it's limited in terms of memory or in terms of compute power because,
18:19 for example, if you need to compare pixels
18:21 on 200 frames in the past on 4K resolutions, it's a huge amount of compute.
18:29 And then when you're showing it, you have to do the decompress of all of that.
18:33 So, you is it the codec that has the encoding
18:37 and the decoding is a is a coupled process that you're developing?
18:41 Exactly, right?
18:41 And those are two different um uh tradeoffs, right?
18:45 Are you going to compress more,
18:47 uh but then it might be more difficult to to to decode?
18:51 Um are you going to comp to make it a codec
18:54 that is more complex to encode and easier to decode?
18:57 Are you going to make a codec that is
18:58 easier to encode because you need to be fast,
19:00 but then the the client side, the player is going to spend more time?
19:04 That's why you have so many different type of codecs.
19:06 Is that it's not always easy.
19:09 A- A- And to make it even more complex,
19:11 modern like AV1, AV2, or VVC are actually not codec's.
19:16 They are a collection of tools, right?
19:18 They are multiple tools,
19:19 multiple codec's in the same codec to depending on the image,
19:23 get the more compression.
19:24 So, just to elaborate, codec's like AV1,
19:28 VVC have a much wide have a wide audience.
19:31 It could be a screen share content, it could be video, it could be animation.
19:36 All of these require different coding tools.
19:40 So, what happens these days is a collection of tools are put in and called AV1,
19:46 called AV2, called VVC to allow for different use cases.
19:49 So, you may be on Zoom and sharing your PowerPoint,
19:53 and then you need to show the audience a video.
19:55 That codec needs to start changing its tool set
19:59 depending on the content to compress in a different way.
20:02 And like you said, there's a a bunch of incredible engineers behind each
20:06 part of that, each part of the tools that make up AV1, for example.
20:09 Sure.