Using YouTube as Cloud Storage
No blog post today, I recommend watching this video! I created this fun project, which touches upon domains in video compression, networking, and encryption. Attached underneath is the transcript of the video.
[00:00:00] you know that feeling when your hard drive gets full and now you have to start debating on which files
[00:00:04] to keep and which to delete You might want to use Google Drive Dropbox or iCloud But no today I'm
[00:00:11] going to show you how to use YouTube to store your files
[00:00:20] You might be asking why Well you can upload a ton of videos to YouTube and each YouTube video can
[00:00:25] be up to 256 GB or 12 hours whichever limit you hit first Instead of paying hundreds of dollars for
[00:00:33] I don't know iCloud why not just capitalize on YouTube's file storage system It should be easy right As long
[00:00:39] as we convert our file into a video we can upload Well the first thing we want to think about
[00:00:45] is how we're going to store our file data into a video container When you look at any video file
[00:00:50] you'll notice it ends with ampp4 mkvov or something similar This is called a video container because it holds all
[00:00:59] the video audio and subtitle tracks that are needed for you to play the video Things like authors descriptions and
[00:01:06] thumbnails can also live inside the container Basically it's just a portable bundle that media players know how to decode
[00:01:14] And that leaves us with a lot of options for storing our file data We could store it in the
[00:01:19] video the audio the subtitles the metadata and honestly anything However we actually can't use most of these options If
[00:01:28] our end goal is storing files on YouTube YouTube actually gets rid of most of these attributes when you upload
[00:01:33] a video For example metadata like the authors and description are stripped out of the file because they aren't relevant
[00:01:39] to YouTube Subtitles aren't stripped out but they could be rejected by YouTube if they are invalid or store too
[00:01:45] much data for the subtitle engine to process That leaves us with two main options video and audio YouTube obviously
[00:01:53] conceptually keeps both of these tracks so that you're able to see and hear the video you're watching And both
[00:01:59] are completely viable But which one should we pick Well let's take an example What is audio Audio is made
[00:02:06] up of sound waves And to represent them digitally we take many samples and measure each signal's amplitude at a
[00:02:13] specific moment in time During playback speakers or headphones use these samples to reconstruct the waveform we perceive as sound
[00:02:22] And some systems might use one channel while others use two channels allowing separate samples to play through the left
[00:02:28] and right ears For example a CD quality stereo audio sampled at 44,100 samples per second with 16 bits each
[00:02:36] sample requires about 1.4 megabits per second of data Video on the other hand is stored in an array of
[00:02:43] pixels For instance a 1080p video is 1,920 pixels wide and 1,080 pixels tall or around 2 million pixels in
[00:02:52] total Each pixel let's say contains R G and B values that are 8 bits for each color channel or
[00:02:59] 24 bits in total That means that each frame of a video is 48 million bits And at a standard
[00:03:06] frame rate like 30 frames per second this means that we are displaying over 1,500 megabits per second of data
[00:03:12] which when uncompressed is over 1,000 times more data than audio It's clear that video is the better option but
[00:03:19] that doesn't rule out audio completely We could theoretically use both at the same time but the amount of data
[00:03:25] that video gives us completely dominates audio to the point where it's not even worth coordinating both at the same
[00:03:31] time So for now we're just going to stick with video to store our file data From here on it
[00:03:37] feels easy Since our file is just bytes we can encode those bytes directly into pixels If we are using
[00:03:43] RGB this means that we are using three bytes one for R one for G and one for B That
[00:03:50] way we can split our file into three bytes segments and just encode each pixel with our file data But
[00:03:56] in reality the situation is much different Earlier I said YouTube keeps the audio and video tracks conceptually I assumed
[00:04:05] that meant it stored the exact file we uploaded but that's not actually what happens In order for YouTube to
[00:04:11] support everyone's massive uploads they compress our videos So when you upload a video YouTube re-encodes it into different resolutions
[00:04:19] and formats so it can stream efficiently on different devices and bandwidths You may have seen this when the video
[00:04:25] quality gets lower when you have bad internet There are many types of video compression but the most important part
[00:04:33] of this is that YouTube's compression is lossy So what does that mean Well it means that YouTube actually tampers
[00:04:40] of your video file data to save space For example small details in your videos might get removed causing pixel
[00:04:47] colors and therefore their bite data to change This causes the video you see to not be the identical video
[00:04:53] you uploaded So this is where our plan starts to fall apart Our approach only works if YouTube keeps our
[00:05:02] video data intact so we can recover it later But YouTube re-encodes everything using lossy compression which means parts of
[00:05:09] the original data are permanently discarded And even worse their compression algorithm isn't public so we don't even know exactly
[00:05:16] what's being changed It's like tossing our file into a black box and then getting back something that only somewhat
[00:05:22] resembles the original file Now this is where most people would give up But I made you guys a promise
[00:05:29] so quitting wasn't an option If YouTube insists on recompressing everything we upload then maybe the solution isn't to fight
[00:05:36] it but instead design around it So let's take a step back Earlier we were writing raw bytes directly into
[00:05:43] pixels splitting the data into tiny three byte chunks But that makes every small change from compression disastrous because even
[00:05:51] a single corrupted pixel can destroy the data it carries So instead of a singular pixel let's think bigger Instead
[00:05:58] of tiny pieces what if we group the data into much larger sections Let's say like 1 megabyte blocks Now
[00:06:05] each big block becomes something we can protect and verify But how can we protect and verify if each block
[00:06:13] is fine or not To think about this I'd like to bring up something surprisingly familiar Something you have definitely
[00:06:21] seen before From posters of lost pets to shipping products Do you know what they are QR codes aren't just
[00:06:29] black and white squares that you can scan They're genius technology designed to survive damage You can scratch them smudge
[00:06:37] them tear them or even print them horribly like the diagrams on your exam but they'd still work So how's
[00:06:43] that even possible Well when you look at a QR code they don't just store the raw data They deliberately
[00:06:50] store extra data alongside it These additional bites act like built-in protection which detects when part of the code has
[00:06:57] been damaged They can also in many cases reconstruct what was lost even when parts are missing using an algorithm
[00:07:04] called read Solomon
[00:07:09] This entire process is called redundancy which is something you see not just in computer science but in aircrafts power
[00:07:17] systems and many other real life applications We can borrow that redundant design here If we treat YouTube's compression as
[00:07:25] a noisy lossy layer like the QR codes subject to damage then each block we store can have its own
[00:07:31] protection But in order to do that we would have to answer two main questions One has this block been
[00:07:38] corrupted by compression And two if it has do we still have enough redundant information to reconstruct the original data
[00:07:46] We want to design our system to assume damage like QR codes so it can be resilient to the video
[00:07:52] compression Now you might be asking why can't we use QR codes directly for each big block we are trying
[00:07:59] to keep errorprone and your question would be valid why shouldn't we use something that has been proven to work
[00:08:06] millions of times in our video storage example the truth is QR codes are overengineered for our specific video compression
[00:08:14] issue we're subject to face different challenges like scratches and tears which are things we wouldn't experience with YouTube's video
[00:08:21] compression in also encoding one QR code for every chunk of our file uses a lot of unnecessary storage which
[00:08:28] doesn't make it an ideal way of storing our blocks Let's go back to the drawing board and address each
[00:08:36] question with a different solution Then our first question is how we can tell if a block has been corrupted
[00:08:41] by compression Well we've actually solved this problem before already except in a different place Network connections face this challenge
[00:08:50] on a daily basis When you download documents videos or any file your connection could be unreliable due to weather
[00:08:57] outages or anything to be honest This could cause certain chunks of your file to be corrupted during transport
[00:09:08] And to fix this networks use something called a cyclic redundancy check or CRC to detect damaged chunks CRC uses
[00:09:17] incredibly clever math to generate a fingerprint for every chunk of data
[00:09:31] Under the hood CRC treats your data like one long stream of bits and runs it through a special division
[00:09:37] process to generate a fingerprint But this isn't normal division with subtraction and carries Instead everything happens in binary using
[00:09:45] Exor which makes the math super fast for computers You can think of it like feeding your data through a
[00:09:51] fixed pattern called a generator And whatever remainder is left over from the division becomes the check sum The remainder
[00:10:00] is strongly tied to the exact bits of our original data So even the tiniest change completely alters the result
[00:10:22] For example suppose we want to send the chunk 11 1 We pick a small generator pattern which is basically
[00:10:30] a constant that determines how many CRC bits we want And for simplicity let's use 1011 as our generator pattern
[00:10:38] First we append a few zeros to the data so we can make space for our remainder This gives us
[00:10:43] 1 1 0 1 0 0 0 Then we perform this xor base division lining up the generator and exoring
[00:10:51] whatever the leading bit is one just like lawn division After a few rounds of shifting and exoring we're left
[00:10:58] with a small remainder Let's just say 001 That remainder becomes our CRC So the value we actually transmit is
[00:11:05] 1 1 0 1 0 0 1
[00:11:10] When the receiver gets this they perform the same division again And let's see what happens if we do it
[00:11:16] on intact data We get a remainder of zero But now let's change a singular bit and do the same
[00:11:25] division again
[00:11:31] The remainder isn't zero anymore even when one value changed This is the secret magic of CRC Tiny changes in
[00:11:39] the data produce completely different fingerprints So instead of checking every bite we can get a super fast mathematical guarantee
[00:11:46] that the block is either intact or corrupted But there's a catch CRC isn't perfect If the corruption itself happens
[00:11:54] to divide evenly by the generator number it would give us a remainder of zero which would give us the
[00:11:59] perception that the block is intact
[00:12:09] In reality though it's super rare especially with a good CRC In my project I use CRC 32 for my
[00:12:16] chunks for error detection which for random corruption has a 1 and 2 to the^ of 32 or 1 in
[00:12:22] 4.3 billion chance of being wrong That's like picking one breasted out of half of Earth's population blindfolded and getting
[00:12:30] it right the first try It's so small that it's essentially negligible for our use case
[00:12:38] CRC is great at catching these errors but it doesn't answer our second design question on whether or not we
[00:12:44] can reconstruct our original data if there is corruption To solve this we need another algorithm on top of CRC
[00:12:51] that's able to recover our data There's many forward or error correction algorithms we could pick like Raptor wheat Solomon
[00:12:59] used by cure codes LDPC but I opted for using a library called wirehair I chose YARE because it scales
[00:13:06] really well for larger files Behind the scenes Yhair uses something called a fountain code which is basically when you
[00:13:13] break your data into small pieces and then generate extra repair chunks for each piece Each repair chunk isn't a
[00:13:20] copy but a smart bit wise mixing of the original chunks together You can think of it like a system
[00:13:26] of equations One chunk might represent A X or C another B X or D another A X or B
[00:13:32] X or D On their own they don't mean much but altogether they form a system of equations that you
[00:13:38] can solve So if some chunks get lost or corrupted which we expect after YouTube's video compression we're able to
[00:13:45] solve and reconstruct our original data as long as we'd have enough total pieces
[00:14:07] Now in reality wire hair is incredibly complex and is filled with beautiful linear algebra proofs that allow it to
[00:14:13] be successful I have attached Wire Hair's repository below where you can see the code for reference if you'd like
[00:14:19] to dive deeper Combining these two methods together we're able to finally create a design that both checks for errors
[00:14:26] within each chunk and can recover if certain parts are lost after video compression In my code I use both
[00:14:33] of these algorithms alongside headers I've created to communicate information from the encoder to the decoder
[00:14:41] However this still doesn't explain how we actually store data inside a video in the first place which turns out
[00:14:47] to be its own challenge Previously we used pixels to represent file data But what can we do with the
[00:14:53] errorproof blocks when encoding them into the actual video file itself Well when encoding our video file with ffmpeg we
[00:15:02] have to keep a few important things in mind First YouTube is going to recompress whatever we upload And second
[00:15:10] even though we don't know how YouTube compression works we know that in general video compression isn't random and that
[00:15:17] most modern codecs don't operate on the raw pixels itself For user compression they often transform each frame into the
[00:15:24] frequency domain by splitting each frame into smaller blocks like 8 by8 pixel squares while trying to represent these squares
[00:15:31] using patterns After transformation low frequencies capture the broad shapes and lighting while high frequencies capture tiny details and noise
[00:15:41] This entire process called a discrete cosine transform or a DCT in signal processing As for an example let's take
[00:15:50] a look at this 8 by8 block of pixels in one channel that is similar in color After running a
[00:15:55] DCT those pixels aren't stored directly anymore Instead they turn into frequency coefficients that might look something like this That
[00:16:05] big number in the top left is the overall brightness of the block And the small numbers nearby describe gradients
[00:16:11] Everything else is tiny detail or noise Video compression aggressively rounds away small details and noise and leaves the big
[00:16:18] number because those are what our eyes notice the most Instead of hiding our data in fragile pixel values we
[00:16:24] hide it directly inside those big coefficients in the top left or the parts of the block that control the
[00:16:29] overall brightness and shapes and the parts of the codec that's least likely to round down We also should hide
[00:16:36] our bits into the signed bits of each value since the magnitude itself will always change but the sign has
[00:16:41] a greater chance of surviving
[00:16:56] to do this For each bit we will gently push one slightly more positive for negative
[00:17:18] When we push our coefficient a little farther away from zero it has more breathing room Even after compression squashes
[00:17:25] it it usually just stays positive or negative In other words the magnitude changes but the sign survives and that's
[00:17:32] enough for us to reliably recover the bit
[00:17:51] There's one more catch though which is what video format we export our frames in If we save the video
[00:17:57] using a lossy video codec like H.264 or AV1 we'd immediately lose precision because the codec itself already uses some
[00:18:06] level of video compression This would mean we are effectively corrupting our data twice once through export and again through
[00:18:14] when we upload our video onto YouTube which goes through their compression algorithm
[00:18:25] That's why we need to export the video using a lossless codec instead preserving each pixel as close to the
[00:18:30] original as possible
[00:18:44] That way the only lossy step that ever touches our data is YouTube's recompression
[00:18:53] Now that's all There's only one thing left to do Let's try it Here I have my ID here and
[00:19:02] you can see my code going to compile it and I'm going to go into my folder and there's the
[00:19:09] uh tool I made for re-encoding and decoding the file Right here I have an input.ext as an example It's
[00:19:17] the BM movie script And I'm going to run this command to to first output it into a video file
[00:19:25] that we can upload onto YouTube And I'm going to speed through this a bit because um it takes a
[00:19:32] bit And here now I'm going to upload my video onto YouTube make it unlisted and do all the stuff
[00:19:39] the logistics I'm going to copy the link here And then once we copy that link we're going to use
[00:19:46] a tool called YTDLP which is an awesome tool that allows us to download the best video quality from our
[00:19:55] YouTube link So we're going to run that command format best video and put in our URL there It's going
[00:20:01] to give us a new video And notice how this video file below is much smaller 1 megabyte versus 20
[00:20:10] megabytes I showed the showed it what it looks like on screen Then let's go back and use our same
[00:20:16] tool to decode the file right from whatever comes out of YouTube And we're going to output it with a
[00:20:24] different name because we don't want it to override our test file control I'm going to speed this up a
[00:20:30] bit so that you guys are able to not don't have to wait for this And you can see that
[00:20:36] once we open our decoded file here and compare it with our context of our original file the control file
[00:20:43] it's exactly the same We got back the B movie script We got back everything exactly the same So yeah
[00:20:50] that's how the tool works in action There's honestly so much I learned throughout this entire research and coding process
[00:20:58] And at first I didn't even think that this would be possible Well there you have it File storage on
[00:21:04] YouTube Is it ideal Probably not I mean the extra file size you need from layers of error proofing is
[00:21:11] likely not worth the hassle but it's cool seeing this working in principle The code's open source so you can
[00:21:17] view it in the GitHub link below Uh it's written in C++ and uses some libraries that you may have
[00:21:22] to install in order to compile it
Comments ()