All language subtitles for [English (auto-generated)] Every Processing Unit Explained in 8 Minutes!

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranî)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu
Would you like to inspect the original subtitles? These are the user uploaded subtitles that are being translated: 1 00:00:00,000 --> 00:00:01,840 CPU, 2 00:00:01,840 --> 00:00:03,760 the central processing unit, the one 3 00:00:03,760 --> 00:00:05,680 chip everyone's heard of, and the one 4 00:00:05,680 --> 00:00:07,640 almost nobody can actually explain 5 00:00:07,640 --> 00:00:09,680 without saying it's the brain. 6 00:00:09,680 --> 00:00:11,040 Here's the thing. 7 00:00:11,040 --> 00:00:13,800 A CPU is a generalist. It handles 8 00:00:13,800 --> 00:00:15,840 anything you throw at it. Opening a 9 00:00:15,840 --> 00:00:18,000 browser, running Excel, decoding a 10 00:00:18,000 --> 00:00:20,600 video, checking if your mouse moved. And 11 00:00:20,600 --> 00:00:22,720 it does it sequentially, one task at a 12 00:00:22,720 --> 00:00:25,920 time, but at insane speeds. Modern ones 13 00:00:25,920 --> 00:00:28,640 run at 3 to 5 GHz. That's billions of 14 00:00:28,640 --> 00:00:30,960 instructions per second. Intel kicked 15 00:00:30,960 --> 00:00:33,480 this whole party off in 1971 with the 16 00:00:33,480 --> 00:00:35,000 4004, 17 00:00:35,000 --> 00:00:38,120 a chip with 2,300 transistors. 18 00:00:38,120 --> 00:00:40,600 Your current CPU has tens of billions 19 00:00:40,600 --> 00:00:43,080 packed into a fingernail of silicon. 20 00:00:43,080 --> 00:00:45,440 Today, you're looking at 4 to 24 cores 21 00:00:45,440 --> 00:00:47,560 in a consumer chip, like Apple's M 22 00:00:47,560 --> 00:00:51,760 series, Intel's Core Ultra, AMD Ryzen. 23 00:00:51,760 --> 00:00:54,440 All these CPU has enough cores to make 24 00:00:54,440 --> 00:00:57,240 your life smooth. Each core is basically 25 00:00:57,240 --> 00:00:59,320 a mini processor that can run its own 26 00:00:59,320 --> 00:01:02,840 task. Think of a CPU as that one kid in 27 00:01:02,840 --> 00:01:05,440 school who was solid at math, history, 28 00:01:05,440 --> 00:01:07,320 English, and gym. 29 00:01:07,320 --> 00:01:09,760 Never the best at any single thing, but 30 00:01:09,760 --> 00:01:12,120 you want them on every group project. 31 00:01:12,120 --> 00:01:15,000 The catch, CPU are terrible at doing the 32 00:01:15,000 --> 00:01:17,280 same small task a million times in 33 00:01:17,280 --> 00:01:18,440 parallel. 34 00:01:18,440 --> 00:01:20,680 Ask one to render a 4K frame and it'll 35 00:01:20,680 --> 00:01:22,840 cry, which is exactly why we invented 36 00:01:22,840 --> 00:01:26,200 the next one, GPU. Graphics processing 37 00:01:26,200 --> 00:01:28,720 unit, originally built for video games, 38 00:01:28,720 --> 00:01:30,800 accidentally became the most important 39 00:01:30,800 --> 00:01:32,720 chip of the AI era. 40 00:01:32,720 --> 00:01:36,600 Where a CPU has 16 billion cores, a GPU 41 00:01:36,600 --> 00:01:38,640 has thousands of simple ones. It's the 42 00:01:38,640 --> 00:01:40,440 difference between hiring one Nobel 43 00:01:40,440 --> 00:01:42,560 laureate and hiring a warehouse full of 44 00:01:42,560 --> 00:01:43,880 interns. 45 00:01:43,880 --> 00:01:46,040 If the job is doing the same small math 46 00:01:46,040 --> 00:01:48,760 problem a million times, the interns win 47 00:01:48,760 --> 00:01:49,960 every time. 48 00:01:49,960 --> 00:01:51,960 And it turns out rendering pixels, 49 00:01:51,960 --> 00:01:53,760 training neural networks, and mining 50 00:01:53,760 --> 00:01:55,800 crypto all look like that exact same 51 00:01:55,800 --> 00:01:59,360 problem, massive parallel matrix math. 52 00:01:59,360 --> 00:02:03,640 Nvidia's RTX 5090 packs over 21,000 CUDA 53 00:02:03,640 --> 00:02:04,720 cores. 54 00:02:04,720 --> 00:02:07,320 AMD's Radeon lineup fights in the same 55 00:02:07,320 --> 00:02:09,679 weight class. These things chew through 56 00:02:09,679 --> 00:02:12,400 ray-traced games, train GPT class 57 00:02:12,400 --> 00:02:16,560 models, and melt your power bill. An RTX 58 00:02:16,560 --> 00:02:20,760 5090 pulls 575 W under load. That's a 59 00:02:20,760 --> 00:02:22,240 microwave. 60 00:02:22,240 --> 00:02:24,240 But here's the dirty secret of the AI 61 00:02:24,240 --> 00:02:27,000 boom, ChatGPT, Midjourney, Stable 62 00:02:27,000 --> 00:02:30,280 Diffusion. None of them run on CPU. They 63 00:02:30,280 --> 00:02:33,720 run on racks of GPU. Nvidia's market cap 64 00:02:33,720 --> 00:02:36,520 crossed $3 trillion because the world 65 00:02:36,520 --> 00:02:38,680 quietly needed a million of their chips 66 00:02:38,680 --> 00:02:41,320 all at once. TPU. 67 00:02:41,320 --> 00:02:44,320 Tensor processing unit. Google's answer 68 00:02:44,320 --> 00:02:46,400 to, what if we built a chip that only 69 00:02:46,400 --> 00:02:48,680 knows how to do AI math and absolutely 70 00:02:48,680 --> 00:02:52,840 nothing else. Released in 2016, the TPU 71 00:02:52,840 --> 00:02:55,160 was Google's admission that even GPU 72 00:02:55,160 --> 00:02:57,160 were overkill for what AI actually 73 00:02:57,160 --> 00:02:58,800 needed. 74 00:02:58,800 --> 00:03:00,760 Neural networks are basically endless 75 00:03:00,760 --> 00:03:03,480 matrix multiplications. So Google built 76 00:03:03,480 --> 00:03:05,680 silicon that only does that at ludicrous 77 00:03:05,680 --> 00:03:08,920 speed. No graphics, no physics, no video 78 00:03:08,920 --> 00:03:11,400 decode, just tensors. 79 00:03:11,400 --> 00:03:13,960 Think of a TPU as a NASCAR car. It can 80 00:03:13,960 --> 00:03:16,280 only turn left, but on an oval track, it 81 00:03:16,280 --> 00:03:18,640 humiliates any Ferrari. 82 00:03:18,640 --> 00:03:22,200 Google's latest TPU v5p powers Gemini, 83 00:03:22,200 --> 00:03:24,320 the AI in your search results, and 84 00:03:24,320 --> 00:03:26,400 almost every neural network inside 85 00:03:26,400 --> 00:03:28,840 Google's empire. You can't buy one. You 86 00:03:28,840 --> 00:03:31,000 rent time on it through Google Cloud. 87 00:03:31,000 --> 00:03:34,000 The tradeoff is flexibility. A TPU is 88 00:03:34,000 --> 00:03:35,840 useless for anything outside machine 89 00:03:35,840 --> 00:03:37,800 learning. But for the thing it was built 90 00:03:37,800 --> 00:03:40,360 for, it's faster per dollar and per watt 91 00:03:40,360 --> 00:03:43,120 than almost anything on Earth. NPU. 92 00:03:43,120 --> 00:03:45,360 Neural processing unit, the quiet chip 93 00:03:45,360 --> 00:03:46,880 sitting inside your phone and your new 94 00:03:46,880 --> 00:03:49,280 laptop doing AI without nuking your 95 00:03:49,280 --> 00:03:51,120 battery or phoning home to a data 96 00:03:51,120 --> 00:03:52,080 center. 97 00:03:52,080 --> 00:03:54,080 Apple slid the first serious one into 98 00:03:54,080 --> 00:03:56,920 the iPhone 10 in 2017, the neural 99 00:03:56,920 --> 00:03:57,880 engine. 100 00:03:57,880 --> 00:04:00,240 Qualcomm followed with Hexagon. Now 101 00:04:00,240 --> 00:04:02,560 every Copilot Plus PC is required to 102 00:04:02,560 --> 00:04:06,240 have an NPU capable of 40 plus tops, or 103 00:04:06,240 --> 00:04:09,280 trillions of operations per second, just 104 00:04:09,280 --> 00:04:11,800 to earn the label. What does it actually 105 00:04:11,800 --> 00:04:14,240 do? It's the reason your phone blurs 106 00:04:14,240 --> 00:04:16,400 backgrounds in real time, transcribes 107 00:04:16,400 --> 00:04:18,480 your voice offline, and unlocks with 108 00:04:18,480 --> 00:04:21,239 your face in 200 milliseconds. 109 00:04:21,239 --> 00:04:23,120 It's also why your new laptop can run a 110 00:04:23,120 --> 00:04:25,000 local language model without the fan 111 00:04:25,000 --> 00:04:26,760 sounding like a truck engine. 112 00:04:26,760 --> 00:04:30,560 An NPU is basically a TPU's pocket-sized 113 00:04:30,560 --> 00:04:33,720 cousin, smaller, dumber, but sips power 114 00:04:33,720 --> 00:04:35,720 like it's rationed. When your iPhone 115 00:04:35,720 --> 00:04:38,560 edits a photo instantly, that's the NPU, 116 00:04:38,560 --> 00:04:40,720 not the CPU, doing the quiet heavy 117 00:04:40,720 --> 00:04:44,520 lifting. LPU. Language processing unit, 118 00:04:44,520 --> 00:04:47,080 a chip so new and so specific that most 119 00:04:47,080 --> 00:04:48,640 people in tech still haven't heard of 120 00:04:48,640 --> 00:04:51,680 it. Grok. Not Elon's Grok, this is the 121 00:04:51,680 --> 00:04:53,440 chip company with a Q. 122 00:04:53,440 --> 00:04:56,360 Built the LPU specifically to run large 123 00:04:56,360 --> 00:04:58,680 language models at speeds nobody thought 124 00:04:58,680 --> 00:05:01,080 were possible. We're talking 500 plus 125 00:05:01,080 --> 00:05:03,520 tokens per second on Llama 3, where a 126 00:05:03,520 --> 00:05:06,160 standard GPU setup might do 50. It 127 00:05:06,160 --> 00:05:08,720 doesn't train models, it just runs them, 128 00:05:08,720 --> 00:05:10,400 faster than anything else alive right 129 00:05:10,400 --> 00:05:11,600 now. 130 00:05:11,600 --> 00:05:13,320 The whole architecture is built around 131 00:05:13,320 --> 00:05:14,800 the fact that language models are 132 00:05:14,800 --> 00:05:17,200 sequential, word by word, token by 133 00:05:17,200 --> 00:05:20,040 token, and GPU, for all their brute 134 00:05:20,040 --> 00:05:22,520 force, weren't built for that rhythm. 135 00:05:22,520 --> 00:05:24,360 If you've ever used Grok's demo and 136 00:05:24,360 --> 00:05:26,000 watched the response fill the screen 137 00:05:26,000 --> 00:05:28,440 instantly, like it's cheating, 138 00:05:28,440 --> 00:05:32,120 yeah, that's the LPU. IPU. Intelligence 139 00:05:32,120 --> 00:05:34,280 processing unit, built by Graphcore, a 140 00:05:34,280 --> 00:05:35,920 British chip company that tried to 141 00:05:35,920 --> 00:05:38,800 outthink Nvidia, and mostly ended up as 142 00:05:38,800 --> 00:05:40,320 a cautionary tale. 143 00:05:40,320 --> 00:05:42,640 The idea was genuinely brilliant. 144 00:05:42,640 --> 00:05:45,040 Instead of GPU that do parallel math in 145 00:05:45,040 --> 00:05:47,320 straight lines, built a chip that thinks 146 00:05:47,320 --> 00:05:48,840 in graphs. 147 00:05:48,840 --> 00:05:51,680 A spiderweb of 1,472 148 00:05:51,680 --> 00:05:54,240 independent processors all talking to 149 00:05:54,240 --> 00:05:56,760 each other at once. Great for certain AI 150 00:05:56,760 --> 00:05:59,480 workloads. Problem is, the world 151 00:05:59,480 --> 00:06:01,840 standardized on Nvidia's CUDA software 152 00:06:01,840 --> 00:06:04,200 before Graphcore could get traction. 153 00:06:04,200 --> 00:06:06,520 The IPU is technically impressive and 154 00:06:06,520 --> 00:06:08,440 commercially struggling. 155 00:06:08,440 --> 00:06:10,360 A reminder that the best chip doesn't 156 00:06:10,360 --> 00:06:12,160 always win. The chip with the best 157 00:06:12,160 --> 00:06:14,200 software ecosystem does. 158 00:06:14,200 --> 00:06:15,640 DPU. 159 00:06:15,640 --> 00:06:18,000 Data processing unit, the bouncer of the 160 00:06:18,000 --> 00:06:19,560 modern data center. 161 00:06:19,560 --> 00:06:23,280 Nvidia's BlueField, AWS Nitro, Intel's 162 00:06:23,280 --> 00:06:25,000 IPU. 163 00:06:25,000 --> 00:06:27,200 They all do the same boring essential 164 00:06:27,200 --> 00:06:29,480 job, handle the infrastructure grunt 165 00:06:29,480 --> 00:06:32,720 work so the CPU and GPU don't have to. 166 00:06:32,720 --> 00:06:35,400 Networking, encryption, storage traffic, 167 00:06:35,400 --> 00:06:37,720 security checks. In a modern cloud 168 00:06:37,720 --> 00:06:40,160 server, the DPU is quietly shuffling 169 00:06:40,160 --> 00:06:42,480 terabytes per second between machines 170 00:06:42,480 --> 00:06:44,920 while the GPU focuses on the actual AI 171 00:06:44,920 --> 00:06:47,200 model. You'll never interact with one 172 00:06:47,200 --> 00:06:50,000 directly, but if you've ever used AWS, 173 00:06:50,000 --> 00:06:53,480 Azure, or Google Cloud, a DPU handled 174 00:06:53,480 --> 00:06:57,120 your packet. QPU. Quantum processing 175 00:06:57,120 --> 00:06:59,200 unit, and this one's going to break your 176 00:06:59,200 --> 00:07:00,480 brain a little. 177 00:07:00,480 --> 00:07:04,240 A regular chip uses bits, zero or one. A 178 00:07:04,240 --> 00:07:07,240 QPU uses qubits, which can be zero or 179 00:07:07,240 --> 00:07:09,440 both at the same time, thanks to a 180 00:07:09,440 --> 00:07:12,000 property called superposition. Instead 181 00:07:12,000 --> 00:07:14,360 of checking every possibility one by one 182 00:07:14,360 --> 00:07:17,840 like a CPU, a QPU explores all of them 183 00:07:17,840 --> 00:07:20,280 simultaneously and collapses into the 184 00:07:20,280 --> 00:07:22,560 right answer at the end. It's less a 185 00:07:22,560 --> 00:07:24,520 faster calculator, more like asking the 186 00:07:24,520 --> 00:07:27,040 universe to do your homework. IBM's 187 00:07:27,040 --> 00:07:29,960 Condor chip crossed 1,121 188 00:07:29,960 --> 00:07:32,400 qubits. Google's Willow chip, released 189 00:07:32,400 --> 00:07:35,240 late 2024, solved a problem in under 5 190 00:07:35,240 --> 00:07:36,680 minutes that would take the fastest 191 00:07:36,680 --> 00:07:38,920 supercomputer on Earth 10 septillion 192 00:07:38,920 --> 00:07:40,960 years. That's a number older than the 193 00:07:40,960 --> 00:07:43,200 universe many times over. 194 00:07:43,200 --> 00:07:45,640 The catch, QPU needs to be cooled to 195 00:07:45,640 --> 00:07:49,760 near absolute zero. -273°C, 196 00:07:49,760 --> 00:07:51,480 colder than outer space. They're the 197 00:07:51,480 --> 00:07:53,440 size of a chandelier. They're wildly 198 00:07:53,440 --> 00:07:55,800 unstable, and they're genuinely useless 199 00:07:55,800 --> 00:07:57,160 for running Chrome. But for 200 00:07:57,160 --> 00:07:59,400 cryptography, drug discovery, and 201 00:07:59,400 --> 00:08:01,600 simulating molecules, they're about to 202 00:08:01,600 --> 00:08:03,040 rewrite what's possible in the next 203 00:08:03,040 --> 00:08:05,960 decade. So, here's the takeaway. 204 00:08:05,960 --> 00:08:08,800 For your everyday computer, the CPU. For 205 00:08:08,800 --> 00:08:11,840 gaming and training AI, the GPU. For 206 00:08:11,840 --> 00:08:15,120 cloud AI at scale, the TPU. For AI on 207 00:08:15,120 --> 00:08:18,200 your phone and laptop, the NPU. For 208 00:08:18,200 --> 00:08:20,640 blazing fast language model responses, 209 00:08:20,640 --> 00:08:23,200 the LPU. For data center plumbing, the 210 00:08:23,200 --> 00:08:25,680 DPU. And for problems we currently can't 211 00:08:25,680 --> 00:08:28,880 solve at all, the QPU is the answer. 212 00:08:28,880 --> 00:08:30,920 Hey, I made a video on some of the most 213 00:08:30,920 --> 00:08:33,400 important data structures. Definitely 214 00:08:33,400 --> 00:08:35,440 check it out, it might really help you. 215 00:08:35,440 --> 00:08:36,919 And with that, I'll see you in the next 216 00:08:36,919 --> 00:08:39,240 one.14936

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.