All language subtitles for [English (auto-generated)] Every Processing Unit Explained in 8 Minutes!
Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
Would you like to inspect the original subtitles? These are the user uploaded subtitles that are being translated:
1
00:00:00,000 --> 00:00:01,840
CPU,
2
00:00:01,840 --> 00:00:03,760
the central processing unit, the one
3
00:00:03,760 --> 00:00:05,680
chip everyone's heard of, and the one
4
00:00:05,680 --> 00:00:07,640
almost nobody can actually explain
5
00:00:07,640 --> 00:00:09,680
without saying it's the brain.
6
00:00:09,680 --> 00:00:11,040
Here's the thing.
7
00:00:11,040 --> 00:00:13,800
A CPU is a generalist. It handles
8
00:00:13,800 --> 00:00:15,840
anything you throw at it. Opening a
9
00:00:15,840 --> 00:00:18,000
browser, running Excel, decoding a
10
00:00:18,000 --> 00:00:20,600
video, checking if your mouse moved. And
11
00:00:20,600 --> 00:00:22,720
it does it sequentially, one task at a
12
00:00:22,720 --> 00:00:25,920
time, but at insane speeds. Modern ones
13
00:00:25,920 --> 00:00:28,640
run at 3 to 5 GHz. That's billions of
14
00:00:28,640 --> 00:00:30,960
instructions per second. Intel kicked
15
00:00:30,960 --> 00:00:33,480
this whole party off in 1971 with the
16
00:00:33,480 --> 00:00:35,000
4004,
17
00:00:35,000 --> 00:00:38,120
a chip with 2,300 transistors.
18
00:00:38,120 --> 00:00:40,600
Your current CPU has tens of billions
19
00:00:40,600 --> 00:00:43,080
packed into a fingernail of silicon.
20
00:00:43,080 --> 00:00:45,440
Today, you're looking at 4 to 24 cores
21
00:00:45,440 --> 00:00:47,560
in a consumer chip, like Apple's M
22
00:00:47,560 --> 00:00:51,760
series, Intel's Core Ultra, AMD Ryzen.
23
00:00:51,760 --> 00:00:54,440
All these CPU has enough cores to make
24
00:00:54,440 --> 00:00:57,240
your life smooth. Each core is basically
25
00:00:57,240 --> 00:00:59,320
a mini processor that can run its own
26
00:00:59,320 --> 00:01:02,840
task. Think of a CPU as that one kid in
27
00:01:02,840 --> 00:01:05,440
school who was solid at math, history,
28
00:01:05,440 --> 00:01:07,320
English, and gym.
29
00:01:07,320 --> 00:01:09,760
Never the best at any single thing, but
30
00:01:09,760 --> 00:01:12,120
you want them on every group project.
31
00:01:12,120 --> 00:01:15,000
The catch, CPU are terrible at doing the
32
00:01:15,000 --> 00:01:17,280
same small task a million times in
33
00:01:17,280 --> 00:01:18,440
parallel.
34
00:01:18,440 --> 00:01:20,680
Ask one to render a 4K frame and it'll
35
00:01:20,680 --> 00:01:22,840
cry, which is exactly why we invented
36
00:01:22,840 --> 00:01:26,200
the next one, GPU. Graphics processing
37
00:01:26,200 --> 00:01:28,720
unit, originally built for video games,
38
00:01:28,720 --> 00:01:30,800
accidentally became the most important
39
00:01:30,800 --> 00:01:32,720
chip of the AI era.
40
00:01:32,720 --> 00:01:36,600
Where a CPU has 16 billion cores, a GPU
41
00:01:36,600 --> 00:01:38,640
has thousands of simple ones. It's the
42
00:01:38,640 --> 00:01:40,440
difference between hiring one Nobel
43
00:01:40,440 --> 00:01:42,560
laureate and hiring a warehouse full of
44
00:01:42,560 --> 00:01:43,880
interns.
45
00:01:43,880 --> 00:01:46,040
If the job is doing the same small math
46
00:01:46,040 --> 00:01:48,760
problem a million times, the interns win
47
00:01:48,760 --> 00:01:49,960
every time.
48
00:01:49,960 --> 00:01:51,960
And it turns out rendering pixels,
49
00:01:51,960 --> 00:01:53,760
training neural networks, and mining
50
00:01:53,760 --> 00:01:55,800
crypto all look like that exact same
51
00:01:55,800 --> 00:01:59,360
problem, massive parallel matrix math.
52
00:01:59,360 --> 00:02:03,640
Nvidia's RTX 5090 packs over 21,000 CUDA
53
00:02:03,640 --> 00:02:04,720
cores.
54
00:02:04,720 --> 00:02:07,320
AMD's Radeon lineup fights in the same
55
00:02:07,320 --> 00:02:09,679
weight class. These things chew through
56
00:02:09,679 --> 00:02:12,400
ray-traced games, train GPT class
57
00:02:12,400 --> 00:02:16,560
models, and melt your power bill. An RTX
58
00:02:16,560 --> 00:02:20,760
5090 pulls 575 W under load. That's a
59
00:02:20,760 --> 00:02:22,240
microwave.
60
00:02:22,240 --> 00:02:24,240
But here's the dirty secret of the AI
61
00:02:24,240 --> 00:02:27,000
boom, ChatGPT, Midjourney, Stable
62
00:02:27,000 --> 00:02:30,280
Diffusion. None of them run on CPU. They
63
00:02:30,280 --> 00:02:33,720
run on racks of GPU. Nvidia's market cap
64
00:02:33,720 --> 00:02:36,520
crossed $3 trillion because the world
65
00:02:36,520 --> 00:02:38,680
quietly needed a million of their chips
66
00:02:38,680 --> 00:02:41,320
all at once. TPU.
67
00:02:41,320 --> 00:02:44,320
Tensor processing unit. Google's answer
68
00:02:44,320 --> 00:02:46,400
to, what if we built a chip that only
69
00:02:46,400 --> 00:02:48,680
knows how to do AI math and absolutely
70
00:02:48,680 --> 00:02:52,840
nothing else. Released in 2016, the TPU
71
00:02:52,840 --> 00:02:55,160
was Google's admission that even GPU
72
00:02:55,160 --> 00:02:57,160
were overkill for what AI actually
73
00:02:57,160 --> 00:02:58,800
needed.
74
00:02:58,800 --> 00:03:00,760
Neural networks are basically endless
75
00:03:00,760 --> 00:03:03,480
matrix multiplications. So Google built
76
00:03:03,480 --> 00:03:05,680
silicon that only does that at ludicrous
77
00:03:05,680 --> 00:03:08,920
speed. No graphics, no physics, no video
78
00:03:08,920 --> 00:03:11,400
decode, just tensors.
79
00:03:11,400 --> 00:03:13,960
Think of a TPU as a NASCAR car. It can
80
00:03:13,960 --> 00:03:16,280
only turn left, but on an oval track, it
81
00:03:16,280 --> 00:03:18,640
humiliates any Ferrari.
82
00:03:18,640 --> 00:03:22,200
Google's latest TPU v5p powers Gemini,
83
00:03:22,200 --> 00:03:24,320
the AI in your search results, and
84
00:03:24,320 --> 00:03:26,400
almost every neural network inside
85
00:03:26,400 --> 00:03:28,840
Google's empire. You can't buy one. You
86
00:03:28,840 --> 00:03:31,000
rent time on it through Google Cloud.
87
00:03:31,000 --> 00:03:34,000
The tradeoff is flexibility. A TPU is
88
00:03:34,000 --> 00:03:35,840
useless for anything outside machine
89
00:03:35,840 --> 00:03:37,800
learning. But for the thing it was built
90
00:03:37,800 --> 00:03:40,360
for, it's faster per dollar and per watt
91
00:03:40,360 --> 00:03:43,120
than almost anything on Earth. NPU.
92
00:03:43,120 --> 00:03:45,360
Neural processing unit, the quiet chip
93
00:03:45,360 --> 00:03:46,880
sitting inside your phone and your new
94
00:03:46,880 --> 00:03:49,280
laptop doing AI without nuking your
95
00:03:49,280 --> 00:03:51,120
battery or phoning home to a data
96
00:03:51,120 --> 00:03:52,080
center.
97
00:03:52,080 --> 00:03:54,080
Apple slid the first serious one into
98
00:03:54,080 --> 00:03:56,920
the iPhone 10 in 2017, the neural
99
00:03:56,920 --> 00:03:57,880
engine.
100
00:03:57,880 --> 00:04:00,240
Qualcomm followed with Hexagon. Now
101
00:04:00,240 --> 00:04:02,560
every Copilot Plus PC is required to
102
00:04:02,560 --> 00:04:06,240
have an NPU capable of 40 plus tops, or
103
00:04:06,240 --> 00:04:09,280
trillions of operations per second, just
104
00:04:09,280 --> 00:04:11,800
to earn the label. What does it actually
105
00:04:11,800 --> 00:04:14,240
do? It's the reason your phone blurs
106
00:04:14,240 --> 00:04:16,400
backgrounds in real time, transcribes
107
00:04:16,400 --> 00:04:18,480
your voice offline, and unlocks with
108
00:04:18,480 --> 00:04:21,239
your face in 200 milliseconds.
109
00:04:21,239 --> 00:04:23,120
It's also why your new laptop can run a
110
00:04:23,120 --> 00:04:25,000
local language model without the fan
111
00:04:25,000 --> 00:04:26,760
sounding like a truck engine.
112
00:04:26,760 --> 00:04:30,560
An NPU is basically a TPU's pocket-sized
113
00:04:30,560 --> 00:04:33,720
cousin, smaller, dumber, but sips power
114
00:04:33,720 --> 00:04:35,720
like it's rationed. When your iPhone
115
00:04:35,720 --> 00:04:38,560
edits a photo instantly, that's the NPU,
116
00:04:38,560 --> 00:04:40,720
not the CPU, doing the quiet heavy
117
00:04:40,720 --> 00:04:44,520
lifting. LPU. Language processing unit,
118
00:04:44,520 --> 00:04:47,080
a chip so new and so specific that most
119
00:04:47,080 --> 00:04:48,640
people in tech still haven't heard of
120
00:04:48,640 --> 00:04:51,680
it. Grok. Not Elon's Grok, this is the
121
00:04:51,680 --> 00:04:53,440
chip company with a Q.
122
00:04:53,440 --> 00:04:56,360
Built the LPU specifically to run large
123
00:04:56,360 --> 00:04:58,680
language models at speeds nobody thought
124
00:04:58,680 --> 00:05:01,080
were possible. We're talking 500 plus
125
00:05:01,080 --> 00:05:03,520
tokens per second on Llama 3, where a
126
00:05:03,520 --> 00:05:06,160
standard GPU setup might do 50. It
127
00:05:06,160 --> 00:05:08,720
doesn't train models, it just runs them,
128
00:05:08,720 --> 00:05:10,400
faster than anything else alive right
129
00:05:10,400 --> 00:05:11,600
now.
130
00:05:11,600 --> 00:05:13,320
The whole architecture is built around
131
00:05:13,320 --> 00:05:14,800
the fact that language models are
132
00:05:14,800 --> 00:05:17,200
sequential, word by word, token by
133
00:05:17,200 --> 00:05:20,040
token, and GPU, for all their brute
134
00:05:20,040 --> 00:05:22,520
force, weren't built for that rhythm.
135
00:05:22,520 --> 00:05:24,360
If you've ever used Grok's demo and
136
00:05:24,360 --> 00:05:26,000
watched the response fill the screen
137
00:05:26,000 --> 00:05:28,440
instantly, like it's cheating,
138
00:05:28,440 --> 00:05:32,120
yeah, that's the LPU. IPU. Intelligence
139
00:05:32,120 --> 00:05:34,280
processing unit, built by Graphcore, a
140
00:05:34,280 --> 00:05:35,920
British chip company that tried to
141
00:05:35,920 --> 00:05:38,800
outthink Nvidia, and mostly ended up as
142
00:05:38,800 --> 00:05:40,320
a cautionary tale.
143
00:05:40,320 --> 00:05:42,640
The idea was genuinely brilliant.
144
00:05:42,640 --> 00:05:45,040
Instead of GPU that do parallel math in
145
00:05:45,040 --> 00:05:47,320
straight lines, built a chip that thinks
146
00:05:47,320 --> 00:05:48,840
in graphs.
147
00:05:48,840 --> 00:05:51,680
A spiderweb of 1,472
148
00:05:51,680 --> 00:05:54,240
independent processors all talking to
149
00:05:54,240 --> 00:05:56,760
each other at once. Great for certain AI
150
00:05:56,760 --> 00:05:59,480
workloads. Problem is, the world
151
00:05:59,480 --> 00:06:01,840
standardized on Nvidia's CUDA software
152
00:06:01,840 --> 00:06:04,200
before Graphcore could get traction.
153
00:06:04,200 --> 00:06:06,520
The IPU is technically impressive and
154
00:06:06,520 --> 00:06:08,440
commercially struggling.
155
00:06:08,440 --> 00:06:10,360
A reminder that the best chip doesn't
156
00:06:10,360 --> 00:06:12,160
always win. The chip with the best
157
00:06:12,160 --> 00:06:14,200
software ecosystem does.
158
00:06:14,200 --> 00:06:15,640
DPU.
159
00:06:15,640 --> 00:06:18,000
Data processing unit, the bouncer of the
160
00:06:18,000 --> 00:06:19,560
modern data center.
161
00:06:19,560 --> 00:06:23,280
Nvidia's BlueField, AWS Nitro, Intel's
162
00:06:23,280 --> 00:06:25,000
IPU.
163
00:06:25,000 --> 00:06:27,200
They all do the same boring essential
164
00:06:27,200 --> 00:06:29,480
job, handle the infrastructure grunt
165
00:06:29,480 --> 00:06:32,720
work so the CPU and GPU don't have to.
166
00:06:32,720 --> 00:06:35,400
Networking, encryption, storage traffic,
167
00:06:35,400 --> 00:06:37,720
security checks. In a modern cloud
168
00:06:37,720 --> 00:06:40,160
server, the DPU is quietly shuffling
169
00:06:40,160 --> 00:06:42,480
terabytes per second between machines
170
00:06:42,480 --> 00:06:44,920
while the GPU focuses on the actual AI
171
00:06:44,920 --> 00:06:47,200
model. You'll never interact with one
172
00:06:47,200 --> 00:06:50,000
directly, but if you've ever used AWS,
173
00:06:50,000 --> 00:06:53,480
Azure, or Google Cloud, a DPU handled
174
00:06:53,480 --> 00:06:57,120
your packet. QPU. Quantum processing
175
00:06:57,120 --> 00:06:59,200
unit, and this one's going to break your
176
00:06:59,200 --> 00:07:00,480
brain a little.
177
00:07:00,480 --> 00:07:04,240
A regular chip uses bits, zero or one. A
178
00:07:04,240 --> 00:07:07,240
QPU uses qubits, which can be zero or
179
00:07:07,240 --> 00:07:09,440
both at the same time, thanks to a
180
00:07:09,440 --> 00:07:12,000
property called superposition. Instead
181
00:07:12,000 --> 00:07:14,360
of checking every possibility one by one
182
00:07:14,360 --> 00:07:17,840
like a CPU, a QPU explores all of them
183
00:07:17,840 --> 00:07:20,280
simultaneously and collapses into the
184
00:07:20,280 --> 00:07:22,560
right answer at the end. It's less a
185
00:07:22,560 --> 00:07:24,520
faster calculator, more like asking the
186
00:07:24,520 --> 00:07:27,040
universe to do your homework. IBM's
187
00:07:27,040 --> 00:07:29,960
Condor chip crossed 1,121
188
00:07:29,960 --> 00:07:32,400
qubits. Google's Willow chip, released
189
00:07:32,400 --> 00:07:35,240
late 2024, solved a problem in under 5
190
00:07:35,240 --> 00:07:36,680
minutes that would take the fastest
191
00:07:36,680 --> 00:07:38,920
supercomputer on Earth 10 septillion
192
00:07:38,920 --> 00:07:40,960
years. That's a number older than the
193
00:07:40,960 --> 00:07:43,200
universe many times over.
194
00:07:43,200 --> 00:07:45,640
The catch, QPU needs to be cooled to
195
00:07:45,640 --> 00:07:49,760
near absolute zero. -273°C,
196
00:07:49,760 --> 00:07:51,480
colder than outer space. They're the
197
00:07:51,480 --> 00:07:53,440
size of a chandelier. They're wildly
198
00:07:53,440 --> 00:07:55,800
unstable, and they're genuinely useless
199
00:07:55,800 --> 00:07:57,160
for running Chrome. But for
200
00:07:57,160 --> 00:07:59,400
cryptography, drug discovery, and
201
00:07:59,400 --> 00:08:01,600
simulating molecules, they're about to
202
00:08:01,600 --> 00:08:03,040
rewrite what's possible in the next
203
00:08:03,040 --> 00:08:05,960
decade. So, here's the takeaway.
204
00:08:05,960 --> 00:08:08,800
For your everyday computer, the CPU. For
205
00:08:08,800 --> 00:08:11,840
gaming and training AI, the GPU. For
206
00:08:11,840 --> 00:08:15,120
cloud AI at scale, the TPU. For AI on
207
00:08:15,120 --> 00:08:18,200
your phone and laptop, the NPU. For
208
00:08:18,200 --> 00:08:20,640
blazing fast language model responses,
209
00:08:20,640 --> 00:08:23,200
the LPU. For data center plumbing, the
210
00:08:23,200 --> 00:08:25,680
DPU. And for problems we currently can't
211
00:08:25,680 --> 00:08:28,880
solve at all, the QPU is the answer.
212
00:08:28,880 --> 00:08:30,920
Hey, I made a video on some of the most
213
00:08:30,920 --> 00:08:33,400
important data structures. Definitely
214
00:08:33,400 --> 00:08:35,440
check it out, it might really help you.
215
00:08:35,440 --> 00:08:36,919
And with that, I'll see you in the next
216
00:08:36,919 --> 00:08:39,240
one.14936
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.