All language subtitles for 009 To Standardize or not to Standardize_en

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranî)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian Download
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

Tutor: Should we standardize?

I avoided preparing this lecture for a couple of days.

Today I was drinking some coffee

and explaining to a colleague of mine,

"I really wanna elaborate on standardization,

but I don't think the students will be interested.

Moreover, there is a dispute on the topic."

My colleague then looked at me and said,

"Then tell that to the students.

Show them both sides"

And that's how he closed the topic and got rid of me.

So to standardize or not to standardize?

That is the question.

Let's explore a simple example.

Here's a scatter plot with four apartments.

The X axis shows the size,

while the Y axis, the price.

That's a very common regression relationship,

but we are doing clustering here,

so instead of causality,

think about how we can group the four observations.

A is a 500 square-foot apartment that is worth $50,000.

B is a 500 square-foot apartment that is worth $100,000.

C is a 1,200 square-foot apartment worth $50,000,

and D has the same size,

but is twice as expensive.

If we were to create two clusters,

just by looking at the plot,

they are likely to be AB and CD, right?

Now, what if we standardize the X axis, size that is?

Without taking you through the calculations,

that's the new situation.

The X axis of these points are either minus 1 or 1.

How would we group them now?

Well, AC and BD looks reasonable, right?

Yes, it does.

Finally, let's also standardize the Y axis or price.

Now, the Y axis only minus 1 or 1, too.

What we see is a perfect square.

We have no way of deciding

if the clusters should be AB and CD, or AC and BD.

So we went from one solution,

through a totally different one

to no solution whatsoever.

Why did that happen?

The ultimate aim of standardization

is to reduce the weight of higher numbers,

and increase that of lower ones.

Now, let's see the first graph once again.

If both axes had the same scale,

so from 0 to 100,000, we would get something like this,

but even more dramatic.

A K-means algorithm would immediately cluster

A with C and B with D,

just because the scale of price was so different

compared to size,

in terms of mere numbers.

So, scale matters.

Finally, the last graph resulted in a square

because there were only two values for each axis.

Logically, every rectangle on a graph

after being standardized

turns into a square.

So no matter how I chose the axes

or how far off they were from each other,

as long as they were in the shape of a rectangle,

the standardized output would've been a square.

With that said,

by standardizing both axes,

we remove the weight introduced by the high price values.

To sum up, if we don't standardize,

the range of the values will serve as weights

for each variable.

Price had much higher values,

which would indicate to K-means

that price is more important.

This would lead to clusters based on price, AC and BD,

the Economy Cluster and the Luxury Cluster.

Note that the clustering would barely,

if at all, care about size.

So, if we don't standardize,

we are not taking advantage of the size data whatsoever.

Therefore, it is a good practice to standardize the data

before clustering, especially for beginners.

The final note I'll leave you with

is when you should not standardize.

As standardization is trying to put all variables

on equal footing,

In some cases, we don't need to do that.

If we know that one variable

is inherently more important than another,

then standardization shouldn't be used.

Our price/size relationship could be one of those.

Most people are affected by the price

much more than the size, aren't they?

If you can't afford the price

you won't care about the size, right?

"How can you know that

prior to clustering," you may ask?

Experience plays a big role,

so practice when the time comes.

We will discuss this a bit more in the next lecture.

Thanks for watching.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.