Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
Tutor: Should we standardize?
I avoided preparing this lecture for a couple of days.
Today I was drinking some coffee
and explaining to a colleague of mine,
"I really wanna elaborate on standardization,
but I don't think the students will be interested.
Moreover, there is a dispute on the topic."
My colleague then looked at me and said,
"Then tell that to the students.
Show them both sides"
And that's how he closed the topic and got rid of me.
So to standardize or not to standardize?
That is the question.
Let's explore a simple example.
Here's a scatter plot with four apartments.
The X axis shows the size,
while the Y axis, the price.
That's a very common regression relationship,
but we are doing clustering here,
so instead of causality,
think about how we can group the four observations.
A is a 500 square-foot apartment that is worth $50,000.
B is a 500 square-foot apartment that is worth $100,000.
C is a 1,200 square-foot apartment worth $50,000,
and D has the same size,
but is twice as expensive.
If we were to create two clusters,
just by looking at the plot,
they are likely to be AB and CD, right?
Now, what if we standardize the X axis, size that is?
Without taking you through the calculations,
that's the new situation.
The X axis of these points are either minus 1 or 1.
How would we group them now?
Well, AC and BD looks reasonable, right?
Yes, it does.
Finally, let's also standardize the Y axis or price.
Now, the Y axis only minus 1 or 1, too.
What we see is a perfect square.
We have no way of deciding
if the clusters should be AB and CD, or AC and BD.
So we went from one solution,
through a totally different one
to no solution whatsoever.
Why did that happen?
The ultimate aim of standardization
is to reduce the weight of higher numbers,
and increase that of lower ones.
Now, let's see the first graph once again.
If both axes had the same scale,
so from 0 to 100,000, we would get something like this,
but even more dramatic.
A K-means algorithm would immediately cluster
A with C and B with D,
just because the scale of price was so different
compared to size,
in terms of mere numbers.
So, scale matters.
Finally, the last graph resulted in a square
because there were only two values for each axis.
Logically, every rectangle on a graph
after being standardized
turns into a square.
So no matter how I chose the axes
or how far off they were from each other,
as long as they were in the shape of a rectangle,
the standardized output would've been a square.
With that said,
by standardizing both axes,
we remove the weight introduced by the high price values.
To sum up, if we don't standardize,
the range of the values will serve as weights
for each variable.
Price had much higher values,
which would indicate to K-means
that price is more important.
This would lead to clusters based on price, AC and BD,
the Economy Cluster and the Luxury Cluster.
Note that the clustering would barely,
if at all, care about size.
So, if we don't standardize,
we are not taking advantage of the size data whatsoever.
Therefore, it is a good practice to standardize the data
before clustering, especially for beginners.
The final note I'll leave you with
is when you should not standardize.
As standardization is trying to put all variables
on equal footing,
In some cases, we don't need to do that.
If we know that one variable
is inherently more important than another,
then standardization shouldn't be used.
Our price/size relationship could be one of those.
Most people are affected by the price
much more than the size, aren't they?
If you can't afford the price
you won't care about the size, right?
"How can you know that
prior to clustering," you may ask?
Experience plays a big role,
so practice when the time comes.
We will discuss this a bit more in the next lecture.
Thanks for watching.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.