Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
Instructor: It's time for a more sophisticated,
yet easy to understand example.
We will look into market segmentation.
I'll open a clean Jupyter notebook
and import all the relevant packages
including the K-means module.
Next, we should load the dataset
located in 3.12 example.csv.
In the csv, we've got data from a retail shop.
There are 30 observations.
Each observation is a client
and we have a score for their customer satisfaction
and brand loyalty.
Let's see how the data was gathered.
Satisfaction is self-reported.
People were basically asked
to rate their shopping experience from one to 10
where 10 means extremely satisfied.
Therefore, satisfaction here is a discrete variable
and takes integer values.
Brand loyalty, on the other hand, is a tricky metric.
There is no widely accepted technique to measure it,
but there are different proxies like churn rate,
retention rate, or customer lifetime value.
In this dataset, loyalty was measured
through the number of purchases from that shop for a year
and several other factors found to be significant.
The range is from around minus two to around two
as the variable is already standardized.
That's something that often occurs
especially when creating latent variables like this one.
Okay, let's plot the data.
We can kind of identify two clusters
by looking at the graph.
This one and that one, right?
Well, before we perform any analysis,
let's reason about the problem for a while.
We can divide this graph into four squares:
low satisfaction, low loyalty,
low satisfaction, high loyalty,
high satisfaction, low loyalty,
and high satisfaction, high loyalty.
So going back to our two-cluster solution,
we realize it didn't make much sense, right?
What would these two clusters represent?
The first one looks like low satisfaction, low loyalty,
but the other one is all over the place.
It seems to me that a two-cluster solution won't cut it,
but enough speculation.
Let's leverage the new knowledge we possess
and learn something new.
First, I'll copy the data into a new variable called x.
Next, I'll use the same code we have used so far,
kmeans equals KMeans of 2,
kmeans.fit(x).
Clusters is a duplicate of x,
and the column cluster_pred of this data frame
will contain the cluster where a particular observation
was predicted to be placed by the algorithm.
Next, I'll quickly plot the data.
What we see are two clusters,
but they aren't the two we imagined, are they?
If you examine the plot closely,
you will realize that there is a cutoff line
at the satisfaction value of six.
Everything on the right is one cluster.
Everything on the left, the other.
This solution may make sense to some,
but most probably the algorithm
only considered satisfaction as a feature.
Why? Because we did not standardize the variable.
The satisfaction values
are much higher than those of loyalty
and K-means more or less disregarded loyalty as a feature.
Whenever we cluster on the basis of a single feature,
the result will look like this,
as if it was cut off by a vertical line.
That's one of the ways
to spot if something fishy is going on.
Okay, satisfaction and loyalty
seem equally important features for market segmentation.
So how can we fix this problem?
How can we give them equal weight?
Yes, by standardizing satisfaction.
There are several ways to do that in sklearn.
The simplest one I am aware of
is through the preprocessing module.
So from sklearn, import preprocessing,
then we will declare a new variable called x_scaled
equal to preprocessing.scale of x.
Scale is a method which scales each variable separately.
In other words, each column will be standardized
with respect to itself, exactly what we need, and it's done.
As you can see, x_scaled is an array
which contains the standardized satisfaction
and the same values for loyalty.
That's because loyalty was already standardized.
It had a mean of zero
and a standard deviation of one that is.
Okay, let's continue in the known way.
Since we don't know the number of clusters needed,
the elbow method will come in handy.
Let's declare a list wcss
for i in range from one to 10:
kmeans equals KMeans of i.
Again, K and M are capital when referring to the method,
kmeans.fit(x)
and we will append the result to the wcss list
using the inertia method.
Great, note that the range is from one to 10
so we will get the wcss for a 1, 2, 3
until nine cluster solutions.
That was an arbitrary decision on my side.
Finally, let's plot wcss versus the number of clusters
as we did in the previous lectures.
The result is an elbow.
Given this graph,
think about the correct number of clusters we should use.
We will explore the different solutions in the next lesson.
Thanks for watching.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.