All language subtitles for 011 Market Segmentation with Cluster Analysis (Part 1)_en

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranî)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian Download
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

Instructor: It's time for a more sophisticated,

yet easy to understand example.

We will look into market segmentation.

I'll open a clean Jupyter notebook

and import all the relevant packages

including the K-means module.

Next, we should load the dataset

located in 3.12 example.csv.

In the csv, we've got data from a retail shop.

There are 30 observations.

Each observation is a client

and we have a score for their customer satisfaction

and brand loyalty.

Let's see how the data was gathered.

Satisfaction is self-reported.

People were basically asked

to rate their shopping experience from one to 10

where 10 means extremely satisfied.

Therefore, satisfaction here is a discrete variable

and takes integer values.

Brand loyalty, on the other hand, is a tricky metric.

There is no widely accepted technique to measure it,

but there are different proxies like churn rate,

retention rate, or customer lifetime value.

In this dataset, loyalty was measured

through the number of purchases from that shop for a year

and several other factors found to be significant.

The range is from around minus two to around two

as the variable is already standardized.

That's something that often occurs

especially when creating latent variables like this one.

Okay, let's plot the data.

We can kind of identify two clusters

by looking at the graph.

This one and that one, right?

Well, before we perform any analysis,

let's reason about the problem for a while.

We can divide this graph into four squares:

low satisfaction, low loyalty,

low satisfaction, high loyalty,

high satisfaction, low loyalty,

and high satisfaction, high loyalty.

So going back to our two-cluster solution,

we realize it didn't make much sense, right?

What would these two clusters represent?

The first one looks like low satisfaction, low loyalty,

but the other one is all over the place.

It seems to me that a two-cluster solution won't cut it,

but enough speculation.

Let's leverage the new knowledge we possess

and learn something new.

First, I'll copy the data into a new variable called x.

Next, I'll use the same code we have used so far,

kmeans equals KMeans of 2,

kmeans.fit(x).

Clusters is a duplicate of x,

and the column cluster_pred of this data frame

will contain the cluster where a particular observation

was predicted to be placed by the algorithm.

Next, I'll quickly plot the data.

What we see are two clusters,

but they aren't the two we imagined, are they?

If you examine the plot closely,

you will realize that there is a cutoff line

at the satisfaction value of six.

Everything on the right is one cluster.

Everything on the left, the other.

This solution may make sense to some,

but most probably the algorithm

only considered satisfaction as a feature.

Why? Because we did not standardize the variable.

The satisfaction values

are much higher than those of loyalty

and K-means more or less disregarded loyalty as a feature.

Whenever we cluster on the basis of a single feature,

the result will look like this,

as if it was cut off by a vertical line.

That's one of the ways

to spot if something fishy is going on.

Okay, satisfaction and loyalty

seem equally important features for market segmentation.

So how can we fix this problem?

How can we give them equal weight?

Yes, by standardizing satisfaction.

There are several ways to do that in sklearn.

The simplest one I am aware of

is through the preprocessing module.

So from sklearn, import preprocessing,

then we will declare a new variable called x_scaled

equal to preprocessing.scale of x.

Scale is a method which scales each variable separately.

In other words, each column will be standardized

with respect to itself, exactly what we need, and it's done.

As you can see, x_scaled is an array

which contains the standardized satisfaction

and the same values for loyalty.

That's because loyalty was already standardized.

It had a mean of zero

and a standard deviation of one that is.

Okay, let's continue in the known way.

Since we don't know the number of clusters needed,

the elbow method will come in handy.

Let's declare a list wcss

for i in range from one to 10:

kmeans equals KMeans of i.

Again, K and M are capital when referring to the method,

kmeans.fit(x)

and we will append the result to the wcss list

using the inertia method.

Great, note that the range is from one to 10

so we will get the wcss for a 1, 2, 3

until nine cluster solutions.

That was an arbitrary decision on my side.

Finally, let's plot wcss versus the number of clusters

as we did in the previous lectures.

The result is an elbow.

Given this graph,

think about the correct number of clusters we should use.

We will explore the different solutions in the next lesson.

Thanks for watching.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.