All language subtitles for 5. EWMA Theory

af Afrikaans
ak Akan
sq Albanian
am Amharic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranรฎ)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

In this lecture, we are going to discuss another kind of moving average, the exponentially weighted

moving average.

Note that some other names for this are exponential smoothing and the low pass filter.

So if you've taken some of my other courses before and you've heard me use those terms, recognize that

this is the same thing, in fact, that this kind of moving average is very applicable in many areas

of machine learning statistics, finance and signal processing.

So you will generally see it pretty often.

So what is the exponentially weighted moving average?

I want to break this lecture up into two parts, the first part is the short summary.

If you want to only watch this part and then skip to the code, that's fine.

The second part of this lecture will be an optional in-depth discussion about why the exponentially

weighted moving average has its name.

You can opt to watch this if you want to get a better understanding of why and how this works.

OK, so what's the short summary, as you know, the arithmetic mean can be calculated by taking all

of your samples, summing them together and then dividing by the number of samples, the exponentially

weighted moving average is calculated differently.

In fact, it's calculated kind of on the fly or in an online manner.

It says that the moving average at time T is equal to some constant alpha times.

The sample at time T plus one minus alpha times the previous moving average at time at T minus one.

In other words, at each step, the new moving average is the weighted some of the new sample and the

old moving average.

All right.

So that's pretty much it.

It's not a terribly complicated calculation.

Of course, without further analysis, it's not clear why this is an average and it's not clear why

it's exponentially weighted.

The next part of our short summary is this how do we do it in code similar to the simple moving average

we call a function on our series or our data frame called IWM.

This returns and IWM object, which is similar to the rolling objects we saw previously.

It has a similar set of functions such as mean variance, covariance and so forth.

To discuss a practical issue, what value of Alpha should we choose Alpha is something like a decay

factor.

Typically, Alpha is chosen to be a small value between a zero and one like zero point one or zero point

to it might help to look at some extreme cases.

So let's say we choose Alpha equals one.

That means set the average to be just the latest value of X..

In this case, all we're doing is copying X and therefore it's not really an average at all.

On the other hand, let's say we set Alpha equal to zero, then all we're doing is copying the previous

average and we're not taking into account any new samples intuitively.

Then if we set off a very close to one that says new samples matter much more in the old average matters,

much less, you can imagine this will lead to a much more noisy time series which will more closely

match the original if we set Alpha very close to zero that says new samples matter much less and the

old average carries much more weight.

In this situation, you'll get a much smoother time series and it will take a much more drastic change

in X to affect the moving average.

OK, so now that the short summary is complete, if you want to know the details behind the exponentially

weighted moving average, keep listening.

Let's suppose we want to calculate the usual arithmetic sample mean using the formula for the sample

mean.

You might suggest that this is quite obvious.

Just take all the values of X that you've collected, add them all together and divide by the total

number of X is that you have.

The question is what's wrong with this?

I'll give you a minute to think about it.

So please pause the video until you think you have the answer.

All right, so hopefully you thought about why calculating the sample mean naively might not be such

a good idea.

What if we have a lot or even an infinite amount of data?

Obviously, our computers or our servers don't have an infinite amount of space.

And even if they did, calculating a summation is of T.

So the more data you have, the longer it will take and that will increase linearly with how much data

you've collected.

Here's my claim.

I claim that you can make the calculation of the sample mean of one on each step in both space and time

complexity, no matter how much data you collect again as an exercise before moving on to the next slide.

I want you to think about how this might be the case.

Please pause the video if you want to take a moment and think.

OK, so hopefully you thought about how you might calculate a sample mean using constant space and time.

The key is that you can calculate a sample mean using the previous sample mean let's call the sample

mean after collecting samples X bar subscript T, this means that the sample mean after collecting T

minus one samples is X, bar subscript T minus one.

We can write down the definition of both of these, which I hope is pretty obvious.

Now that you know the metric, let's again make this an exercise.

Can you express Esbati in terms of X, bar T minus one, please pause the video until you've tried this

on your own.

OK, so here's what you can do.

First, you take Esbati and split up the summation so that you only sum up to T minus one.

Then you leave X subscript T by itself.

This is just the last sample you've collected.

The next step is to realize that the sum of the ex towers from one up to T minus one can be expressed

in terms of X bar subscripts, T minus one.

We just have to rearrange the equation from earlier.

It's clear that this sum is just T minus one times X, bar T minus one.

We can substitute this into our expression for Esbati to get the sample mean at time t in terms of the

sample mean a time T minus one.

One interesting thing you can do, although it's not totally clear why you'd want to do this at this

time, is split up the formula as follows.

The first step is to multiply out the one over Tetum that gives us T minus one over T as the first coefficient

and one over T as the second coefficient.

The second step is to simplify T minus one over T to one, minus one over T.

At this point we can just leave this as is this is the form that we want.

We have one term with the previous sample mean and we have one term with the latest sample.

What's important to recognize about this equation is that we have discovered a way to calculate the

sample mean that does not depend on carrying around all of the samples you've ever collected.

All you need to have is the previous sample mean the latest sample and the number of samples you've

seen in total.

The next question to consider is, what if we believe that recent data matters more than past data?

If we look at our equation carefully, we see an interesting characteristic.

Remember that as we collect more and more samples, the value of tea is increasing.

That means as we collect more and more samples, the weight that we give to the latest sample decreases.

We can see that the weight that we give to the sample is exactly one over tea.

Now, although this might make you think that the influence of each sample somehow decays over time,

remember that this is not true because this is still just a regular arithmetic mean.

But what if we want recent data to matter more, what would happen if instead of making the way one

over tea, we simply make it a constant alpha?

Well, then this is exactly the exponentially weighted moving average.

The basic idea is instead of giving less and less weight to each new sample, we now give a constant

weight to each new sample.

Let's see how this affects the influence of each sample overall.

The next question we want to answer is, how does this update actually implement an exponentially weighted

moving average?

Can we show that this is true?

And in fact, it's not too difficult at this point.

What we can do is just keep recursively plugging in.

Older and older values of the sample mean so we can replace X, bar T minus one with its representation

in terms of X bar at T minus two.

Then we can multiply out the one minus alpha term so that we get X bar at T minus two by itself.

Now we have three terms X bar T minus two, the sample at T minus one and the sample at time T.

The next step is, of course, to replace X, bar T minus two with its representation in terms of X,

bar T minus three.

From there we can do the same thing, multiply out the one minus alpha and get each of the terms by

themselves.

At this point you should see a pattern.

The number of individual samples keeps growing and the power on the one minus alpha term also keeps

growing.

If we keep repeating this pattern tee times, we end up with this expression involving a summation over

all the past samples from one up to T.

And of course, these weights are exactly exponentially decaying since Alpha is a number between zero

and one, one minus Alpha is also a number between zero and one.

And when you raise a number between zero and one to OPOWER, it gets smaller and smaller exponentially

as K gets larger and larger.

So how can we summarize what we've learned in this lecture, we've extended the concept of the mean

to include the exponentially weighted mean.

We can picture this by assigning weights to each of our samples with the arithmetic average.

Each of the weights is just constant, with equal weight for each sample.

With the exponentially weighted average, the weights decay exponentially, going backwards in time.

This means that the latest sample matters the most.

The second latest sample matters less.

The third latest sample matters even less and so forth.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.