Afrikaans
Akan
Albanian
Amharic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
In this lecture, we are going to discuss another kind of moving average, the exponentially weighted
moving average.
Note that some other names for this are exponential smoothing and the low pass filter.
So if you've taken some of my other courses before and you've heard me use those terms, recognize that
this is the same thing, in fact, that this kind of moving average is very applicable in many areas
of machine learning statistics, finance and signal processing.
So you will generally see it pretty often.
So what is the exponentially weighted moving average?
I want to break this lecture up into two parts, the first part is the short summary.
If you want to only watch this part and then skip to the code, that's fine.
The second part of this lecture will be an optional in-depth discussion about why the exponentially
weighted moving average has its name.
You can opt to watch this if you want to get a better understanding of why and how this works.
OK, so what's the short summary, as you know, the arithmetic mean can be calculated by taking all
of your samples, summing them together and then dividing by the number of samples, the exponentially
weighted moving average is calculated differently.
In fact, it's calculated kind of on the fly or in an online manner.
It says that the moving average at time T is equal to some constant alpha times.
The sample at time T plus one minus alpha times the previous moving average at time at T minus one.
In other words, at each step, the new moving average is the weighted some of the new sample and the
old moving average.
All right.
So that's pretty much it.
It's not a terribly complicated calculation.
Of course, without further analysis, it's not clear why this is an average and it's not clear why
it's exponentially weighted.
The next part of our short summary is this how do we do it in code similar to the simple moving average
we call a function on our series or our data frame called IWM.
This returns and IWM object, which is similar to the rolling objects we saw previously.
It has a similar set of functions such as mean variance, covariance and so forth.
To discuss a practical issue, what value of Alpha should we choose Alpha is something like a decay
factor.
Typically, Alpha is chosen to be a small value between a zero and one like zero point one or zero point
to it might help to look at some extreme cases.
So let's say we choose Alpha equals one.
That means set the average to be just the latest value of X..
In this case, all we're doing is copying X and therefore it's not really an average at all.
On the other hand, let's say we set Alpha equal to zero, then all we're doing is copying the previous
average and we're not taking into account any new samples intuitively.
Then if we set off a very close to one that says new samples matter much more in the old average matters,
much less, you can imagine this will lead to a much more noisy time series which will more closely
match the original if we set Alpha very close to zero that says new samples matter much less and the
old average carries much more weight.
In this situation, you'll get a much smoother time series and it will take a much more drastic change
in X to affect the moving average.
OK, so now that the short summary is complete, if you want to know the details behind the exponentially
weighted moving average, keep listening.
Let's suppose we want to calculate the usual arithmetic sample mean using the formula for the sample
mean.
You might suggest that this is quite obvious.
Just take all the values of X that you've collected, add them all together and divide by the total
number of X is that you have.
The question is what's wrong with this?
I'll give you a minute to think about it.
So please pause the video until you think you have the answer.
All right, so hopefully you thought about why calculating the sample mean naively might not be such
a good idea.
What if we have a lot or even an infinite amount of data?
Obviously, our computers or our servers don't have an infinite amount of space.
And even if they did, calculating a summation is of T.
So the more data you have, the longer it will take and that will increase linearly with how much data
you've collected.
Here's my claim.
I claim that you can make the calculation of the sample mean of one on each step in both space and time
complexity, no matter how much data you collect again as an exercise before moving on to the next slide.
I want you to think about how this might be the case.
Please pause the video if you want to take a moment and think.
OK, so hopefully you thought about how you might calculate a sample mean using constant space and time.
The key is that you can calculate a sample mean using the previous sample mean let's call the sample
mean after collecting samples X bar subscript T, this means that the sample mean after collecting T
minus one samples is X, bar subscript T minus one.
We can write down the definition of both of these, which I hope is pretty obvious.
Now that you know the metric, let's again make this an exercise.
Can you express Esbati in terms of X, bar T minus one, please pause the video until you've tried this
on your own.
OK, so here's what you can do.
First, you take Esbati and split up the summation so that you only sum up to T minus one.
Then you leave X subscript T by itself.
This is just the last sample you've collected.
The next step is to realize that the sum of the ex towers from one up to T minus one can be expressed
in terms of X bar subscripts, T minus one.
We just have to rearrange the equation from earlier.
It's clear that this sum is just T minus one times X, bar T minus one.
We can substitute this into our expression for Esbati to get the sample mean at time t in terms of the
sample mean a time T minus one.
One interesting thing you can do, although it's not totally clear why you'd want to do this at this
time, is split up the formula as follows.
The first step is to multiply out the one over Tetum that gives us T minus one over T as the first coefficient
and one over T as the second coefficient.
The second step is to simplify T minus one over T to one, minus one over T.
At this point we can just leave this as is this is the form that we want.
We have one term with the previous sample mean and we have one term with the latest sample.
What's important to recognize about this equation is that we have discovered a way to calculate the
sample mean that does not depend on carrying around all of the samples you've ever collected.
All you need to have is the previous sample mean the latest sample and the number of samples you've
seen in total.
The next question to consider is, what if we believe that recent data matters more than past data?
If we look at our equation carefully, we see an interesting characteristic.
Remember that as we collect more and more samples, the value of tea is increasing.
That means as we collect more and more samples, the weight that we give to the latest sample decreases.
We can see that the weight that we give to the sample is exactly one over tea.
Now, although this might make you think that the influence of each sample somehow decays over time,
remember that this is not true because this is still just a regular arithmetic mean.
But what if we want recent data to matter more, what would happen if instead of making the way one
over tea, we simply make it a constant alpha?
Well, then this is exactly the exponentially weighted moving average.
The basic idea is instead of giving less and less weight to each new sample, we now give a constant
weight to each new sample.
Let's see how this affects the influence of each sample overall.
The next question we want to answer is, how does this update actually implement an exponentially weighted
moving average?
Can we show that this is true?
And in fact, it's not too difficult at this point.
What we can do is just keep recursively plugging in.
Older and older values of the sample mean so we can replace X, bar T minus one with its representation
in terms of X bar at T minus two.
Then we can multiply out the one minus alpha term so that we get X bar at T minus two by itself.
Now we have three terms X bar T minus two, the sample at T minus one and the sample at time T.
The next step is, of course, to replace X, bar T minus two with its representation in terms of X,
bar T minus three.
From there we can do the same thing, multiply out the one minus alpha and get each of the terms by
themselves.
At this point you should see a pattern.
The number of individual samples keeps growing and the power on the one minus alpha term also keeps
growing.
If we keep repeating this pattern tee times, we end up with this expression involving a summation over
all the past samples from one up to T.
And of course, these weights are exactly exponentially decaying since Alpha is a number between zero
and one, one minus Alpha is also a number between zero and one.
And when you raise a number between zero and one to OPOWER, it gets smaller and smaller exponentially
as K gets larger and larger.
So how can we summarize what we've learned in this lecture, we've extended the concept of the mean
to include the exponentially weighted mean.
We can picture this by assigning weights to each of our samples with the arithmetic average.
Each of the weights is just constant, with equal weight for each sample.
With the exponentially weighted average, the weights decay exponentially, going backwards in time.
This means that the latest sample matters the most.
The second latest sample matters less.
The third latest sample matters even less and so forth.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.