Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome back to the course on deep learning in today's tutorial we're talking about gradient
2
descent.
3
What we learned previously was that in order for a neural network to learn what needs to happen is back
4
propagation and that is when the error the difference or the sum of squared differences between y hat
5
and Y is back propagated through the neural network and the weights are adjusted accordingly.
6
So we saw that and today we're going to learn exactly how these weights are adjusted.
7
So let's have a look.
8
This is our very simple version of a neural work a percept Trauner a single letter feedforward neural
9
network and what we can see here is this whole process in action where we've got some input value then
10
we've got to wait then a activation function is applied.
11
We have we get y hat and then we compare it to the actual value we calculate the cost function.
12
So how can we minimize the cost function.
13
What can we do about it.
14
Well one approach to do it is a brute force approach where we just take all lots of different possible
15
weights and look at them and see which one looks best and what we do is for instance we would try out
16
for let's say for example a thousand weights and we'd try them out that would get something like this
17
for the cost function and this is a chart of on the Y axis of cross-functional the vertical axis on
18
the horizontal axis of y hat.
19
And because you can see the formulas I had minus Y squared.
20
This is what the cost function would look something like that.
21
And basically you'd find the best one is over here.
22
So very simple very intuitive approach.
23
Why not do this brute force method.
24
Why not just try out a thousand different cost for a thousand different parameters or inputs for weights
25
and see which one works best.
26
You'll find the best one that way.
27
Well if you have just one way to optimize this might work but as you increase the number of weights
28
increase the number of Synopsys in your network you have to face the curse of dimensionality.
29
And so what is the cause of dimensionality.
30
The best way to describe this or explain it is to just look at a practical example.
31
So remember this example we had when we were talking about how neural networks actually work where we
32
were building or running a neural network for a property valuation.
33
So this is what it looked like when it was trained up already well when it's not trained before it's
34
trained before we know which one what are the weights.
35
The actual neural network looks like this.
36
Right because we have all these different possible synopses and we still have to train up the weights
37
and here we have a total of 25 weights so four times five at the start plus five more from the hit out
38
there 25 weights total.
39
And let's see how we could possibly brute force 25 ways.
40
This is a very simple neural network right here.
41
Very simple just one hit in there and how could we brute force our way through a neural network of this
42
size.
43
Well there's some simple mathematical calculations.
44
We have 25 weights.
45
So that means if we have a thousand combinations that we're going to solve for every weight the total
46
number of combinations is 1000 to the power 25 or a thousand or 10 to parse any five different combinations.
47
Now let's see how Sun the way to tohu light the world's Fosse's supercomputer as of June 2016 what how
48
would it approach this problem.
49
Right so Sunway tie who light.
50
It looks like this is a whole huge building pretty much for this one supercomputer and it got the Guinness
51
World Record for being the Fosses supercomputer.
52
Right now it is the fastest supercomputer in the world and some way tie lights can operate at a speed
53
of 93 of flops.
54
Flop stands for floating operation per second.
55
So it can do ninety three to the power oil.
56
Times ten to the power of 15 floating operations per second.
57
That's how quick it is in comparison.
58
Average computers right now they do like just over several gigaflops and so on.
59
So it like kind of those ranges.
60
Less than TEI Sunway type light.
61
So suddenly it's all a lie it is in the forefront of technology.
62
And let's say hypothetically that it can do one one test one combination of four on your own network
63
in one floppy disk and one floating operation that is not possible that is not practical because you
64
need multiple floating operations to test out a single weight in your own little.
65
But even Let's let's give it a head start.
66
Let's say that it can do it in a ideal world it can do that in one floating operation it can do one
67
test per one floating operation.
68
That means it will Doddridge still require tend to of any five.
69
Divide by ninety three times ten to about 15 seconds to come to run all of those tests to brute force
70
through that network.
71
So that means one or approximate tend to power 58 seconds and that is the same as tend to the power
72
of 50 years.
73
That is a huge number that is longer than the universe has existed and that is definitely not going
74
to simply this number is so huge its just definitely not going to work for us at all in our optimization.
75
So there we go.
76
This is a no no.
77
Even on the world's fastest supercomputer Sunway tail light.
78
So we have to come up with a different approach how are we going to find the optimal weight.
79
By the way this our neural network was very simple what about if the neural networks looks like something
80
like this or even a greater than that then yeah its just not going to happen at all ever.
81
So the method were going to be looking at is called gradient descent and you may have heard of it already.
82
If not we will find out what it is right now.
83
So theres our cost function and now we go into see how we can foster for kind of a faster way to find
84
the best option.
85
So lets say we start somewhere you're going to start somewhere.
86
So we start over there.
87
And from that point in the top left what we're going to do is we're going to look at the angle of our
88
cost function at that point so we're just going to basically that's what's called gradient because you
89
have to differentiate.
90
We're not going to look at the mathematical equations.
91
We will provide some tips on additional reading at the end of the next lecture.
92
But basically you just need to differentiate find out what the slope is in that specific point and find
93
out if the slope is positive or negative.
94
If the if the slope is negative like in this case means that you're going downhill so to the right is
95
downhill to the left is uphill.
96
And from there it means you need to go right.
97
Basically you need to go downhill.
98
And that's what we're going to do.
99
Boom takes a step right.
100
The ball rolls down again.
101
Same thing.
102
You calculate the slope and the slope is positive meaning writer's uphill left is downhill and you need
103
to go left and you're on the ball down.
104
And again you calculate the slope and you're all the bull right there you go so that's how you find
105
in simple terms that's how you find the best WAITES The best situation that minimizes your cost function.
106
Of course it's not going to be like a ball rolling is going to be a very zigzag type of approach but
107
it's easier to remember or kind of it is more fun to look at it as a ball rolling.
108
But in reality yes you just it's going to be like a step by step approach is going to be a zigzag type
109
of method.
110
Yeah and also there's there's lots of other elements to it.
111
There's things like for instance why like why does it go down why does it not go way over the line so
112
it could have jumped out of this gone upwards instead of downwards and things like that so there are
113
parameters that you can tweak.
114
And again we will mention where you can find out more on that.
115
And plus we'll have this in practical application but in the simplest intuitive approach this is what
116
is happening.
117
We are getting to the bottom by just understanding which way we need to go.
118
Instead of brute forcing through thousands and thousands and millions and billions and quadrillions
119
of combinations.
120
We can just simply every time have a look at where is where which way is it sloping so right like your
121
or you imagine you're standing on a hill.
122
Which way does it feel that it's going downwards and whichever way it is going down and you just keep
123
walking that way you like take 50 steps away and then you assess again OK which way is it going downwards
124
this way.
125
OK and I'll take 50 steps or less take 40 steps that way.
126
So it gets less and less and less as you get closer.
127
So here's an example of gradient descent applied in a two dimensional space.
128
So that was a one dimensional example.
129
Here we have a two dimensional space for the gradient descent as you can see it's getting closer to
130
the minimum and it's also called gradient descent because you're descending into the minimum of the
131
cost function and find that he has a gradient descent applied in three dimensions.
132
This is what it looks like if you projected onto two dimensions you can see zigzagging its way into
133
the minimum.
134
So there you go that it was gradient descent index of Tauriel We'll talk about stochastic.
135
Gradient descent is really a continuation of this tutorial.
136
And I look forward to seeing you there.
137
And so next time enjoy deep learning.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.