All language subtitles for 021 Action Selection Policies-subtitle-en

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranรฎ)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

1

Hello and welcome back to the course on artificial intelligence.

2

I hope you're enjoying the course so far.

3

And today we're talking about action the selection policies.

4

All right let's get straight into it.

5

Previously we talked about adding a neural network to our simple learning and so far we are getting

6

quite into deep learning.

7

We've talked about the learning part quite a bit including adding some elements to it.

8

And today we're talking about this part we're talking about the acting.

9

So let's have a look.

10

So here we've got what we discussed about acting that once you input the values the parameters are the

11

vector describing the state agent is clearly in that environment then that is after all the learning

12

is done or even before the learning is done.

13

Basically we get all the q values so we're not interested in the learning right now we insist on acting

14

so once we have these key values how do we understand which one we need to use.

15

Well if you think about it.

16

Q values are simply predictions for the cube.

17

So as we did in the simple learning algorithm what did we do we just selected the one with the best

18

of the highest value.

19

Once we have the one with the highest IQ value we just take that action because it just brings us the

20

highest value and that we know that Duval's calculator's immediate reward that we expect to receive

21

Plus the DK factor times the value of the next date.

22

And it's a recursive calculation so why not why wouldn't you take the best value and that's kind of

23

the end of it.

24

But as you can see here it's not as simple here we're using a soft max function and this is where we're

25

going to talk about actual selection policies.

26

So here in reality we don't have to have just a software function.

27

We can have different action selection policies for example we've got Epsilon greedy Epsilon's soft

28

and we've got the soft Macs and those are kind of like the most commonly used action selection policies

29

of course there are others.

30

For instance the most basic one is a very simple action sociables it just select the best.

31

The one with the highest Q value.

32

But why doesn't that action pulse fly and why do we have different types of action pulse action selection

33

policies.

34

Well it all boils down to exploration versus exploitation.

35

And that is the core of reinforcement learning because we already talked about this a little bit that

36

your agent when it's operating in an environment it might predict certain queue values which might be

37

good and it might turn out great it might turn out that those are available and will be forced to explore.

38

So if we for instance in this case predict that Q2 is the best one and then it takes Q To takes action

39

to and it.

40

So from here to Section 2 and then it gets it gets a very negative reward.

41

Then the environment is forcing the agent to go and explode because now he's going to learn that oh

42

actually I thought Q2 was going to be very good but it turned out very bad.

43

So the results are not very bad.

44

So the networks can update itself so next time he's in the state he's going to probably eat my soul

45

just get to it.

46

You know like if it is a very very favorable so you might think that that's like you know you might

47

need a couple of times a couple of penalties or punishments in order to learn it is about action.

48

But maybe he'll already soon learn that I'm going to take a different action and take the wrist action

49

because now it has the best value.

50

So sometimes the environment forces the agent to take different to explore different actions but sometimes

51

the agent might get it find itself stuck in a local maximum it might find that it followed through through

52

its initial exploration and found that oh this is a pretty cool action like I'm going to go right here.

53

And that d'esprit collection.

54

But the problem is that it thinks is the best action simply because it hasn't explored is explored going

55

up his nose or going left is explore going right but it hasn't explored going down from that specific

56

state that it's in and now that it's kind of like biased towards this action and think thinks a good

57

action is going to keep taking it is going to keep getting.

58

He's going to keep taking is actually going to keep getting a good reward.

59

But what if this action would have been even better if this action would have been so much better that

60

if it knew about this action it would actually switch to this action but because it got stuck in a local

61

maximum is getting these good rewards is just going to be reinforced.

62

This is going to keep reinforcing itself that or the violence going to reinforce it that this is a good

63

action to take keep doing that.

64

But really the reality is that there's this other action that hasn't found yet or hasn't even explored.

65

That would have been much better.

66

So what we want to do is we want to come up with an actual selection policy that allows our agent not

67

to get stuck in a local maximum.

68

Yes it's important to you know keep doing the good actions that's the exploitation part.

69

We won't exploit what we've found.

70

But at the same time we still want to explore we never want to stop exploring as like in life you never

71

want to stop learning you stop learning you die.

72

That's things like that that when you're not growing you're dying or something got so you want to keep

73

learning and your agent wants to keep learning.

74

And that's where these action selection policies come in.

75

So we've got three you listed here so the first one is Epsilon greedy it's a very simple one it sounds

76

pretty complex in the sense that like it's got a cool name and usually things with surgical names.

77

It's actually not.

78

So basically what it does is it will select the one with the best Q value and epsilon like Epsilon you

79

might hear other places it's just like a selection policy.

80

So in this case we're using it to slick so our out of Al-Q values are by sales like the one with the

81

highest Q value all the time except for Epsilon percent of the time.

82

So for instance if you set epsilon to 10 percent then you're going to or 0.1 than 10 percent of the

83

time that the action is going to be selected at random.

84

So 90 percent of the time you're still going to be selecting the best action based on the highest value.

85

But 10 percent of the time is going to be selecting a random action.

86

Uniform it is going to be absolutely randomly taking an action or if you said epsilon to zero point

87

five for 0.05 that means that 95 percent of the time the agent is going to be taking the action with

88

the highest value.

89

But 5 percent of the time it's still going to be selecting and random action.

90

So it's going to be going out there and exploring.

91

So Epsilon's soft is very similar to the way that does kind of like why it's called FCL greedy because

92

then you're greedily selecting the action the good action except for that little episode.

93

Some of the time.

94

So the lower the EPS deal they'll lower the Lepp Epsilon the more greasily you're selecting that kind

95

of the action that is the optimal action and the less you're leaving the less chances you leaving for

96

exploration Epsilon's soft is the opposite.

97

So basically you're selecting at random you're selecting one minus Epsilon cent of the time.

98

So if you epsilons like 0.1 to 10 percent then only 10 percent of the time you're taking this action.

99

And 90 percent of the time you're selecting a random action.

100

So very very simple just inverted algorithms and a soft Max is kind of like the next step from or it's

101

it's a more advanced version I would say over epsilon of epsilon greedy algorithm although they both

102

have merit and they both have a place.

103

We're going to be using self-finance in our coding in our practical sort of thing.

104

So that's what we're going to talk in a bit more detail about soft max.

105

So let's have a look.

106

So let's move on to your next hopefully.

107

It's pretty clear about Ebsen agrees it's a pretty straightforward algorithm.

108

Select this one.

109

Most of the time except for sometimes go and explore.

110

And now we also see why it's important to do that exploration so that we don't end up in local maximums

111

in our in our optimization process so now we're going to talk a bit more about soft Macs.

112

There's a tutorial on soft marks at the end of the course.

113

I think it's an annex number two where we talk about the concept of Maxim's because you refresh a little

114

bit here so there we're talking about neural networks and by the way we're all going to be covering

115

convolutional.

116

We're not covering evolution neural networks in this section.

117

Of course in this section we're still using a vector.

118

But in the next section of the course when we're we're creating an AI to play Doom we are going to be

119

using convolutional neural network so it could be beneficial for you to look at in relational neural

120

networks and then take a self max function or you can learn a bit more about soft Max.

121

After you take the convolutional neural networks and of course later on.

122

But here's a quick refresher So here we've got our convolutional neural network which decides whether

123

it's a dog or cat.

124

So here we've got the voting process between these neurons and this one says that it's a it's got the

125

features you know the fluffy ears What's the pointed pointed face type of thing and the kind of the

126

features are the types of eyes the eye with eyes look all these features that belong to a dog.

127

So it's a 95 percent chance that it's a dog and the 5 percent chance that it's a cat.

128

But the question is how did we get in that Tauriel we're talking about how do we get these values to

129

add up to one.

130

Well whatever convolutional all our whole neural networks are the convolutional neural network plus

131

the fully connected Lares whatever it's bad out whatever the values that we apply to soft max function

132

are here.

133

This is where we introduced the formula for the soft next function.

134

Is what it looks like.

135

And then we got these values.

136

And so basically that's a quick refresher.

137

This is the formula for the soft Max.

138

It's what it does is it takes however many outputs you have doesn't matter.

139

It will take them and it will squash them all into values between 0 and 1 regardless of how big they

140

are just by it's for me you can see that there's a total sum at the bottom so these devices are going

141

to be zero and in.

142

And also the all these values are going to add up to one always.

143

And so that's that's very beneficial for us because when we're using the soft max function what happens

144

is we get these values we select this best view value.

145

But in reality what happens is these values that we get there are actual numbers right.

146

So this is some kind of numbers.

147

They don't have to all add up to one and don't have to be between 0 and 1.

148

Just some numbers.

149

But when we apply soft Max we don't just select the best one we actually get numbers like that so we

150

get our numbers in the range between 0 and 1 and that are also that also add up to 1.

151

And so what other thing do we know that adds up to one.

152

Well probabilities we know that probabilities always have to add up to 1 so that is why we can say here

153

we've got q values but here all of a sudden we've got soft or we've got probabilities.

154

So we can say that the likelihood of this being the best action is 90 percent.

155

This lesbian section 5 percent 2 percent 3 percent because we know the higher your value the better

156

the action.

157

So if we squash them to 0 to 1 then these become possibilities and we can deal with them as such.

158

And therefore now is when the action is selected and that's how we come up with Q2.

159

But if you look at it closely this isn't a strict 100 percent and these are not Saroo 0 percent.

160

So this is a 5 percent to 3 percent.

161

So the most natural way to apply the soft Max in order to preserve exploration in the algorithm is to

162

use these exact probabilities as how often we're going to be taking that action.

163

So these probabilities actually present the distribution of these actions that we're taking so basically

164

soft Max makes it very easy for us to come up with a way to combine exploitation and exploration.

165

So the best the best action will always have the high probability because it has highest Q value and

166

therefore here we're going to be just going to use these as our distribution or we're going to say okay

167

we're going to be taking Q2 90 percent of the time but 5 percent of the time we still get to be taking

168

Q1 and 2 percent of the time we get to 3 and 3 percent of the time we're going to be taking Q4.

169

And the beauty here is also that as these values update as and as the agent goes through the network

170

more and more and more it becomes more familiar with with the environment and therefore these updates

171

so this value for instance might become like it might might ascertain that this value is actually less

172

or this actually is higher and so these probabilities will also change as an agent goes through.

173

So even though here we've got Choo-Choo.

174

Nobody is to say that sometimes 5 percent of the time to be more precise we'll be selecting Q1 as the

175

action to take and sometimes or action one will be taking action one.

176

Sometimes will be taking action through a two action three two percent of the time and action for will

177

be taking about 3 percent.

178

So every action has a chance to play in this process as long as we have enough iterations an agent goes

179

through lots and lots of times through these states that they're in.

180

And that's that's how this that's how any kind of deep learning algorithm works that you want to do

181

this many many times so that you learn from experience and therefore as you can see here it's a very

182

natural transition to.

183

We're not just randomly like an Epson angry algorithm and not just randomly selecting the actions we're

184

selecting them based on their soft max values which makes it makes it like has some logic behind it

185

not just not just that random 10 percent of the time we're selecting a random action but there's some

186

logic behind how we're doing it and based on the key values that we've explored.

187

And so that's the action selection policy that we're going to be using in this course.

188

You're welcome to definitely check out Ebsen greedy action section Polsce if you like but we're going

189

to be predominately using the soft Max action section policy and I've got an interesting reading for

190

you.

191

So this is called adaptive Epsilon greedy exploration in reinforcement learning based on value differences

192

it's the 2010 article.

193

And it's interesting because Mike Michel I'm not sure how to pronounce Michelle and Miquel toxic introduces

194

a different type of Algren's and adjusted Epsilon greedy algorithm and called the VDB VDB algorithm

195

or epsilon greedy VDB algorithm you can see here.

196

And he actually compares compares to the Ebsen greedy and soft Max and it's an absolute greedy algorithm

197

which basically the main idea behind it is to adjust the value of epsilon depending on the state the

198

agent is in.

199

So if if the agent is very certain about the state in then Epsilon should be smaller so they should

200

be less exploration if the agent is answered Epson's should be higher should be more exploration.

201

So it is a 2010 article.

202

I'm not sure if it's if this new proposed algorithm is widely used or is as being accepted in the community

203

or or if artificial Times has kind of a way from this this suggestion.

204

But nevertheless it will definitely help you reinforce your knowledge about action selection policies

205

which we discussed the Epsom Ingredion the soft Naxal help you ill give you an opportunity to compel

206

Subha site and also see in which direction people actually think when they want to improve artificial

207

intelligence so if you're ever planning on creating really interesting algorithms that are pushing the

208

edge of Elche artificial intelligence and pushing the envelope in this space then this could be a good

209

way for you to see in which direction people think sometimes when they're trying to improve the norms

210

of artificial intelligence or the norms that existed back then in 2010.

211

So there we go.

212

Hopefully you enjoyed today's tutorial about the action selection policies and we learned about abseil

213

greedy Epson salt and the soft Macs and now you're even more prepared for the practical side of things.

214

And on that note I look forward see your next step.

215

And until then enjoy AI.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.