Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome to this new tutorial.
2
All right now things are going to get slightly more difficult.
3
We're going to get into the heart of that as is immoral.
4
So make sure you understand this model first.
5
I highly recommend to watch Carol's intuition lectures first.
6
So if that's the case and you're ready to go and you're all fresh let's tackle this.
7
So first of all let's remind the context we got the output of the neural network SSD which is why.
8
And then from this output why we extracted the important informations that we need and we extract them
9
in this detections variable which is a tensor by taking the data of Y which corresponds to the tensor
10
part of the torch variable.
11
Remember the torch marble is composed of two elements a torched answer and a gradient.
12
And by taking the data attribute of the output Y which is a variable.
13
Well we get the first element of this torche variable which is the torch tensor.
14
And now we're going to see what this tensor contains exactly.
15
So what is this information exactly.
16
What is this detections tensor and what does it contain.
17
So this detection sensor I'm going to try to come in here to make sure everybody understands this detections
18
tensor right now contains four elements.
19
So I'm going to put them into brackets.
20
The first element is the Bachche because we created this fake dimension of the batch with this and squeeze
21
function here.
22
So the first element of this detections tensor is this batch.
23
So I know I told you about a batch of inputs but that's the same for the output.
24
We also have batch of outputs associated to the same batch of inputs.
25
That is the batch of inputs contains several inputs.
26
Of these inputs we get the outputs and these upwards are all into about and that's where this first
27
element of detections is you can write sections here detections equal.
28
All right.
29
So for us Almont the batch.
30
Now the second element is the number of classes.
31
So what do I mean by classes.
32
I simply mean the objects that can be detected.
33
So for example one class will be the dog.
34
Another class will be the plane.
35
Another class will be there but another class will be the car.
36
So each class corresponds to each object that can be detected.
37
And the second element of this detection sensor is the number of classes that is the number of objects
38
that was detected in the input image.
39
So a number of classes and now the third element is the number of occupants of the class.
40
I'm going to write it down number of utterance of the class.
41
So what does it mean.
42
Well it means that for example for Class Number two the class number two which let's say a response
43
to the dog.
44
Well the number of arguments will be the number of the utterance of the dog.
45
So in the video of the funny dog that we watched in the first title of this module there was one dog.
46
But imagine there are several dogs to detect in the video.
47
Let's say there are two dogs.
48
Well we would have two numbers of occupants.
49
The first occupants of the dog that is the first dog on the video and the second argument of the dog
50
corresponding to the second dog in the video.
51
So we would have the actor in zero and the occupants won.
52
But in the video we only have one dog.
53
So we will only have one number of utterance when accurate.
54
All right so that's the third element of this detection sensor and the fourth element will be a couple.
55
It's up all of five elements.
56
So I'm going to write them now or we will get crazy.
57
This is a couple of the following five elements.
58
First Ullman's is the score.
59
Second element is x 0 third element is y zero then fourth element is X.
60
And the last element is why when and why does this double the belt.
61
Well for each occurrence of each class in the batch we will get a score for this argument.
62
And of course the coordinates of the upper left corner of the rectangle defect in the end and the lower
63
right corner.
64
And what are these scores about.
65
Well these scores are going to go from low to high.
66
We will have a score for each arguments of each class such that if the score is lower than 0.6 then
67
the arguments of the class won't be found in the image if it is higher than 0.6 then it will consider
68
it to be found.
69
So that's what this course is about.
70
It's like a threshold for each occupants of each class.
71
We will get a score if it is more than 0.6 then no arguments will be found.
72
And if this course hadn't 0.6 the arguments will be found and we will get the coordinates and the upper
73
left corner and the lower right corner of the detected object.
74
All right so now that we clearly understand what's inside the detections tensor we can move on to a
75
for loop.
76
Yes there we go we have to make a follow because we have to iterate through all the classes and then
77
through all the arguments of the classes we're going to look for a certain number of occupancies for
78
each class.
79
We're going to get the score for each of these other answers and we'll make an IF condition to say that
80
if discours had an open six we keep the utterance and if the score is lower than 0.6 we reject the accurate.
81
All right.
82
Are you ready.
83
Perfect.
84
So let's start this for you.
85
So for then as I just said we're going to iterate through all the classes and therefore I'm going to
86
add here.
87
I for I in range detections that size 1.
88
So detections that size one is exactly the number of classes.
89
So this that I highlighted is the number of classes so we're just making a full loop from 0 to this
90
number of classes.
91
There is a number of detectable object.
92
And so for all these classes we're going to go inside the for loop and we're going to start by introducing
93
the variable J which will correspond to the utterance exactly the utterance.
94
So why is the class and j will be the utterance of the class.
95
And now we're going to start a second loop not a fluke this time it's going to be a while loop because
96
you know we're going to put that condition into the loop which is actually more efficient.
97
So it's like a loop combined with the if condition at the same time.
98
And the trick to do that you're going to see is very intuitive is to do this well loop.
99
And since we only want to keep the occupancies for which the score is higher than 0.6.
100
Well we simply need to take the detections of the utterance J of the class I.
101
But then let's not forget the Bache zero zero and then we're going to get the score which is the first
102
element of this double here.
103
So we simply need to add zero therefore since zero corresponds here to the index of the score Jaker
104
response to the index of the occupants of the class I.
105
Well this detections of 0 0 is exactly the score of the occurrence J of the class.
106
And therefore Well the score of the current state of the class is larger than 0.6 than what are we going
107
to do.
108
We're going to keep this accurate and how are we going to keep it.
109
Well we are going to keep it in the variable that we're going to call Peetie for point because we're
110
going to keep that argument by keeping the point and therefore the coordinates of the upper left corner
111
and the lower right corner of the rectangle detecting arguments J.
112
The class II.
113
So Peetie will be the detections and we're going to take the same is zero.
114
That is the Bache then the class II then the arguments J and then be careful try to guess what I'm going
115
to type here.
116
Well it's not going to be zero because we're no longer interested in the score.
117
We are now interested in these four coordinates.
118
Exit row 1 0.
119
That is the coordinates of the upper left corner of the rectangle and X one y one that is the coordinates
120
of the lower right corner of the rectangle.
121
And therefore we want to take the last four elements of this table.
122
And so we're going to take the range from one to the end and the trick to take the range from one to
123
the end is to use this the range from one colon and nothing and nothing means to the end.
124
Perfect.
125
So here with this trick we're taking X you are wise you are x 1 and white 1.
126
So now we have our coordinates.
127
That's perfect.
128
But now remember that we created this scale tensor to do this normalization of the coordinates between
129
0 and 1.
130
And that's exactly where we're going to use the scale tensor and to use it we simply need to multiply
131
all this by this scale tensor and that will apply to normalization which will give us the coordinates
132
of these points at the scale of the image.
133
And finally we need to do one last thing we need to transform this tensor for coordinates that is all
134
this you know all this is a tensor right now it's a torch tensor.
135
Well since now we're going to use open Svea to draw the rectangles thanks to the upper left corner and
136
the lower right corner coordinates that we have.
137
Well we need to put that back into an umpire.
138
Because open sea works with an array.
139
We're going to use the rectangle function you know to draw the rectangles exactly like what we did in
140
the first module but this rectangle function works only with non-pay arrays and not with torch tensors.
141
So we just need to convert that back into an umpire.
142
And to do this there is nothing more simple.
143
We just used a pi function like that.
144
And here we go we have Arnon by Array containing the four normalized coordinates of the upper left corner
145
and the lower right corner of the occurrence J of the class.
146
And now since we have these coordinates Well we can draw the rectangle we're going to do that still
147
in the well loop obviously.
148
And you know how to do that.
149
We take open city CB2 then the rectangle function and remember the arguments we first need to input
150
frame the image.
151
Then the second argument is zero.
152
And that is Peachi of index 0 because pittie contains x 0 1 0 X Y N Y one.
153
So PITI 0 will be zero.
154
Then the third argument to be y 0 and that is P-T 1.
155
Then the next argument is x 1 and that is Peachi to and the last one is why one.
156
And that is P-T three because zero pity one 52 and three are exactly these four coordinates in the same
157
or x y z x y y y.
158
And now we're just going to add a safety.
159
We're going to convert these values of the coordinates into integers.
160
It's always safer to do that and to do this we're going to use the int function to convert them int
161
int
162
and and now be careful.
163
The second argument of this rectangle function from open city is actually the coordinates of the upper
164
left corner of the rectangle.
165
So these two coordinates here should go into the same second argument.
166
So I'm going to put some parentheses around these two coordinates and same for these two coordinates
167
that should go into one same argument.
168
That's the third argument of the rectangle function and these correspond of course to the coordinates
169
of the lower right corner of the rectangle that is detecting the object.
170
So again I'm putting some parenthesis around them.
171
All right so first argument frame second argument the coordinates of the upper left corner.
172
And third argument the coordinates of the lower right corner of the detector rectangle.
173
Now next argument there is two more arguments to go the next one is the color of the rectangle.
174
And we're going to choose the red color that is coded in the RGV code by 255 0 0.
175
All right so that's our next argument.
176
And now the final arguments of our rectangle function is.
177
Remember the thickness of the text to display and as in the Mudgal one we're going to choose to.
178
All good.
179
And now let's move on to the next step which is to print the label onto the rectangle because we will
180
have several objects to detect.
181
So we want to print the label Doug onto the rectangle that is detecting the drug and then the label
182
person onto the rectangle that is detecting the person.
183
So it's indeed quite useful to print a label so to print these labels we're going to use open city again
184
CB2 and then we're going to use another function.
185
You can see that open city has many functions but the function we want right now is called put text.
186
There we go.
187
This one put text and this function takes several arguments.
188
Not exactly the same as a rectangle function but pretty close.
189
The first one is of course our frame.
190
The second one is the text to display and to get this text we need to use the label map shortcut name
191
that we gave to classes remember what lessors is this dictionary that maps the names of the classes
192
with numbers and we can get the label of the text we're interested in.
193
That is the class we're dealing with right now which is a Class I.
194
Well we get the index I minus one y minus one it's because indexes in Python started 0.
195
So the index minus one is actually the index of the ice class.
196
All right.
197
And then a third argument will be the position of the text where we want to display the label.
198
And we're going to display it at the upper left corner of the rectangle just above the upper left corner
199
and therefore we need to take these coordinates because they correspond exactly to the coordinates of
200
the upper left corner.
201
So that's our next argument then we need to choose a fund.
202
That's our next argument.
203
And we're going to pick the following fund which we get from our open library.
204
And the name of the fund is called fund her say and not complex but simplex.
205
That's a nice one right.
206
Then we need to choose a size of the text.
207
We're going to choose size two then a color of the text.
208
We're going to choose the following color 2 5 5 5 5 and 2 5 5 then a thickness of the text.
209
We're going to choose two again.
210
And finally we want our text to be displayed continuously you know with continuous lines and not little
211
dots.
212
And to make sure we get this we need to hear the last argument which is going to be C-v to that line.
213
Right that will just give us some continuous lines to display the text and not little dots and that's
214
it.
215
We displayed a nice label onto our detector rectangles and now eventually we have one last thing to
216
do inside this while loop.
217
And then after that one final thing to do over all this one I think you have to do inside this while
218
loop is of course to increment J because right now we're dealing with the arguments 0 because Jaycar
219
0 and we did not do a for loop we did a while loop and then a while loop we must not forget to increment
220
the iterative variable that is J.
221
So in other words we just need to deal with the next occupants of the class.
222
I and to do this we just need to increment J like that J plus equals 1.
223
Perfect.
224
So now we can get out of this while loop and also get out of this for loop because we did exactly what
225
we had to do for all the objects that can be detected.
226
And now the final thing that we need to do is just to return the frame and that's it.
227
We have the frame returned with the detector rectangles on any object that is part of the training and
228
as is the model.
229
So you know that as the model was trained to detect between 30 to 40 objects here in this loop we look
230
for all these objects and several possible occupancies of these objects.
231
We approach them as core and just matching scores high enough.
232
We keep it and therefore we have several objects detected on this return frame with the rectangles and
233
the labels.
234
So congratulations.
235
That was the hardest part.
236
I hope that's OK.
237
Now the rest will be much easier and soon enough we should be able to see that output video with the
238
several detected object.
239
I can't wait to show you this.
240
And until then enjoy computer vision.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.