Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome back to the course on computer vision got an exciting section ahead and we're going
2
to kick off by talking about how SS D or the single shot multi box detection algorithm is different
3
to others.
4
This looking at this image.
5
So here is an image and.
6
Yeah.
7
So we can see some sheep walking around on a field.
8
And the question is how would normally a computer or object detection algorithms detect these sheep
9
Well it's the sheep as you can see quite small and you not only need to look at the image as we learned
10
in the Anik some computer unconventional neural networks there where we're just looking at image and
11
saying if something is present on the image or not here we actually need to detect exactly where the
12
sheep are and put those boxes around them to you know say that the sheep is here or the sheep is here
13
and so on and so what normally these algorithms would do is they would try to save time or try to save
14
the computational cost by through certain tricks and hacks and of course all of them even less as D
15
has its own trick and hacks.
16
But what changes is the tricks and hacks that those things change and so previously most are all like
17
a huge portion of the algorithms.
18
What they would do is they would engage in something called an object proposal methodology.
19
They would have some object proposal techniques and basically you've got an image and of course is a
20
huge image Normally these images are resized to smaller images and they were like the algorithms work
21
with smaller images and then maybe they resized back in and so on.
22
But it is like even on a small image you still have the same problem that if you were to for instance
23
convert this whole image with rectangle and then try to you know try to guess in which one there's a
24
sheep or not let's say here's a rectangle Here's a rectangle you can see that most of them don't actually
25
have the sheep and you would be spending a lot of time finding these these sheep and a lot of can be
26
additional cost because of competition cost and Dalgard would be slow and that's a big problem.
27
And so what they would normally do is do this thing called object proposal where they would come up
28
with a way to break down the image segmented into parts to suggest where they could potentially be objects
29
and where they could not be objects for instance by looking at the gradients in color.
30
So if you look at this big chunk over here it's all green.
31
There's not much there's not much gradient there's not much changes not much chare like it's all gradual
32
and there's no red there's no edges there's no there's a little bit of edges but they're all kind of
33
like it's all about the same kind of texture same same color more or less.
34
And so there's no.
35
Nothing changes radically even from like one neighboring pixel to the next and therefore you could make
36
a proposal that there is no object in here.
37
And then you would take this and you know maybe you would make a proposal that like it depends on how
38
you break it up break it up like I can this you might say you know there's an object here then you might
39
say that there's an object here because you know you can see the shadow.
40
And again we're not just detecting sheep we're detecting idealy in the most complex of all challenges
41
predicting any types of objects that were officers so for instance you know sheep or we could be looking
42
for mice and we could be looking for cars and we could look looking for giraffes anything.
43
So it's very dangerous to discard anything any object because then you might miss out.
44
And then on the other hand if you take this segment off if it's segmented like that or however it's
45
segment you might you would see that you would first calculate the gradient that there is some kind
46
of change happening over here and there is an object proposal for this segment and then we will dig
47
into this further.
48
And so that's that way they save time.
49
But as you can imagine it's they would sacrifice accuracy because there's all lots of other details
50
to it but in short they would say can be additional cost but they would sacrifice accuracy and you know
51
Miss objects sometimes and when it's kind of like it might be OK in some implication in other applications
52
the like very very high percentage of accuracy especially in self-driving cars and things like that
53
you can't afford to sacrifice accuracy but at the same time you also need for computational efficiency.
54
So because it needs to happen all in all these happen real time.
55
Yeah and so what did S's come up with.
56
Well let's have a look at this slide is the they are the authors of SSD.
57
They came up with a brand new solution.
58
Well it's a very interesting solution where they do everything in one shot.
59
That's why it's called the single shot.
60
Box detection algorithm.
61
It it can.
62
It just looks at the image once it doesn't have to go back to it.
63
It doesn't have to do this object proposal it doesn't have to then run many convolutional neural networks.
64
That's the part where the communicational cost comes in.
65
So if you go back one slide here if you like break it down and then you run a of neural network for
66
every single one of those rectangles to detect.
67
OK is there a ship there is there a ship there is there a ship there is this ship that that's very can
68
be really costly and that's where you need to do these jobs proposals to cut out parts that are definitely
69
out of the question so you don't have to run Minicom the conventional network many times.
70
But in the single shot multi-book detection algorithm what they did is they only look at it once there's
71
only one algorithm one other algorithm in the world that does the same thing.
72
It's called the YOLO algorithm.
73
It's stands for you only look once and in their paper they actually compare the two.
74
This is does a better job.
75
They prove empirically that Issas is better.
76
But apart from that apart from yellow and the there is no algorithm other algorithm in the world that
77
or at the time there was no other algorithm in the world that actually only looks at the image ones
78
and that that's it.
79
It only does that one can volitional neural networks all of them would come back to the image several
80
times or many times.
81
And one second.
82
OK.
83
And.
84
Yeah.
85
And so now we have this and that's we're going what did they do.
86
So basically all of those boxes as you will we will discuss in the further trials will the boxes and
87
boxes that are up here on the image all of them go through the network at the same time.
88
And moreover not just to go through the network at the same time network also remembers which boxes
89
dealing with the boxes.
90
They train up on detecting objects and detecting the borders or objects correctly.
91
And moreover in this same network and now the really cool thing that they introduced is that there's
92
many convolutions to reduce the image size so you can see stars with 300 image 38 pixels 19 10 5 3 1
93
as the image size of the juices.
94
And that helps deal with different sizes of objects on the original image and we'll talk about that
95
as well so they combined lots of hacks in this one algorithm and probably one of the most important
96
ones is of course that they do everything in a single shot.
97
All these it all happens in one huge convolution as it goes through this network and that really helps
98
save on the communication efficiency.
99
So that's in a nutshell what how this is different.
100
So it does everything in a single shot.
101
It can deal with training of identifying objects and also identifying the borders and it can deal with
102
scale all all at the same time.
103
And there's only one other algorithm that rival the D in that architecture and that set up it's called
104
the yoa algorithm.
105
And yeah and they talk about more in the paper so that's I'm not sure we'll go into those components
106
in more detail in upcoming tutorials but we'll go into them in a very intuitive sense if you would like
107
to get the math behind it all the more strict approach to what's going on.
108
The best place to look at is of course the original paper by way you and others is called Single Shot
109
multi box detector you can find our archive there's a link.
110
And yeah this is quite a it's a very new algorithm and so therefore there isn't much written about it
111
but nevertheless this paper is not as you can imagine the ultimate source of truth and into its Gregorie
112
and on that note I hope you enjoyed this tutorial.
113
Hope you're excited about the trials coming up in the section and I look for the next step.
114
And until then enjoy computer vision.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.