Afrikaans
Akan
Albanian
Amharic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
So in this lecture we will discuss the squad data set, which is the dataset we'll be using for this
task.
Squad simply stands for Stanford.
Question Answering Data Set.
Now, as you already know, the task of question answering is still constrained in a few ways.
We can't yet give a neural network a big database of knowledge and just ask it any question we want.
Instead, we do what is called extractive.
Question Answering.
What this means is that we're going to give the network a pair of texts, namely the question and the
context which contains the answer to the question.
In that way, the answer is always a substring of the context.
Note also that because of this, the network never has to actually generate any text, so we don't require
an encoder decoder setup.
Now there are some details that will become important when you want to actually write the code.
Firstly, let's look at how we will load in the data.
As you can see, we just call the standard function load data set passing in the string squad.
The data set comes with five columns which are ID title, context, question and answers.
Note that the title is pretty much irrelevant.
Interestingly, ID is not irrelevant, which may seem strange since we've ignored it up until this point.
We'll discuss this more when the time comes.
So let's look at some examples.
Here's a context.
It says, Architecturally, the school has a Catholic character.
Atop the main building is Gold Dome, is a golden statue of the Virgin Mary, and so on and so forth.
The question is what is in front of the Notre Dame main building?
And the corresponding answer is a copper statue of Christ.
Note that the answer seems to have a funny format.
Firstly, the text is stored in a list.
We'll see why that makes sense shortly.
Furthermore, we see that in addition to the text, we also get the position of the start of the answer
in terms of characters.
As you recall, a string is simply an array of characters.
So if you think of the context as an array of characters, this would be the index of the start of the
answer.
So for this example, the corresponding title happens to be University of Notre Dame.
As you can see, this is irrelevant for finding the answer to the question.
What should strike you as interesting is that the answer column is plural.
This implies that there can potentially be multiple answers to the same question.
This also explains why the answer data is stored in lists.
Now, how can this be?
Well, consider the question where did Super Bowl 50 take place?
One possible answer is Santa Clara, California.
This is a true fact.
But another possible answer is Levi's Stadium.
This is also a true fact.
So how can one question have multiple answers?
Well, here's the context where this came from.
It says The game was played on February seven, 2016 at Levi's Stadium in the San Francisco Bay area
at Santa Clara, California.
Depending on how you interpret this question, both answers would be valid.
So this is an example of where the same question can have multiple valid answers.
Now, oddly, this data set is built such that for some questions, the exact same answer can appear
multiple times.
I'm not sure why that is.
Finally, note that this only happens for the validation set for the train set.
Although the column is called Answers.
There is only one answer per sample.
Since our neural networks loss function is only built for one target per input, this is a good thing
since it means we don't have to do any extra work to split up multiple answers into separate training
samples.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.