All language subtitles for 007 Term Frequency Inverse Document Frequency (TFIDF)_en[UdemyIran.Com]

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranรฎ)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

Now, of course, indices aren't quite that simple.

An index is actually what's called an inverted index and this is basically the mechanism by which pretty

much all search engines work.

As an example, imagine I have a couple of documents in my index that contain text to data.

Let's say I have one document that contains: Space the final frontier,

these are the voyages, and maybe I have another document that says: he's bad,

he's number one, he's a space cowboy with a laser gun,

and if you understand what both of those are references to, then you and I have a lot in common. Now an

inverted index wouldn't store those strings directly,

instead, it sort of flips it on its head. A search engine, such as the elastic search, actually splits each

document up into its individual search terms,

and in this example, we'll just split it up for each word and we'll convert them to lowercase just to

normalize things.

Then what it does is map each search term to the documents that those search terms occur within.

So in this example, the word space actually occurs in both documents, meaning the inverted index would

indicate that the word space occurs in both documents one and two, the word

the also appears in both documents,

so that will also map to both documents one and two, and the word, final, only appears in the first document,

so the inverted index would match the word, final, as a search term to document one.

Now it's a little bit more complicated than that in practice and in reality it actually stores not only

what documents end but also the position within the document that it's in.

But at a high conceptual level, this is the basic idea. An inverted index is what you're actually getting

with a search index, where it's mapping things that you're searching for to the documents of those things

live within, and of course it's not even quite that simple.

So how do I actually deal with the concept of relevance?

Let's take - for example - the word the,

how do I deal with that?

The word the is going to be a very common word in every single document.,

so how do I make sure that only documents where the is a special word are the ones that I get back,

if I actually search for the word the? Well that's where TF IDF comes in, that stands for a term frequency

times inverse document frequency, it's a very fancy-sounding term but it's actually a very simple concept.

So let's break it down. Term frequency is just how often a given search term appears within a given document.

So if the word space occurs very frequently in a given document, it would have a high term frequency.

The same applies if the word appears frequently to the document,

it would also have a high term frequency. Now document frequency is just how often a term appears in

all of the documents in your entire index.

So here's where things get interesting.

So the word space probably doesn't occur very often across the entire index, so it would have a low document

frequency.

However, the word does appear in all documents pretty frequently,

so it would have a very high document frequency. So if we divide term frequency by document frequency,

that's the same as multiplying by the inverse document frequency,

mathematically we get a measure of relevance.

So we see how special this term is to the document.

It measures not only how often does this term occur within the document, but how does that compare to

how often this term occurs in documents across the entire index?

So with that example, the word, space, in an article about space would rank very highly.

However, the word the wouldn't necessarily rank very highly at all,

that's a common term found in every other document as well,

and this is the basic idea of how search engines work. If you're searching for a given term, it will try

to give you back results in the order of their relevancy. Relevancy is loosely based at least on the

concept of TF-IDF,

it's not really that complicated.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.