Afrikaans
Akan
Albanian
Amharic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
Alright. Before crunching any numbers and making decisions,
we should introduce some key definitions.
The first step of every statistical analysis you will perform is to determine whether the data you are
dealing with is a population, or a sample. A population is the collection of all items of interest
to our study and is usually denoted with an uppercase N. The numbers we've obtained when using a population
The numbers we've obtained when using a population are called parameters.
A sample is a subset of the population and is denoted with a lowercase n, and the numbers we've obtained
when working with a sample are called: statistics.
Now you know why the field we are studying is called statistics.
Let's say we want to make a survey of the job prospects of the students studying in the New York University.
What is the population? You can simply walk into NYU and find every student, right?
Well probably that would not be the population of NYU students.
The population of interest includes not only the students on campus, but also the ones at home, on exchange,
abroad,
distance education students, part time students, even the ones who are enrolled but are still at high school.
Though exhaustive, even this list misses someone. Point taken.
Populations are hard to define and hard to observe in real-life.
A sample, however, is much easier to contact.
It is less time consuming and less costly. Time and resources are the main reasons we prefer drawing samples
compared to analyzing an entire population.
So, let's draw a sample then. As we first wanted to do, we can just go to the NYU campus.
Next let's enter the canteen because we know it will be full of people.
We can then interview 50 of them.
Cool.
This is a sample.
Good job! But what are the chances of these 50 people provide us answers that are a true representation
of the whole university?
Pretty slim, right?
The sample is neither random, nor representative. A random sample is collected when each member of the
sample is chosen from the population strictly by chance.
We must ensure each member is equally likely to be chosen.
Let's go back to our example.
We walked into the university canteen and violated both conditions.
People were not chosen by chance.
They were a group of NYU students who were there for lunch.
Most members did not even get the chance to be chosen as they were not on campus.
Thus we conclude the sample was not random.
What about the representativeness of the sample?
A representative sample is a subset of the population that accurately reflects the members of the entire population.
A representative sample is a subset of the population that accurately reflects the members of the entire population.
Our sample was not random but was it representative?
Well, it represented a group of people but definitely not all students in the university to be exact.
It represented the people who have lunch at the university canteen.
Had our survey been about job prospects of NYU students who eat in the university canteen we would have done well.
Had our survey been about job prospects of NYU students who eat in the university canteen we would have done well.
By now, you must be wondering how to draw a sample that is both random and representative.
Well, the safest way would be to get access to the student database and contact individuals in a random manner.
Well, the safest way would be to get access to the student database and contact individuals in a random manner.
However, such surveys are almost impossible to conduct without assistance from the university.
We said populations are hard to define and observe.
Then we saw that sampling is difficult. But samples have two big advantages.
First, after you have experience, it is not hard to recognize, if a sample is representative.
And second statistical tests are designed to work with incomplete data.
Thus, making a small mistake while sampling is not always a problem.
Don't worry.
After completing this course samples and populations will be a piece of cake for you!
Thanks for watching!
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.