All language subtitles for 13. TF Syntax Basics - Part One - Preparing the Data

af Afrikaans
ak Akan
sq Albanian
am Amharic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranรฎ)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

Welcome back everyone.

In this lecture we're going to begin learning about the very syntax basics for the caris API for tensor

flow.

Let's head over to a notebook and get started.

OK.

Here I am at the notebook.

I'm going to begin with a couple of imports will import pansies.

PD will also import pi as MP.

And since we're going to be doing just a little bit of visualization I will import seaborne as s an

s.

Now to actually focus on the syntax of cares for this particular lecture we're going to be using a very

simple and actually fake data set and in the next lecture we'll be using a realistic dataset and we'll

focus a lot more on feature engineering.

Let's go ahead and read in this file we'll use PD that read CSC and then underneath our data folder

there is a file called fake underscore our e.g. for regression that CSB and then we'll check out the

head of the state of frames so as data frame it's very very simple.

It simply has a price and then two corresponding features.

So we're going to treat this as a regression problem where based off feature 1 and feature 2 will attempt

to predict the price so we can imagine that maybe these are measurements of some rare gemstones where

the gemstone has feature 1 and in feature 2 and we're trying to predict the price.

So here we have historical information which means this is a supervised learning problem.

Our main goal is to build a model that when we pick a new Gemstone from the ground we can measure its

features feature 1 and feature two and predict what price we should be selling this at the market due

to the fact that we have historical information on the price sold based off these two features.

So a very simple dataset I want to quickly show you how we could explore this dataset.

We'll say create a pair plot of the data frame run that and then we'll be able to see the features versus

the price and you'll notice that especially feature 2 it seems to have a very high correlation with

the actual price.

So this kind of is another indicator that this is fake data.

So in a realistic dataset which we're going to do in the very next lecture we would take a lot of time

exploring this data doing what's known as exploratory data analysis performing a lot of visualizations

as well as possibly performing some feature engineering trying to extract other features from features

that we can't quite use.

However let's focus really on the main workflow for using Caris and tensor flow for deep learning.

So step number one is to read in your data and then once you have done feature engineering or data once

you've explored your data the next step is to create a test train split and we can do this from psychic

learns model selection has a train test split functionality which is really easy to use.

So what we're gonna do is we're gonna use this train test split function to split our data into a training

set and a test set.

So we'll train on the training set and then evaluate our model's performance on the test set.

So first what we want to do is we want to grab the features that we're going to use.

So in this case that will be feature 1 and feature 2 and because of the way tensor flow works we have

to actually pass in num pi arrays instead of Panda's data frames or pan the series so I can simply add

dot values to the end of a series or data frame and I'll return it back as a num pi array.

So what we're gonna do is we're gonna graph features and set that as X and by convention we use capital

X because typically the feature matrix is two dimensional to indicate that for capital and then the

label that we're going to predict our y is the price column and same thing here will grab values and

again by convention since the price is essentially a one dimensional vector we have lower case Y for

that.

So that's why we have upper case x and lowercase y it essentially stems from the way you would write

this down mathematically on paper.

So now that we have our actual num pi arrays.

So if we take a look at X it's just a num pi array of the same information we had in that data frame

for each one to feature two.

It's just num pints that a panda's it's time for our train test split.

So the way I like to do this is simply called train test split.

And after you've imported it you should be able to do shift tab to see the documentation string expand

on it go ahead and scroll down and eventually at the bottom you'll see something called examples and

to save myself a little bit of time.

I like to just copy and paste this line from the example which is essentially showing you how you would

actually use this.

So I'm going to paste that in and then put this all on one line and essentially explain what's going

on here.

Recall that when we do a train test split we both split our features into X train next test as well

as our labels into y train and white test.

Make sure you review the machine learning section of the course in case you have any questions on what

these four parameters or variables actually represent.

Then for train to split you pass and all your features as X your labels as y.

And then you choose a percentage as your test size.

So typically use maybe around 30 percent of your data.

So if I said zero point three that's going to be 30 percent of my total data will be used for the test

set and you can always make this smaller if you have really large data sets and then there's the random

state.

So the train test split is going to perform this split randomly so it's going to grab random rows and

then split them into the training side and the test side.

If you want to repeat the actual results of the split the same every time then you would set a random

state to a specific number.

The number itself is just an arbitrary arbitrary choice.

Have to make sure to choose the same one each time.

So go ahead and choose random state is equal to forty two and that way you will get the same random

split that I do.

So we'll go ahead and run this.

And now we've split up our data and we can actually check this by checking the shape of this.

So notice X train that shape is now seven hundred by two features and x test that shape is three hundred

by two features.

So here's 70 percent of our data as the train set and 30 percent as the test set.

Since our total size of the original data was 1000 rows.

OK now typically the next step is to actually normalize or scale your data because we're working with

weights and biases inside of a neural network.

If we have really large values in our feature set that could cause errors with the weights and later

on we'll talk about vanishing and exploding gradients that could be an issue.

But one way to try to avoid any issues when train your network is to normalize and scale your feature

data so psychic learn actually allows us to do this quite simply by saying from as K learned that pre

processing import and there's actually lots of different ways you can normalize or scale your data one

simple way is to use what's known as min max scaling.

So we'll go ahead and import min max scalar and if you call help on min max scalar it will actually

describe what this is doing.

So essentially it's going to transform your data based off the standard deviation of your data as well

as the men in the max values.

So you can see here the actual formulation that it's running for us.

So all we're gonna do is show you how you can scale your data which is very typical in your workflow

for dealing with neural networks.

Now you don't have to actually scale the label and if you take a look at our provided notebook we have

a link explaining why we don't need to scale the label.

We really only need to scale the features since that's essentially what's being passed through the actual

network.

The final label is just a comparison done at the end so to use a scalar with psychic learn all we do

is we first create an instance of it.

So we choose some variable name typically scalar and then we create an instance of our min max scalar

open close princes.

So now we have this instance of this scalar and what I'm going to do is I need to actually fit the scalar

onto my training data so I will say fit on X train and what it does is it simply calculates the parameters

it needs to perform the actual scaling later on.

So if we recall from calling help on min max scalar the min max scalar is dependent on the standard

deviation the minimum value and the maximum value within that particular dataset.

So what it does is it essentially calculates the stern deviation.

The men and Max.

So that's what it does to 1 fit on our training set.

And the reason we only run it on the training set is because we want to prevent what's known as data

leakage from the test set.

You don't want to assume that we have prior information of the test set.

So we only fit our scalar to the training set to not try to cheat and look into the test set.

Then what we need to do ifs transform our training data so we'll say extreme is now equal to scalar

that transform on X train that actually performs a transformation.

So there's essentially two steps here we fit which is calculate what's needed for the transformation

to occur and then we actually perform the transformation and we'll do the same for the test set.

So we'll say scalar transform on X test.

And now if we take a look at these values for x train you'll notice they have been scaled.

So if we take a look at what the max value on X train is it's now 1 and then the minimum value is now

0.

So everything's been scaled to now be between 0 and 1.

And again we're only fitting on the train set to not ascertain information from the test set because

that's essentially cheating.

Okay.

So now that we've scaled the data it's time to actually show you how to create your neural network.

So in part 2 of this lecture we'll begin creating our neural network.

I'll see you there.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.