Afrikaans
Akan
Albanian
Amharic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
So in this lecture, we'll be investigating something that came up earlier, which is why the fine tune
model outputs only generic label names.
As you recall in the previous lecture, we solved this in kind of a hacky way, which was to modify
the config file after calling the save method.
If you're a programmer, this might make you recoil in horror.
Luckily, there is a slightly better way.
Unfortunately, what you would hope existed doesn't actually exist.
Specifically, it would be nice if you could just pass in specific label names into the Franprix Pre-Trained
method, just as you can specify the number of labels.
This would be ideal and in my opinion makes the most sense.
But unfortunately, this currently isn't possible.
However, there is a way to achieve a similar effect, which is what we'll look at in this lecture in
particular.
Hugging face also has config objects.
We'll pass in this config object into the from pre-trained method.
So it pretty much works like our ideal scenario.
Note that these config objects are model specific like tokenisation.
So you can have a better config, a GP2 config and so forth.
As you might expect, there is also an auto config which automatically chooses the right config object
based on the checkpoint you give it.
Just as we can have auto tokenization and auto models.
We also have auto configs.
Now please note that most of this notebook is the same as the previous one.
So we'll skip to the relevant parts.
So we'll begin by importing auto config along with the auto model and the trainer class.
So recall that earlier in this notebook we've loaded in the data sets, converted them into the correct
format and so forth.
So the next step is to load up a config by calling from pre-trained passing in our checkpoint.
The next step is to print out our config just to see what it looks like.
So as you can see, it's sort of like a dictionary with keys and values.
Importantly, notice how there's nothing here about label names.
Now, if you check the attributes of the config object, you'll see that there are two relevant attributes
corresponding to labels, i.e. to label and label to ID.
So let's look at ID to label.
As you can see, this is a dictionary mapping an integer label ID to the corresponding label name.
As you may have expected.
The next step is to look at the label to ID.
So as you can see, we get the reverse mapping, which again, you may have expected.
So it should be evident that what we need to do is overwrite these 82 label in label to ID attributes.
Now the API for this isn't too great.
In my opinion, there should be a function for doing this so that you don't have to manually overwrite
attributes yourself.
For example, you can just pass in gibberish and it would break your config.
But since no such method exists, we'll just stick with what we can get.
So you can see here that we're basically assigning these the targeted map from earlier in this notebook.
As you recall, the targeted map had our desired label names mapped to corresponding integer IDs.
The next step is to call from Pre-Trained with their auto model to get back a model object.
The difference between what we did before and what we are doing now is that we are now passing in the
config object we just looked at.
Okay.
So essentially all of the remaining steps are the same as the previous notebook, so I won't bother
to explain them again.
Okay.
So at this point we fine tune our model, saved it and loaded it back in as a pipeline object.
At this point, we can just pass in some strings and get back predictions.
Okay.
So our first input is JetBlue.
Thank you.
Predictably, the prediction is positive and the label shows up as the string positive instead of something
like label one.
So passing in the config object was a success.
We no longer had to manually modify the config file.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.