sumegha19
I am trying to build similar model for Hindi(Devnagri). The pipeline in Deepmoji from building vocab to tokenizing, will it be the same or can we go for use of pre-trained word embeddings as provided by Fasttext and elmo to be used here. I am just curious as why building our own vocab appealed as a better way to target the problem than using word embeddings. I had this notion that once we have converted the words to their respective numeric representation, the model will work accordingly. But your team has specifically gone for buildling dataset specific words, then tokenizing them and finally training. What does the model do when an Out-of-vocabulary(OOV) word comes in a test case? Sorry for so many doubts. I am just beginning to understand emotion analysis. Thanks!!