lvaleriu/language-models

pre-trained Language Models

★ 0Forks 0GitHub ↗Compare

README

Language Models

Repository of pre-trained Language Models.

WARNING: a Bidirectional LM model using the MultiFiT configuration is a good model to perform text classification but with only 46 millions of parameters, it is far from being a LM that can compete with GPT-2 or BERT in NLP tasks like text generation. This my next step ;-)

Note: the training times given below are the sum of fastai Databunch creation time + model training time on 10 epochs.

Portuguese

I trained 1 Portuguese Bidirectional Language Model with the MultiFit configuration.

MultiFiT configuration (architecture 4 QRNN with 1550 hidden parameters by layer / tokenizer SentencePiece (15 000 tokens))

accuracy perplexity training time
forward 39.68% 21.76 8h
backward 43.67% 22.16 8h

French

I trained 3 French Bidirectional Language Models but the best is the one trained with the MultiFit configuration.

1. MultiFiT configuration (architecture 4 QRNN with 1550 hidden parameters by layer / tokenizer SentencePiece (15 000 tokens))

accuracy perplexity training time
forward 43.77% 16.09 8h40
backward 49.29% 16.58 8h10

2. Architecture QRNN / tokenizer SentencePiece

accuracy perplexity training time
forward 40.99% 19.96 5h30
backward 47.19% 19.47 5h30

3. Architecture AWD-LSTM / tokenizer spaCy

accuracy perplexity training time
forward 36.44% 25.62 11h
backward 42.65% 27.09 11h

Contributors

piegu

Issues