Hello team!
I was wondering how can we go about deploying the deepmoji model on mobile. The optimised size is around 22 MB. For deployment purpose on client side we need model size about 3-4MB. Ant tips on how can we go about it or how can we go about compressing the size of the model ?
Thanks in advance!
Good question. What have you done already?
The embedding layer can be reduced massively while retaining performance so I’d suggest you start there. There’s various papers on this :)
Hello
I started off with quantization of model,basically two techniques post training - pruning and quantisation. The quantisation as suggested by https://pytorch.org/tutorials/advanced/dynamic_quantization_tutorial.html, but hardly any change was observed. Next I thought of considering https://github.com/NervanaSystems/distiller/blob/master/examples/word_language_model/quantize_lstm.ipynb.
If it's not too much trouble, can you share some links which I can refer to ?