pip install pandas transformers sklearn nltk
Start the Jupyter notebook server
jupyter notebook
There are two files for initial mapping of Charts and Stocks datasets
Code/Para_Extracter.py- Extracting paragraphs from news datasetCode/ParaPrice_Mapper.py- Map stock price chart with extracted paragraphs
Run these two files in order to create the mapped dataset, which will be used in all further notebooks
There are three notebooks:
Code/Twitter Stock volume extraction.ipynb- Extract recent tweets for a particular stock symbol.Code/Twitter Breakout.ipynb- Detecting breakout hours from extracted tweetsCode/Web Content Extraction.ipynb- Extract news URL contents scraped from tweets
Code/Stock_Models.ipynb- Contains code for training the following algorithms: XGBoost, Decision tree, SVC, Naive Bayes, Random forest, Logistic regressionCode/gpt2classifier_recent_news.ipynb- GPT-2 classifierCode/lstmcnnGlove.ipynb- LSTM-CNN with GloVe Embedding classifier
Code/Model earnings.ipynb- Logic for evaluating ROIs for any given prediction output
- https://machinelearningmastery.com/sequence-classification-lstm-recurrent-neural-networks-python-keras/
- https://stackabuse.com/python-for-nlp-word-embeddings-for-deep-learning-in-keras/
- https://datascience.stackexchange.com/questions/45165/how-to-get-accuracy-f1-precision-and-recall-for-a-keras-model
- https://towardsdatascience.com/text-classification-on-disaster-tweets-with-lstm-and-word-embedding-df35f039c1db