Scraping news sources
- Google News
- The Hindu
- Hindustan Times
- Deccan Chronicle
- Telegraph
- Times of India
- News18
- TimesNow
- NDTV
- CNN-IBM
- Yahoo News
- AajTak
- Mid Day
- Rediff
- National Herald
- News Today
- DD News
- Indian Express
- PTI
- Scraping(Beautiful Soup)
- News APIs(requests)
- Selenium
Initial plan is to look into various feasible ways to extract the headlines from news websites.
pipenv install
OR
pipenv --three
pipenv shell
pip install -r requirements.txt
pipenv shell
cp .env.example .env
EDITOR=nano $EDITOR .env
If you have setup mongodb then you can run the following scripts. Running them will store the scraped content to the database for further analysis.
python3 ./update_db_for_archived_news.py
python3 ./update_db_for_trending_news.py
If you don't have mongodb then just use the following command:
python3 ./scraper/dd-news.py
Here dd-news.py can be replaced with any other news scraper too. This will just give the scraped output and won't store it.
Make sure the name of the python source file is in lowercase and doesn't contain punctuation characters. If the news source name is News Source then the corresponding filename should be news-source.py and should be kept in scraper/ directory.
Reusing the-hindu.py is the best option to start writing a new parser for a news source.
Also add details of the source in sources.py.
You will have to keep a get_headlines(url) function in the python module, else running update_db_for_*_news.py will throw error. For consistency you can keep get_headline_details and get_all_content also which are used to find headline details and news content respectively.