This project is a web scraper written in Go that fetches JSON data from multiple URLs concurrently and saves the data to Cassandra.
- Fetch JSON data from multiple URLs concurrently.
- Use goroutines and channels for concurrency.
- Use WaitGroup to wait for all goroutines to finish.
- Save the fetched data to Cassandra.
- Retry Mechanism
- Kafka integration for producing and consuming messages.
- Go 1.23.3 or higher
- Docker and Docker Compose
github.com/joho/godotenvpackagegithub.com/confluentinc/confluent-kafka-go/v2package
-
Install dependencies:
go mod tidy
-
Create a
.envfile in the root directory with the following content:BASE_URL=https://jsonplaceholder.typicode.com/posts NUM_PAGES=10 TIMEOUT=10000 WORKERS=5 RETRY_ATTEMPTS=3 KAFKA_BROKER=kafka:9092 KAFKA_TOPIC=posts MAX_MESSAGES=100 MAX_DURATION=60000 CASSANDRA_KEYSPACE=web_scraper
-
Build and run the Docker containers:
docker-compose up --build
-
The fetched data will be saved to Cassandra.
To view the scraped posts stored in Cassandra, follow these steps:
-
Open a new terminal and connect to the Cassandra container:
docker exec -it <cassandra_container_name> cqlsh
Replace
<cassandra_container_name>with the actual name of your Cassandra container. The default iscassandra. -
Use the
web_scraperkeyspace:USE web_scraper; -
Query the
poststable to view the scraped posts:SELECT * FROM posts;
This will display the data in the posts table, allowing you to verify that the posts have been inserted correctly.
This project uses Kafka for message brokering. The docker-compose.yml file includes a Kafka service setup. Ensure you have Docker and Docker Compose installed.
The docker-compose.yml file should include the following services:
zookeeper: Required by Kafka for broker coordination, fault tolerance, configuration management and distributed synchronization.kafka: The Kafka broker.web-scraper: The Go application.cassandra: The NoSQL storage system.
- Load environment variables from the
.envfile. - Fetch JSON data from URLs concurrently using goroutines and channels.
- Produce and consume messages to/from Kafka.
- Save the fetched data to Cassandra.
- Fetch JSON data from a given URL.
- Decode the JSON data into a slice of
Poststructs. - Send the data to a channel.
- Represents a post with
UserID,ID,Title, andBodyfields.
Refer to learning-notes.md for a cheatsheet on concurrency in Go, including usage of goroutines, channels, buffered channels, select statements, WaitGroup, and Mutex.
This project is licensed under the MIT License.