A web-based tool that allows users to input one or more URLs, ingest their content, and ask questions based on the information from those pages. The tool uses Gemini for question answering and has a user-friendly interface built with Streamlit (frontend) and FastAPI (backend).
final.mov
- URL Ingestion: Input one or more URLs to scrape and ingest their content.
- Question Answering: Ask questions based on the ingested content.
- Gemini Integration: Uses Gemini for accurate and context-aware question answering.
- Minimal UI: Clean and intuitive user interface for seamless interaction.
- Framework: Streamlit (Python)
- Styling: Streamlit components
- Hosting: Streamlit Cloud
- Framework: FastAPI (Python)
- Web Scraping:
requests+BeautifulSoup - Question Answering: Gemini API
- Hosting: Render
- Frontend: https://web-content-app-python-aadi71.streamlit.app/
- Backend: https://web-content-qa-python.onrender.com
Follow these steps to set up and run the project on your local machine.
- Python: Ensure you have Python installed (v3.8 or higher).
- pip: Python package manager.
- Gemini API Key: Get your API key from Google AI Studio.
Clone the repository to your local machine:
git clone https://github.com/Aadi71/web-content-qa-tool-python.git
cd web-content-qa-tool-python-
Navigate to the
backenddirectory:cd backend -
Install dependencies:
pip install -r requirements.txt
-
Create a
.envfile in thebackenddirectory and add your Gemini API key:GEMINI_API_KEY=your-gemini-api-key
-
Start the backend server:
uvicorn main:app --host 0.0.0.0 --port 8000
The backend will run on
http://localhost:8000.
-
Navigate to the
frontenddirectory:cd ../frontend -
Install dependencies:
pip install -r requirements.txt
-
Create a
.envfile in thefrontenddirectory and add the backend URL:BACKEND_URL=http://localhost:8000
-
Start the frontend development server:
streamlit run app.py
The frontend will run on
http://localhost:8501.
Open your browser and navigate to http://localhost:8501 to use the Web Content Q&A Tool.
web-content-qa-tool-python/
├── backend/ # FastAPI backend
│ ├── main.py # FastAPI application
│ ├── scraper.py # Web scraping logic
│ ├── gemini.py # Gemini integration
│ ├── requirements.txt # Backend dependencies
│ └── .env # Environment variables
├── frontend/ # Streamlit frontend
│ ├── app.py # Streamlit application
│ ├── requirements.txt # Frontend dependencies
│ └── .env # Environment variables
├── README.md # Project documentation
└── .gitignore # Git ignore file
GEMINI_API_KEY: Your Gemini API key.
BACKEND_URL: The URL of the backend server (e.g.,http://localhost:8000).
-
URL Ingestion:
- The user inputs one or more URLs.
- The backend scrapes the content of the URLs using
requestsandBeautifulSoup.
-
Question Answering:
- The user asks a question based on the ingested content.
- The backend sends the question and context to the Gemini API.
- Gemini generates an answer based on the provided context.
-
Response Display:
- The frontend displays the answer to the user.
fastapi: Backend framework.uvicorn: ASGI server to run FastAPI.requests: HTTP client for web scraping.beautifulsoup4: HTML parsing library.google-generativeai: Gemini API client.python-dotenv: Environment variable management.
streamlit: Frontend framework.requests: HTTP client for API calls.