Based on MCP, this agent extracts policies and rules from PDF, HTML, and TXT documents and produce structured risk categories governed by corresponding rule definitions.
deep policy explorationfeature: Designed for automatically exploring comprehensive policies among lengthy PDF documents or distributed HTML websites via a priority-based search algorithm that explores document subsections in-depth.
- ๐ Supports multiple document types (PDF, HTML, TXT)
- ๐ Deep exploration of document sections and links via tree-search
- โ๏ธ Prioritizes sections based on likelihood of containing target policies
- ๐ Parallel processing support for faster extraction
- ๐ Comprehensive output including document trees and visualizations
- ๐ท๏ธ Automatic rules extraction and risk categorization
# Clone the repository
git clone https://github.com/BillChan226/ShieldAgent.git
cd ShieldAgent
# Install required dependencies
pip install -r requirements.txtpython policy_extractor_async.py [OPTIONS]| Argument | Short | Description |
|---|---|---|
--document-path |
-d |
Path to document file (PDF, TXT) or URL (HTML) |
--organization |
-org |
Name of the organization whose policies are being extracted |
--input-type |
-t |
pdf |
| Argument | Short | Default | Description |
|---|---|---|---|
--organization-description |
-org-desc |
"" |
Description of the organization |
--target-subject |
-ts |
User |
Target subject of the policies (e.g., "User", "Customer") |
--initial-page-range |
-ipr |
1-5 |
Initial page range to extract from PDF |
--deep-policy |
-dp |
False |
Flag to automatically explore policy documents in-depth |
--output-dir |
-o |
./output/deep_policy |
Directory to save output files |
--debug |
False |
Enable debug mode | |
--user-request |
-u |
"" |
Specific user request to guide the policy extraction |
--async-num |
-a |
1 |
Number of policy extraction tasks to run in parallel (1-3) |
--extract-rules |
-er |
False |
Extract concrete rules from policies after extraction |
The tool supports three input types:
- Supports page ranges with
-ipr(e.g., "1-5" for pages 1 through 5) - Automatically extracts text from specified pages
- Useful for official policy documents, terms of service, etc.
- Extracts content from web pages
- Can follow links for deep policy extraction when
-dpis specified - Ideal for online policy repositories and terms of service pages
- Directly processes plain text files
- Simple and fastest processing option
- Good for pre-extracted content or manually curated policy text
This example extracts policies from Reddit's rules page, follows links to related pages, and extracts rules targeting users:
python policy_extractor_async.py \
-d https://redditinc.com/policies/reddit-rules \
-t html \
-org "Reddit" \
-u "Focus on rules that target users and customers" \
-dp \
-erThis example extracts policies from the EU AI Act PDF, focusing on pages 2-8:
python policy_extractor_async.py \
-d ./policy_docs/eu_ai_act_art5.pdf \
-t pdf \
-org "EU AI ACT" \
-ipr "2-8" \
-erThis example extracts policies from a text file containing EU AI Act rules:
python policy_extractor_async.py \
-d ./eu_ai_act_rules/art_5.txt \
-t txt \
-org "EU AI ACT" \
-erFor large documents, you can enable parallel processing of sections:
python policy_extractor_async.py \
-d large_policy_document.pdf \
-t pdf \
-org "Large Organization" \
-a 3 \
-erThe tool generates several output files in the specified output directory:
| File | Description |
|---|---|
{organization}_all_extracted_policies.json |
All extracted policies |
{organization}_all_extracted_rules.json |
Concrete rules extracted from policies |
{organization}_risk_categories.json |
Categorized rules with risk categories |
{organization}_document_tree.json |
Hierarchical representation of document sections with policies |
{organization}_extraction_report.json |
Detailed report of the extraction process |
{organization}_policy_rule_mapping.json |
Mapping between policies and rules |
Provide a detailed description of the organization to improve extraction quality:
python risk_extraction_doc/policy_extractor_async.py \
-d ./company_policies.pdf \
-org "Acme Corp" \
-org-desc "A multinational technology company specializing in AI services" \
-erFocus extraction on specific subjects within policies:
python policy_extractor_async.py \
-d platform_guidelines.pdf \
-org "Social Platform" \
-ts "Content Creator" \
-u "Focus on monetization policies" \
-er- If you encounter errors with PDF processing, ensure you have all dependencies installed
- For HTML extraction issues, check if the website allows scraping
- For large documents, consider using page ranges (
-ipr) to process specific sections - Enable debug mode (
--debug) for detailed logging information
- Document Loading: The tool loads the document based on input type
- Content Extraction: Text is extracted from the document
- Section Analysis: The system analyzes sections for policy content
- Policy Extraction: Policies are identified and extracted
- Deep Exploration: If enabled, the system explores other related links or subsections
- Rules Extraction: Concrete rules are extracted from policies
- Risk Categorization: Rules are categorized by risk type
- Output Generation: Results are saved to JSON files
Policy structure:
{
"policy_id": "p123",
"organization": "Example Org",
"policy_text": "Users must not share personal information of others without consent.",
"policy_source": "Privacy Policy, Section 3.2",
"extraction_timestamp": "2023-10-25T14:30:00"
}Rule structure:
{
"rule_id": "r456",
"source_policy_ids": ["p123"],
"rule_text": "Do not share another user's personal information without their explicit consent",
"risk_category": "Privacy Violation",
"severity": "High"
}[Your license information here]