A web application to display intelligence evaluation scores for open-source AI models, similar to the structure on artificialanalysis.ai/models/open-source.
- ๐ Display model performance across multiple benchmarks (MMLU, AALCR, SciCode, ฯยฒ-Bench, Telecom, etc.)
- ๐ Search and filter models
- ๐ Color-coded scores (high/medium/low performance)
- ๐จ Modern, responsive UI
- ๐ Refresh data functionality
Test_AA/
โโโ index.html # Main web interface
โโโ styles.css # Styling
โโโ app.js # Frontend JavaScript
โโโ data_fetcher.py # Data fetching script (sample data)
โโโ api_client.py # API client for artificialanalysis.ai
โโโ web_scraper.py # Web scraper for the website
โโโ evaluations_data.json # Data file (generated)
โโโ README.md # This file
-
Install dependencies:
pip3 install requests beautifulsoup4
-
Fetch data:
# Try API first (recommended) python3 api_client.py # Or try web scraping (for JavaScript-rendered pages) python3 web_scraper.py # Or use Selenium scraper (requires ChromeDriver) pip install selenium python3 selenium_scraper.py # Or use sample data (fallback) python3 data_fetcher.py
-
Open the web interface:
-
Option 1: Use the included server (recommended):
python3 server.py
This will automatically open your browser
-
Option 2: Use Python's built-in server:
python3 -m http.server 8000
Then visit
http://localhost:8000 -
Option 3: Simply open
index.htmldirectly in your browser
-
Your API key is configured in the scripts:
api_client.py:aa_OXmwOTJvjVHpPnJQsOgimbFMwsPoVgOTdata_fetcher.py:aa_OXmwOTJvjVHpPnJQsOgimbFMwsPoVgOT
- MMLU - Massive Multitask Language Understanding
- AALCR - AALCR Benchmark
- SciCode - Scientific Code Understanding
- ฯยฒ-Bench - Tau Squared Benchmark
- Telecom - Telecom Benchmark
- HellaSwag - HellaSwag Benchmark
- ARC - AI2 Reasoning Challenge
- TruthfulQA - TruthfulQA Benchmark
- GSM8K - Grade School Math 8K
- Winogrande - Winogrande Benchmark
The evaluations_data.json file follows this structure:
{
"benchmarks": [
{
"id": "mmlu",
"name": "MMLU",
"full_name": "Massive Multitask Language Understanding"
}
],
"models": [
{
"id": "deepseek-v3.2",
"name": "DeepSeek V3.2",
"provider": "DeepSeek",
"scores": {
"mmlu": 0.85,
"aalcr": 0.69,
"scicode": 0.42
}
}
]
}-
View Evaluations:
- Open
index.htmlin your browser - The table displays all models with their scores
- Open
-
Search Models:
- Use the search box to filter models by name or provider
-
Refresh Data:
- Click the "Refresh Data" button to reload from
evaluations_data.json
- Click the "Refresh Data" button to reload from
To get actual scores from artificialanalysis.ai:
-
API Method:
- Update
api_client.pywith correct API endpoints if they differ - Run
python3 api_client.py
- Update
-
Scraping Method:
- If the website uses JavaScript rendering, you may need Selenium/Playwright
- Run
python3 web_scraper.py
-
Manual Method:
- Manually update
evaluations_data.jsonwith real scores - The web interface will automatically display the updated data
- Manually update
- Add more benchmarks: Edit the benchmarks list in
data_fetcher.pyorevaluations_data.json - Add more models: Add entries to the models array in
evaluations_data.json - Styling: Modify
styles.cssto change colors, fonts, or layout - Functionality: Extend
app.jsfor additional features
- The API endpoints may require authentication or have different URLs
- If the website structure changes, update
web_scraper.pyaccordingly - Scores are displayed as percentages (0.85 = 85%)
- Color coding: Green (โฅ80%), Yellow (60-79%), Red (<60%)
This project is for educational and evaluation purposes.