Embedding model selection for RAG database - RSPEED-1798
Select an embedding model for use in our RAG (Retrieval-Augmented Generation) database that will be shipped with the system for document embedding and query processing.
- Install Prerequisites:
pip install -r requirements.txt - Download CSV: Visit https://huggingface.co/spaces/mteb/leaderboard
- Apply Filters: Set "Number of Parameters" to <100M
- Export Data: Click "Download CSV" button
- Save File: Place CSV in
/statsdirectory with date format:hugging_face_stats_YYYY_MM_DD.csv(script will auto-select most recent) - Run Analysis:
./run_full_analysis.sh
Rather than relying solely on overall MTEB rankings, we prioritized tasks most relevant to RAG:
Primary RAG Tasks:
- Retrieval: Finding relevant documents for queries (core RAG functionality)
- Semantic Textual Similarity (STS): Measuring text passage similarity
Secondary Tasks:
- Clustering, Instruction Retrieval (useful for document organization)
Excluded Tasks:
- Classification, Multilabel Classification, Pair Classification, Bitext Mining (not relevant to RAG use case)
- Source: HuggingFace MTEB leaderboard
- Filter: Models with <100M parameters
- Focus: RAG-optimized performance metrics (Retrieval + STS scores)
We initially attempted to use the official MTEB Python library to fetch live benchmark data:
python3 fetch_with_mteb_lib.py # Available for referenceProblems Discovered:
- Ranking Discrepancies: Models ranked differently in raw MTEB data vs. official HuggingFace leaderboard
- Evaluation Bias: Newer models evaluated on fewer tasks appeared to score higher due to selection bias
- Data Quality Issues: Raw MTEB data includes experimental/unvalidated evaluations not shown on official leaderboard
- Performance: Loading complete MTEB database is slow and resource-intensive
Example Issue: prdev/mini-gte showed high scores in raw MTEB data but appears with empty scores on the official leaderboard, indicating the raw data contains unvalidated results.
The HuggingFace MTEB leaderboard uses curated, validated evaluation results. For reliable model selection: