A lightweight, OpenAI-compatible embedding & rerank service. Embedding aligns with OpenAI /v1/embeddings, rerank aligns with Cohere /v1/rerank.
π δΈζ
| Feature | Description |
|---|---|
| π OpenAI Embedding | Drop-in replacement for OpenAI /v1/embeddings β works with the official Python SDK with zero code changes |
| π Cohere Rerank | /v1/rerank endpoint, Cohere-compatible. Dual-mode: cosine (lightweight) or cross-encoder (bge-reranker-base, 1GB) for high accuracy |
| π§ Dual Backend | BGE-M3 via FlagEmbedding (dense 1024-dim), general models via SentenceTransformer β auto-selects the optimal loader per model type |
| π Broad Model Support | 8+ pre-registered models (Embedding: e5-small/large, BGE series, MiniLM; Reranker: bge-reranker-base), plus auto-discovery for any HuggingFace model |
| π¦ Auto Download | Downloads models via CDN + aria2c with real-time progress bars. Cached locally, never re-downloads |
| π Hot Swap | Switch models on the fly by passing a different model parameter. Controlled via ALLOW_MODEL_SWITCH (recommended off in production) |
| π― Device Auto-detect | CUDA β MPS β CPU, with fp16 on GPU/MPS and fp32 on CPU |
| π€ Eager Loading | Preloads both Embedding + Rerank models at startup (PRELOAD_MODEL=true). Set to false for lazy loading on first request |
| π§΅ Thread Safe | threading.Lock guards model load/unload critical sections. PyTorch inference is natively thread-safe for concurrent requests |
| π Dimension Truncation | Supports OpenAI's dimensions parameter β truncate vectors to the first N dimensions for downstream storage needs |
| π³ Docker Ready | Multi-stage Docker build, built-in health check, compose one-liner, model cache bind-mounted to local dir |
| π Swagger Docs | Visit /docs after startup for full interactive API docs β all schemas include descriptions and examples |
On Apple Silicon (MPS), lightweight models deliver dramatically better throughput than heavy ones:
| Model | Size | Dims | Memory | Throughput | Suitable For |
|---|---|---|---|---|---|
e5-small |
120MB | 384 | ~0.1GB | β‘ Very High | εΌε / δ½θ΅ζΊη―ε’ |
all-minilm |
80MB | 384 | ~0.1GB | β‘ Very High | θ±ζεΊζ― |
bge-small-zh |
95MB | 512 | ~0.1GB | β‘ High | δΈζεΊζ― |
e5-large |
560MB | 1024 | ~0.5GB | Medium | ι«θ΄¨ιε€θ―θ¨ |
bge-large-zh |
390MB | 1024 | ~0.4GB | Medium | ι«θ΄¨ιδΈζ |
bge-m3 |
2GB+ | 1024 | ~2GB | π Slow / OOM | ιθ¦ dense+sparse+colbert ζΆ |
bge-reranker-base |
1GB | - | ~1GB | Medium | Reranker δΈη¨ cross-encoder |
π‘ Recommendation: Start with
e5-small(default) for embeddings,bge-reranker-basefor reranking. Both pre-loaded bypython preload.py.
The service supports two rerank modes β cosine (embedding-based, instant) and cross-encoder (dedicated reranker, high accuracy). The examples below demonstrate the critical difference on the same query.
π‘ Run
./compare_rerank.shto reproduce these results on your own machine.
Query: "How to lose weight effectively?"
| # | Document | Cosine (e5-small) | Cross-encoder (bge-reranker) |
|---|---|---|---|
| 1 | Exercise 30 minutes daily and eat less sugar | 0.8789 β | 5.4955 β |
| 2 | My friend tried everything but still failed to lose weight | 0.8429 | -4.4139 |
| 3 | The price of Bitcoin dropped 5% today | 0.7811 | -10.1932 |
π Cosine: all three scores cluster in a narrow range (0.78β0.88) β the model struggles to penalize the irrelevant Bitcoin doc because it shares surface-level features with the other texts. π― Cross-encoder: score spread is 15.69 points β clearly separates relevant (+5.50) from irrelevant (-10.19), making top-1 selection unambiguous.
Query: "θΉζε ¬εΈζζ°δΊ§εζ―δ»δΉοΌ" (What is Apple's latest product?)
| # | Document | Cosine (e5-small) | Cross-encoder (bge-reranker) |
|---|---|---|---|
| 1 | Apple announced the new iPhone 16 with AI features | 0.8455 | -0.7275 β |
| 2 | θΉζζ―δΈη§θ₯ε »δΈ°ε―ηζ°΄ζοΌε«ζδΈ°ε―ηη»΄ηη΄ C | 0.8720 β | -3.5546 |
| 3 | δ»ε€©ε€©ζ°εΎε₯½ιεεΊε»η© | 0.8202 | -10.1942 |
π Cosine ranks fruit above iPhone! The fruit doc (0.8720) outscores the correct answer (0.8455) β because cosine similarity treats all documents as bags of words and "θΉζ" (apple) appears heavily in both the query and the fruit doc. π― Cross-encoder correctly places the iPhone announcement first, understanding that "θΉζε ¬εΈ" (Apple Inc.) refers to the company, not the fruit.
| Dimension | Cosine (e5-small) | Cross-encoder (bge-reranker-base) |
|---|---|---|
| Speed | β‘ Instant (~10ms) | π’ Slower (~200ms per pair) |
| Memory | Free (reuses embedding model) | ~1GB extra |
| Score range | [-1, 1] (cosine similarity) | Unbounded (raw logit) |
| Discriminative power | Low β scores cluster tightly | High β wide score spread |
| Best for | Quick filtering, head queries | Precision reranking, tail queries |
| Ambiguity handling | Weak β "θΉζ" = fruit | Strong β "θΉζε ¬εΈ" = Apple Inc. |
π‘ Recommendation: Use cosine for fast pre-filtering of large candidate sets, then cross-encoder for precise top-N reranking. The
compare_rerank.shscript helps you evaluate which mode fits your use case.
| Method | Path | Description |
|---|---|---|
GET |
/ |
Service info: current model, dimension, backend type, load time |
GET |
/health |
Health check: {"status": "healthy"} |
GET |
/v1/models |
List models (OpenAI-compatible format) |
POST |
/v1/embeddings |
Generate embeddings (fully OpenAI-compatible) |
POST |
/v1/rerank |
Rerank documents by query relevance (Cohere-compatible) |
GET |
/docs |
Swagger UI |
GET |
/redoc |
ReDoc |
GET |
/openapi.json |
OpenAPI spec |
{
"model": "bge-m3",
"input": "The quick brown fox jumps over the lazy dog",
"encoding_format": "float",
"dimensions": null
}{
"object": "list",
"model": "bge-m3",
"data": [
{
"object": "embedding",
"index": 0,
"embedding": [0.0123, -0.0456, 0.0789, "..."]
}
],
"usage": {
"prompt_tokens": 10,
"total_tokens": 10
}
}Dual-mode: cosine (lightweight, uses embedding model) or cross-encoder (professional reranker).
// Request: POST /v1/rerank
{
"model": "bge-reranker-base",
"query": "What is machine learning?",
"documents": [
"Deep learning is a subset of ML",
"The weather is nice today",
"Machine learning is a branch of AI"
]
}
// Response
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"model": "bge-reranker-base",
"results": [
{"index": 2, "relevance_score": 0.98, "document": "Machine learning is a branch of AI"},
{"index": 0, "relevance_score": 0.82, "document": "Deep learning is a subset of ML"},
{"index": 1, "relevance_score": 0.12, "document": "The weather is nice today"}
],
"meta": {"api_version": "1.0"}
}
β οΈ Run preload first! Downloads both embedding and reranker models (~1.2GB total). Skip this and the first API call will block until the download completes.
# 1. Download models (required β do this before starting!)
cd lite-emb
uv sync
python preload.py# 2. Start the server
uv run python main.pyOptional: customize models or port in .env:
cp .env.example .env
# Edit .env to change MODEL_NAME, DEV_PORT, etc.The service runs at http://localhost:8000. Visit /docs for interactive API docs.
# Health check
curl http://localhost:8000/health
# Service info (current model, dimension, etc.)
curl http://localhost:8000/
# Model list
curl http://localhost:8000/v1/models
# Single text embedding
curl -X POST http://localhost:8000/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "bge-m3", "input": "Hello, world!"}'
# Batch embedding
curl -X POST http://localhost:8000/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "bge-m3", "input": ["text one", "text two", "text three"]}'
# Rerank (cosine mode β instant)
curl -X POST http://localhost:8000/v1/rerank \
-H "Content-Type: application/json" \
-d '{"model":"e5-small","query":"What is ML","documents":["Deep learning","Nice weather","AI branch"]}'
# Rerank (cross-encoder mode β higher accuracy)
curl -X POST http://localhost:8000/v1/rerank \
-H "Content-Type: application/json" \
-d '{"model":"bge-reranker-base","query":"What is ML","documents":["Deep learning","Nice weather","AI branch"]}'from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed" # lite-emb requires no API key
)
response = client.embeddings.create(
model="bge-m3",
input="Your text here"
)
print(response.data[0].embedding)| Alias | HuggingFace ID | Dims | Max Length | Backend | Languages |
|---|---|---|---|---|---|
bge-m3 |
BAAI/bge-m3 |
1024 | 8192 | BGEM3FlagModel | 100+ |
bge-large-zh |
BAAI/bge-large-zh-v1.5 |
1024 | 512 | SentenceTransformer | Chinese |
bge-small-zh |
BAAI/bge-small-zh-v1.5 |
512 | 512 | SentenceTransformer | Chinese |
e5-large |
intfloat/multilingual-e5-large |
1024 | 512 | SentenceTransformer | 100+ |
all-minilm |
sentence-transformers/all-MiniLM-L6-v2 |
384 | 256 | SentenceTransformer | English |
all-minilm-l12 |
sentence-transformers/all-MiniLM-L12-v2 |
384 | 256 | SentenceTransformer | English |
bge-micro |
BAAI/bge-micro |
384 | 512 | SentenceTransformer | 100+ (~17MB) |
e5-large |
intfloat/multilingual-e5-large |
1024 | 512 | SentenceTransformer | 100+ |
Any HuggingFace sentence-transformers compatible model works out of the box β no pre-registration needed. The system will:
- Check if it is a known model β use the registered backend
- Model name contains
bge-m3β use BGEM3FlagModel backend - Everything else β use SentenceTransformer backend
# Use any HuggingFace model directly
curl -X POST http://localhost:8000/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "e5-base", "input": "hello"}'# Download the default model (BAAI/bge-m3)
uv run python preload.py
# Download a specific model
uv run python preload.py --model sentence-transformers/all-MiniLM-L6-v2
# Start in offline mode after pre-downloading
HF_HUB_OFFLINE=1 uv run python main.pyRequest arrives at POST /v1/embeddings
ββ Model already loaded? β Return cached instance
ββ Different model requested?
β ββ ALLOW_MODEL_SWITCH=true β Unload old model β Load new model
β ββ ALLOW_MODEL_SWITCH=false β Return 400 error
ββ First load?
ββ Check local cache
ββ Not cached β Download from HuggingFace (via hf-mirror)
ββ Load via BGEM3FlagModel / SentenceTransformer into memory
| Mode | Command | Use Case |
|---|---|---|
| Development | uv run python main.py |
Local dev, single process |
| Hot Reload | uv run uvicorn lite_emb.app:app --reload |
Auto-reload on code changes |
| Multi-worker | WORKERS=4 uv run python main.py |
Multi-core concurrency |
| Lazy Loading | PRELOAD_MODEL=false uv run python main.py |
Defer model loading to first request |
| Offline | HF_HUB_OFFLINE=1 uv run python main.py |
No network (preload first) |
| Custom Model | MODEL_NAME=all-MiniLM-L6-v2 uv run python main.py |
Switch default model |
| Docker | docker compose up -d |
Containerized deployment |
| Docker (GPU) | Uncomment GPU section in docker-compose.yml, then docker compose up -d |
NVIDIA GPU acceleration |
# Build and start
docker compose build --build-arg USE_APT_MIRROR=true --build-arg USE_PYPI_MIRROR=true
docker compose up -d
# View logs
docker compose logs -f
# Stop
docker compose downπ China users: Configure Docker Desktop β Settings β Resources β Proxies first, otherwise pulling
python:3.11-slim-bookwormbase image will timeout. The proxy only affects the Docker daemon pulling base images β build-stage apt/pip/model downloads still go direct to domestic mirrors.
All settings are managed via .env file or environment variables:
| Variable | Default | Description |
|---|---|---|
MODEL_NAME |
e5-small |
Embedding model (alias or HuggingFace ID) |
RERANK_MODEL_NAME |
bge-reranker-base |
Rerank model (alias or HuggingFace ID) |
DEV_HOST |
0.0.0.0 |
Listen address |
DEV_PORT |
8000 |
Listen port |
WORKERS |
1 |
Uvicorn workers (set to CPU count in production) |
MAX_BATCH_SIZE |
32 |
Max batch size per encoding call |
MAX_SEQUENCE_LENGTH |
512 |
Max input tokens |
PRELOAD_MODEL |
true |
Preload both models at startup |
ALLOW_MODEL_SWITCH |
false |
Allow dynamic model switching per request |
EMBEDDING_NORMALIZE |
true |
L2-normalize output vectors |
HF_ENDPOINT |
https://hf-mirror.com |
HuggingFace mirror endpoint |
HF_HUB_ENABLE_HF_TRANSFER |
true |
Enable hf-transfer for faster downloads |
HF_HUB_OFFLINE |
0 |
Offline mode (1 = local cache only) |
LOG_LEVEL |
INFO |
Log level (DEBUG/INFO/WARNING/ERROR) |
LOG_FILE |
logs/app.log |
Log file path (auto rotation + compression) |
| Component | Technology | Notes |
|---|---|---|
| Web Framework | FastAPI | Async, auto-generated OpenAPI docs |
| Server | Uvicorn | ASGI server, multi-worker support |
| Configuration | pydantic-settings | Type-safe, .env loading |
| Logging | loguru | Structured logging, file rotation & compression |
| BGE-M3 Backend | FlagEmbedding | Native dense/sparse/colbert vector support |
| General Backend | sentence-transformers | Compatible with all HuggingFace embedding models |
| Model Download | huggingface-hub + hf-transfer | Parallel high-speed downloads |
| Deep Learning | PyTorch | CUDA/MPS/CPU adaptive |
| Package Manager | uv | Rust-based, fast resolution & installation |
| Testing | pytest + httpx | 24 tests covering all modules |
| Containerization | Docker multi-stage | Minimized image size |
MIT