Upload a CSV, pick which columns matter, and search it by meaning rather than keywords. Embeddings run in your browser (ONNX Runtime Web) once a small model is downloaded; until then, an identical search runs server-side as a fallback. Everything works inside a private network with no internet access — the model is served by this app, not a CDN.
[Upload CSV] → [Choose columns] → [Prepare embeddings] → [Search] → [Browse / export]
Two layers, one repo:
semantic_diff/
├── quick_embed.py # SemanticEngine: tokenizer + ONNX session (offline, reused as-is)
├── models/ # model_quantized.onnx (+ .onnx_data), tokenizer.json
├── VERSION # plain-text cache-bust key, e.g. "1.0.0"
├── requirements.txt # server + test deps
├── src/
│ ├── server/
│ │ ├── server.py # FastAPI: serves model files, /search fallback, static frontend
│ │ └── tests/ # pytest suite (health, models, search)
│ └── frontend/
│ ├── index.html # app shell + styles
│ ├── app.js # UI orchestration + state machine
│ ├── model.js # Web Worker owner: model state, init, embedTexts
│ ├── embed.worker.js # tokenizer + ONNX inference, off the main thread
│ ├── cache.js # OPFS / IndexedDB storage, model download + version check
│ ├── csv.js # CSV parse, column typing, row-text + hashing
│ ├── search.js # BM25 pre-filter + cosine re-rank, client/server dispatch
│ ├── setup_vendor.mjs# copies browser libs into vendor/
│ └── vendor/ # self-hosted ORT Web WASM + Papa Parse (no CDN)
└── PLAN.md # the full design + phased build plan
The browser caches the model in OPFS (or IndexedDB) and the per-CSV embedding matrix in IndexedDB, so repeat visits and repeat searches are instant.
- BM25 lexical pre-filter over all row texts → top N candidates (the
"Candidate pool size" /
pre_filter). - Embed the query (candidates are already embedded client-side; the server embeds them on the fly).
- Cosine re-rank (vectors are pre-normalized → dot product).
- Return top K.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
uvicorn src.server.server:app --host 0.0.0.0 --port 8000This serves the model files, the /search fallback, and the frontend at /.
Open http://localhost:8000.
The browser libraries are self-hosted (no CDN). After cloning, populate vendor/:
cd src/frontend
npm install
npm run setup-vendor # copies ORT Web WASM + Papa Parse into vendor/vendor/ can be committed so deployment needs no npm step.
source .venv/bin/activate
pytest tests/ -vSeveral browser APIs the on-device path relies on are only available in a secure
context — that means HTTPS, or http://localhost / http://127.0.0.1:
| API | Used for | If missing (plain-HTTP LAN IP) |
|---|---|---|
crypto.subtle |
cache keys (CSV / column hashes) | falls back to a non-crypto FNV-1a hash |
navigator.storage (OPFS) |
fast model file storage | falls back to IndexedDB |
navigator.storage.estimate() |
storage usage in Settings | shows "unavailable" |
SharedArrayBuffer / threads |
multi-threaded WASM inference | single-threaded (slower) |
The app works over plain HTTP (fallbacks kick in automatically), but for the best on-device experience, serve it over HTTPS or reach it via localhost. If on-device search won't initialize, that's the first thing to check.
The VERSION file is the single source of truth for cache invalidation. Clients
fetch /models/version on every load and re-download the model files when the
string changes.
Whenever you change anything in
models/, you MUST bumpVERSION. Otherwise browsers keep serving the old cached model and silently use stale embeddings.
echo "1.0.1" > VERSION # patch
echo "1.1.0" > VERSION # minor
# restart the serverA bumped version evicts the cached model in every browser on next load and triggers a fresh download.
Unauthenticated microservice. Base URL http://<host>:<port> (default port 8000).
| Endpoint | Purpose |
|---|---|
GET /health |
{ status, model_loaded, version } |
GET /models/version |
version string + per-file size & SHA-256 |
GET /models/{filename} |
serves model_quantized.onnx, model_quantized.onnx_data, tokenizer.json (else 404) |
POST /search |
BM25 + embedding re-rank fallback; see below |
{ "query": "who works in finance",
"chunks": [{ "id": 0, "text": "Alice, Head of Finance" }],
"top_k": 10, "pre_filter": 20 }| Field | Default | Constraints |
|---|---|---|
query |
required | — |
chunks |
required | 1–10 000 items, {id:int, text:str} |
top_k |
10 | 1–100 |
pre_filter |
20 | 1–500, auto-clamped up to top_k |
Response: { results: [{ id, score, bm25_rank }], mode: "server", pre_filter, query_ms }.
See PLAN.md for the complete design, phased build plan, and test matrix.