JayQuimby/semantic_diff

★ 0Forks 0JavaScriptGitHub ↗Compare

README

Semantic Document Intelligence

Upload a CSV, pick which columns matter, and search it by meaning rather than keywords. Embeddings run in your browser (ONNX Runtime Web) once a small model is downloaded; until then, an identical search runs server-side as a fallback. Everything works inside a private network with no internet access — the model is served by this app, not a CDN.

[Upload CSV] → [Choose columns] → [Prepare embeddings] → [Search] → [Browse / export]

Architecture

Two layers, one repo:

semantic_diff/
├── quick_embed.py          # SemanticEngine: tokenizer + ONNX session (offline, reused as-is)
├── models/                 # model_quantized.onnx (+ .onnx_data), tokenizer.json
├── VERSION                 # plain-text cache-bust key, e.g. "1.0.0"
├── requirements.txt        # server + test deps
├── src/
│   ├── server/
│   │   ├── server.py       # FastAPI: serves model files, /search fallback, static frontend
│   │   └── tests/          # pytest suite (health, models, search)
│   └── frontend/
│       ├── index.html      # app shell + styles
│       ├── app.js          # UI orchestration + state machine
│       ├── model.js        # Web Worker owner: model state, init, embedTexts
│       ├── embed.worker.js # tokenizer + ONNX inference, off the main thread
│       ├── cache.js        # OPFS / IndexedDB storage, model download + version check
│       ├── csv.js          # CSV parse, column typing, row-text + hashing
│       ├── search.js       # BM25 pre-filter + cosine re-rank, client/server dispatch
│       ├── setup_vendor.mjs# copies browser libs into vendor/
│       └── vendor/         # self-hosted ORT Web WASM + Papa Parse (no CDN)
└── PLAN.md                 # the full design + phased build plan

The browser caches the model in OPFS (or IndexedDB) and the per-CSV embedding matrix in IndexedDB, so repeat visits and repeat searches are instant.

Search algorithm (same in both modes)

  1. BM25 lexical pre-filter over all row texts → top N candidates (the "Candidate pool size" / pre_filter).
  2. Embed the query (candidates are already embedded client-side; the server embeds them on the fly).
  3. Cosine re-rank (vectors are pre-normalized → dot product).
  4. Return top K.

Running it

1. Python server

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
uvicorn src.server.server:app --host 0.0.0.0 --port 8000

This serves the model files, the /search fallback, and the frontend at /. Open http://localhost:8000.

2. Frontend vendor libraries (one-time)

The browser libraries are self-hosted (no CDN). After cloning, populate vendor/:

cd src/frontend
npm install
npm run setup-vendor   # copies ORT Web WASM + Papa Parse into vendor/

vendor/ can be committed so deployment needs no npm step.

3. Tests

source .venv/bin/activate
pytest tests/ -v

Secure context matters

Several browser APIs the on-device path relies on are only available in a secure context — that means HTTPS, or http://localhost / http://127.0.0.1:

API Used for If missing (plain-HTTP LAN IP)
crypto.subtle cache keys (CSV / column hashes) falls back to a non-crypto FNV-1a hash
navigator.storage (OPFS) fast model file storage falls back to IndexedDB
navigator.storage.estimate() storage usage in Settings shows "unavailable"
SharedArrayBuffer / threads multi-threaded WASM inference single-threaded (slower)

The app works over plain HTTP (fallbacks kick in automatically), but for the best on-device experience, serve it over HTTPS or reach it via localhost. If on-device search won't initialize, that's the first thing to check.


Updating the model

The VERSION file is the single source of truth for cache invalidation. Clients fetch /models/version on every load and re-download the model files when the string changes.

Whenever you change anything in models/, you MUST bump VERSION. Otherwise browsers keep serving the old cached model and silently use stale embeddings.

echo "1.0.1" > VERSION   # patch
echo "1.1.0" > VERSION   # minor
# restart the server

A bumped version evicts the cached model in every browser on next load and triggers a fresh download.


API

Unauthenticated microservice. Base URL http://<host>:<port> (default port 8000).

Endpoint Purpose
GET /health { status, model_loaded, version }
GET /models/version version string + per-file size & SHA-256
GET /models/{filename} serves model_quantized.onnx, model_quantized.onnx_data, tokenizer.json (else 404)
POST /search BM25 + embedding re-rank fallback; see below

POST /search

{ "query": "who works in finance",
  "chunks": [{ "id": 0, "text": "Alice, Head of Finance" }],
  "top_k": 10, "pre_filter": 20 }
Field Default Constraints
query required —
chunks required 1–10 000 items, {id:int, text:str}
top_k 10 1–100
pre_filter 20 1–500, auto-clamped up to top_k

Response: { results: [{ id, score, bm25_rank }], mode: "server", pre_filter, query_ms }.

See PLAN.md for the complete design, phased build plan, and test matrix.

Contributors

JayQuimby

Issues