An AI-powered interview preparation assistant that runs 100% offline on your local machine. Built with Foundry Local, SQLite, and JavaScript β no cloud, no API keys, no internet required.
Upload your CV/resume and a job description, and Interview Doctor generates tailored interview questions with coaching tips, all powered by a local LLM using Retrieval-Augmented Generation (RAG).
New to RAG? Retrieval-Augmented Generation is a pattern where an AI model's answers are grounded in your own documents. Instead of relying solely on what the model learnt during training, RAG retrieves relevant chunks from your uploaded CV and job description and feeds them to the model as context. This dramatically reduces hallucination and makes the questions specific to your experience.
flowchart TD
A[π Start] --> B[π Upload CV as PDF]
B --> C[πΌ Enter Job Title & Level]
C --> D{π Job Description?}
D -->|Yes| E[Chunk & Index in SQLite]
D -->|No| F[Skip]
E --> G[π§ RAG Retrieval + Foundry Local LLM]
F --> G
G --> H[π 5-7 Tailored Interview Questions]
H --> I[π¬ Interactive Follow-up Chat]
How a query flows:
- The user uploads a CV (PDF) and enters a job title/level
- The PDF text is extracted, chunked, and stored with TF-IDF vectors in SQLite
- When generating questions, the engine retrieves the most relevant CV chunks
- Those chunks are injected into the prompt as context for the local LLM
- Foundry Local generates a response using Phi-3.5 Mini, grounded in the retrieved context
- The response streams back to the user via SSE in the web UI
- 100% offline β no internet, no cloud, no API keys, no outbound calls
- RAG-powered β answers grounded in your actual CV and job description
- PDF support β upload your CV as a PDF; text is extracted automatically
- Streaming responses β real-time SSE streaming in the web UI
- Model loading progress β visual progress bar whilst the model initialises
- Document management β upload additional documents, view indexed docs
- Edge/compact mode β toggle for constrained devices with limited resources
- Interactive follow-up β ask follow-up questions after initial question generation
- Privacy-first β all data stays on your machine; nothing leaves the device
Before you begin, make sure you have:
- Node.js β₯ 20 β Download here
- Foundry Local β Microsoft's on-device AI runtime
# Windows winget install Microsoft.FoundryLocal # macOS brew install microsoft/foundrylocal/foundrylocal
- The phi-3.5-mini model (auto-downloaded on first run via the SDK, approximately 2 GB)
Tip: Run
foundry model listto check which models are already cached on your machine.
# 1. Clone the repository
git clone https://github.com/leestott/interview-doctor-js.git
cd interview-doctor-js
# 2. Install dependencies
npm install
# 3. Ingest the sample interview guide documents
npm run ingest
# 4. Start the web server
npm start
# Open http://127.0.0.1:3000 in a browsernpm run ingestreads every.mdfile indocs/, splits them into overlapping chunks, computes TF-IDF vectors, and stores everything indata/rag.db(SQLite).npm startlaunches Foundry Local, downloads and loads the Phi-3.5 Mini model (with progress shown in the UI), opens the vector store, and starts the Express server on port 3000.
- Open
http://127.0.0.1:3000in your browser - Wait for the model to finish loading (a progress bar is shown)
- Upload your CV (PDF, Markdown, or text file)
- Enter the job title and seniority level
- Optionally paste a job description for more targeted questions
- Click Generate Interview Questions
- Use the follow-up chat or quick-action buttons for deeper preparation
| Button | What it Does |
|---|---|
| π‘ Coaching Tips | Detailed tips for answering each generated question |
| πͺ My Strengths | Identifies your strongest talking points from your CV |
| π Gap Analysis | Highlights gaps between your CV and the job requirements |
| π Mock Interview | Generates a mock interview with sample answers |
| π§ Behavioural Qs | Focuses on behavioural/STAR-method questions |
You can also ingest a PDF directly:
npm run ingest -- path/to/your-cv.pdf| Method | Path | Description |
|---|---|---|
POST |
/api/chat |
Non-streaming chat completion |
POST |
/api/chat/stream |
Streaming chat via SSE |
POST |
/api/upload |
Upload a document (PDF/MD/TXT) to the knowledge base |
GET |
/api/docs |
List indexed documents |
GET |
/api/init-status |
Model initialisation status (SSE) |
GET |
/api/health |
Health check |
interview-doctor-js/
βββ docs/ # Interview guide documents (RAG knowledge base)
β βββ 01-interview-categories.md
β βββ 02-star-method.md
β βββ 03-technical-prep.md
βββ public/
β βββ index.html # Web UI (single-file, no build step)
βββ src/
β βββ chatEngine.js # Foundry Local + RAG orchestration
β βββ chunker.js # Document chunking + TF-IDF vector computation
β βββ config.js # App configuration (model, paths, chunk sizes)
β βββ ingest.js # Batch document ingestion script
β βββ pdfParser.js # PDF text extraction (offline)
β βββ prompts.js # System prompts (full + compact/edge)
β βββ server.js # Express server + API endpoints
β βββ vectorStore.js # SQLite-backed local vector store
βββ test/ # Unit tests (Node.js test runner)
β βββ chunker.test.js
β βββ config.test.js
β βββ prompts.test.js
β βββ server.test.js
β βββ vectorStore.test.js
βββ data/ # Generated at runtime
β βββ rag.db # SQLite vector database
βββ uploads/ # Uploaded CV files
βββ package.json
βββ AGENTS.md # AI agent instructions for this codebase
βββ CONTRIBUTING.md # Contributor guidelines
βββ blog_post.md # Developer tutorial blog post
βββ README.md
Reads .md files from docs/ and PDF files, parses optional YAML front-matter, then splits the content into overlapping chunks (default: approximately 200 tokens with 25-token overlap). Each chunk is stored with its TF-IDF vector in SQLite.
A lightweight vector store backed by SQLite (via sql.js, pure JavaScript with no native compilation needed). Stores document chunks alongside their TF-IDF vectors. At query time, it cosine-similarity-ranks all chunks against the query vector and returns the top-K results.
Orchestrates the full RAG flow:
- Converts the user's question into a TF-IDF vector
- Retrieves the top-K most relevant chunks from SQLite
- Builds a prompt with system instructions + retrieved context + user question
- Sends it to the local Phi-3.5 Mini model via the native
ChatClient - Streams the response back chunk-by-chunk
Extracts plain text from PDF files using pdf-parse, working entirely offline with no external API calls.
Two prompt variants:
- Full mode: detailed instructions for interview-focused, structured responses
- Edge mode: minimal prompt for constrained devices with limited context windows
npm testTests use the built-in Node.js test runner (no extra dependencies). They cover the chunker, vector store, config, prompts, and server API contract.
| Name | Command | Description |
|---|---|---|
| Ingest | npm run ingest |
Chunk and index all docs into SQLite |
| Ingest PDF | npm run ingest -- cv.pdf |
Ingest a specific PDF file |
| Start | npm start |
Start the web server (production) |
| Dev | npm run dev |
Start with auto-restart on file changes |
| Test | npm test |
Run unit tests |
Foundry Local is Microsoft's on-device AI runtime. It lets you run small language models (SLMs) like Phi-3.5 Mini directly on your laptop or workstation β no GPU required, no cloud dependency. The JavaScript SDK manages model discovery, download, and loading, then provides a native ChatClient for inference.
import { FoundryLocalManager } from "foundry-local-sdk";
const manager = FoundryLocalManager.create({ appName: "my-app" });
const model = await manager.catalog.getModel("phi-3.5-mini");
await model.load();
const chatClient = model.createChatClient();
const response = await chatClient.completeChat([
{ role: "user", content: "Hello!" },
]);TF-IDF (Term FrequencyβInverse Document Frequency) is a classic information retrieval technique. Each document chunk is converted into a numeric vector based on how important each word is within that chunk relative to all chunks. At query time, the user's question is vectorised the same way and compared against all stored vectors using cosine similarity.
This project uses TF-IDF instead of embedding models to keep everything lightweight and offline β no embedding API or large model needed for retrieval.
For small-to-medium document collections (a CV plus job descriptions), SQLite is fast enough for brute-force cosine similarity search and adds zero infrastructure. No need for Pinecone, Qdrant, or Chroma β just a single .db file on disc.
This project implements a lightweight, fully offline RAG pipeline using TF-IDF vectors and cosine similarity β no embedding models, no vector databases, no external services.
Document β Chunk β TF-IDF Vector β Store in SQLite
β
Query β TF-IDF Vector β Cosine Similarity Search β Top-K Chunks β LLM Prompt
Documents are split into overlapping chunks of approximately 200 whitespace-delimited tokens with a 25-token overlap between consecutive chunks. The overlap ensures important context is not lost at chunk boundaries. These values are configurable in src/config.js:
| Parameter | Default | Effect |
|---|---|---|
chunkSize |
200 tokens | Larger = more context per chunk, fewer chunks. Smaller = more precise retrieval. |
chunkOverlap |
25 tokens | Higher overlap = better boundary coverage, but more storage. |
topK |
5 | Number of chunks retrieved per query. Higher = more context, but may dilute relevance. |
Each chunk is converted into a term-frequency (TF) map β a dictionary of {word: count} after:
- Lowercasing all text
- Removing punctuation and non-alphanumeric characters
- Filtering out approximately 100 common English stopwords (the, is, at, which, etc.)
This produces a sparse vector that captures which meaningful terms appear in each chunk and how often. No IDF weighting is applied at index time β the cosine similarity comparison inherently accounts for term specificity.
At query time, the user's question is vectorised the same way and compared against every stored chunk using cosine similarity:
This measures the angle between two vectors, returning a score between 0 (no overlap) and 1 (identical term distribution). The top-K highest-scoring chunks are selected as context.
Retrieved chunks are formatted and injected into the system prompt alongside the user's question. The local LLM (Phi-3.5 Mini) then generates a response grounded in those specific document excerpts, dramatically reducing hallucination.
| Factor | TF-IDF (this project) | Embedding Models |
|---|---|---|
| Offline | Yes β pure maths, no model needed | Requires an embedding model |
| Speed | Instant vectorisation | Model inference per chunk |
| Storage | Sparse vectors (small) | Dense vectors (larger) |
| Quality | Good for keyword-heavy docs (CVs, JDs) | Better for semantic similarity |
| Dependencies | Zero | Requires embedding model download |
For interview preparation (CVs, job descriptions, interview guides), TF-IDF works well because these documents are keyword-rich and the queries are typically about specific skills, technologies, or job requirements.
To start fresh with a clean database, delete the SQLite file and re-ingest:
# Remove the existing database
rm data/rag.db # macOS/Linux
del data\rag.db # Windows (Command Prompt)
Remove-Item data\rag.db # Windows (PowerShell)
# Re-ingest the interview guide documents
npm run ingest
# Optionally ingest a new CV at the same time
npm run ingest -- path/to/your-cv.pdfTo clear only the uploaded documents but keep the interview guides, you can also delete files from uploads/ and re-ingest:
# Remove uploaded files
rm -rf uploads/* # macOS/Linux
Remove-Item uploads\* -Force # Windows (PowerShell)
# Remove DB and re-ingest base docs only
rm data/rag.db
npm run ingestTip: The database is a single file at
data/rag.db. You can back it up by simply copying it, or version different knowledge bases by renaming the file.
This project is designed to be adapted:
- Replace the documents in
docs/with your own domain-specific content - Edit the system prompt in
src/prompts.jsto match your domain and tone - Adjust chunk sizes in
src/config.jsβ smaller chunks for precise retrieval, larger for more context - Swap the model β change
config.modelto any Foundry Local-supported model (runfoundry model list) - Customise the UI β the frontend is a single HTML file with inline CSS, easy to modify
| Feature | Go (Original) | JavaScript (This Version) |
|---|---|---|
| AI Backend | GitHub Copilot SDK (cloud) | Foundry Local (offline) |
| Connectivity | Requires internet | 100% offline |
| Document Storage | None (single-use) | SQLite RAG database |
| CV Processing | Copilot SDK attachment | pdf-parse + TF-IDF chunking |
| Follow-up Chat | No | Yes (interactive) |
| Document Upload | No | Yes (PDF, MD, TXT) |
MIT β this is a learning sample. Fork it and make it yours.





