Super Vault is a self-hosted source-ingestion and retrieval system for Hermes Agent. It collects content from sources you choose, stores the original text and metadata locally, indexes that content for search, and lets Hermes answer questions using the stored sources.
For a machine-readable, step-by-step install procedure—including prerequisite checks, Qdrant, API keys, source connectors, indexing, visualization, and verification—use INSTALL_WITH_AGENT.md. It includes a copy-paste prompt that tells a Hermes agent exactly what to install and what evidence it must return before claiming success.
For local versus hosted Qdrant, embedding collections, and background-job setup, read VECTOR_AND_BACKGROUND.md. It explicitly separates the starter’s current automatic behavior from modules that an agent must configure and test.
Use Super Vault when you want an agent to keep track of information you read or save over time.
It can store:
- 🌐 Web pages and articles — fetch the page, extract readable text, and save it as Markdown.
- 🎥 YouTube videos — fetch a transcript with Supadata or
youtube-transcript-api. - 📰 RSS and Atom feeds — monitor publications with
feedparserand ingest new articles. - ✉️ Substack posts — optionally fetch posts from subscribed publications using an authenticated session.
- 🔖 X bookmarks — optionally read a user’s saved X posts through
xurl. - 📝 Pasted notes, PDFs, and GitHub READMEs — save material that does not come from a tracked feed.
For each source, the system stores the full source text, URL, title, source type, tags, capture time, and a content hash. It can then search that material by keywords, by semantic similarity, or by relationships between extracted entities.
A Vault Pulse is a scheduled digest of new sources added since the last pulse. It is intended to show what entered the vault, not to replace the original source material.
After a connector is configured and manually verified, Hermes can run it in the background with cron jobs. A reader checks a feed, newsletter, or bookmark list; skips IDs and URLs already recorded in SQLite/state files; fetches full unseen sources; saves them through the same scan pipeline; then queues indexing. INSTALL_WITH_AGENT.md defines the required state, job order, retries, and Pulse behavior.
Super Vault installs as a Hermes skill. The skill recognizes @scan in any Hermes surface that supports chat messages, including CLI, Discord, and Telegram.
@scan https://example.com/article
@scan https://youtube.com/watch?v=VIDEO_ID
@scan A pasted note or research excerpt
When you use @scan, Hermes should:
- Fetch the source text or accept the pasted text.
- Convert it to clean text when needed.
- Normalize the source URL and check whether it was already saved.
- Write a Markdown file with YAML metadata.
- Add the content to full-text, vector, and graph indexes.
- Return the saved or duplicate status and the source location.
The repository includes the local storage bootstrapper, canonical Markdown saver, web/YouTube scanner, Hermes skill definition, Qdrant Docker service, and local graph viewer. RSS, Substack, X bookmark polling, scheduled pulses, and full Qdrant/LightRAG indexing are connector modules to enable as the project expands.
The full agent-oriented procedure is in INSTALL_WITH_AGENT.md. For a normal local install:
git clone https://github.com/ianlapham/super-vault.git
cd super-vault
python3 scripts/setup.py --root ~/super-vault-data
# Install the Hermes instruction layer, then start a fresh Hermes session.
hermes skills tap add ianlapham/super-vault
hermes skills install super-vaultsetup.py creates a Python environment, installs requirements.txt, copies .env.example to local .env, creates the source/SQLite layout, and starts Qdrant. It stops with an error if Docker is missing; use --no-qdrant only when intentionally setting up without vector search.
Super Vault works as plain local Markdown by default; Obsidian is optional. To make the data folder easier to recognize/open in Obsidian, opt in during setup:
python3 scripts/setup.py --root ~/super-vault-data --obsidianThis only adds an OBSIDIAN.md marker. It does not install Obsidian, create a cloud account, or turn on cross-device sync. If you use Obsidian, open the same --root folder as a vault. For multi-device sync, separately choose Obsidian Sync, iCloud/Dropbox, or a private Git workflow.
Job: track data sources, fetch their contents, and convert each source into clean text that can be stored and indexed.
| Source type | How it is fetched and parsed | Technology |
|---|---|---|
| Web page | Download HTML, remove navigation/scripts/styles, keep article/main text | requests, Beautiful Soup |
| YouTube | Request a transcript, then join transcript segments into text | Supadata Transcript API, youtube-transcript-api fallback |
| RSS / Atom | Poll feed XML, compare article URLs against saved state, ingest new URLs | feedparser, SQLite |
| Substack | Use an authenticated Substack session to request post HTML, then convert HTML to text | Substack API, Python urllib, Beautiful Soup |
| X bookmarks | Read authenticated bookmarks and post text, then send each item through the same scan pipeline | xurl |
| Pasted text | Save the supplied text without fetching a URL | Hermes @scan trigger, Python |
All source types should use the same scan pipeline. This gives every source the same metadata format, deduplication rule, storage layout, and indexes.
Job: smartly store the contents and information for any data source.
- Markdown files store the source content. Each source is written as a separate Markdown file. The body contains the full extracted text. YAML frontmatter contains title, original URL, canonical URL, source type, capture time, tags, and SHA-256 content hash.
- SQLite stores source records and full-text search. SQLite records which URLs were scanned and prevents duplicate saves. Its FTS5 extension indexes titles and source text for fast keyword search.
- Qdrant stores vector embeddings. The system splits each source into small text chunks (typically about 200 words), creates an embedding for each chunk, and stores that embedding plus source metadata in Qdrant. This supports meaning-based search, such as finding documents about “agent memory” even when they do not use those exact words.
- Raw and derived data stay separate. Original source text, Markdown, database records, vector embeddings, summaries, and graph data are separate files or stores. If an index is deleted, it can be recreated from the Markdown corpus.
Job: find the source text most relevant to a user question and give Hermes enough context to answer with citations.
- Keyword search: SQLite FTS5 finds exact titles, names, phrases, and terms.
- Semantic search: Qdrant finds text chunks with similar embedding vectors to the query.
- Reranking:
sentence-transformersscores the candidate chunks again withcross-encoder/ms-marco-MiniLM-L-6-v2and keeps the most relevant results. - Answer generation: Hermes receives the ranked source passages and their URLs/files, then generates an answer that points back to those sources.
Example: a search for “how are AI agents evaluated?” can return a saved article about benchmarks, a newsletter about eval harnesses, and a note that uses the phrase “testing agent behavior,” even if it does not use the exact query wording.
Job: extract entities and relationships from source text so the system can answer connection questions across multiple documents.
LightRAG reads saved text and extracts entities such as people, companies, products, concepts, and claims. It also extracts relationships between them.
Saved newsletter
└─ mentions → OpenAI
└─ provides → embedding model
└─ indexes text in → Qdrant
A graph query can answer questions such as: “What sources connect OpenAI, embeddings, and Qdrant?” LightRAG combines relationship traversal with vector retrieval, so it can use both graph connections and relevant text passages.
For visualization, NetworkX reads the entity graph, Louvain community detection groups related entities, and the included HTML5 Canvas viewer displays nodes, clusters, and search. The visualization export contains entity labels, types, community IDs, and edges only. It must not include source text, source URLs, filesystem paths, API keys, or embeddings.
This repository contains code and configuration templates only. It does not contain:
- Saved sources, notes, PDFs, raw documents, or personal graph data
- API keys, browser cookies, X sessions, or
.envfiles - Local SQLite databases, Qdrant data, LightRAG storage, or logs
MIT