Mcp for omnisimplemem

#70 · closed · 5 comments

View on GitHub ↗

iamvince

Any update on omnisimplemem for all mcp tool update?

Comments

Jiaaqiliu

Hi, thanks for the interest. Could you say a bit more about what you're after? To clarify the current state: - The MCP server (multi-tenant memory over Streamable HTTP + a legacy SSE transport) lives under `MCP/` and its mirror `simplemem/integrations/`, and exposes the memory tools (`memory_add`, `memory_query`, etc.). - `OmniSimpleMem/` is the multimodal (text/image) memory package. If you're asking for the full Omni/multimodal memory pipeline to be exposed through the MCP tool interface (i.e. multimodal `memory_add`/`memory_query` over MCP), that isn't wired up yet. If that's the ask, let me know your concrete use case and I'll track it as a feature request — otherwise, please share which specific tools you need and I'm happy to point you to the right entry points.

iamvince

Thanks for your reponse. I have been waiting for the full mcp version of omnisimplemem as currently multi modal are python only. I have been waiting for full mcp support for it with url support for minio and s3 and Google drive direct link sipport and file upload as input for image, audio, video and doc for memory and seperate isolated memory cluster for agents. And advance mem search to segmenting object, characters etc in multiframe image like segment anything meta to map graph search for similar objects, person etc from image search. Along with art style search to know like retro, anime etc. Similarly advance audio search and video search. In previous version I had worked on sampling In my custom simplemem but have been eagerly waiting for full mcp support with these features.

iamvince

Any update?

Jiaaqiliu

Update: the MCP server for Omni-SimpleMem is now on `main`. Multimodal memory is no longer Python-only — you can drive it from Claude Desktop or any MCP client over the stdio transport: ```bash cd OmniSimpleMem python -m omni_mcp --data-dir ~/.omni_simplemem/mcp ``` Full setup, Claude Desktop config and tool reference: [`OmniSimpleMem/omni_mcp/README.md`](https://github.com/aiming-lab/SimpleMem/blob/main/OmniSimpleMem/omni_mcp/README.md) ### What shipped, mapped to your asks **Multimodal ingestion over MCP** — `omni_add_text`, `omni_add_image`, `omni_add_audio`, `omni_add_video`, `omni_add_document` (`.txt/.md/.json/.csv/.yaml/.pdf/.docx`), plus `omni_query`, `omni_answer`, `omni_stats`, `omni_list_events`, `omni_consolidate`, `omni_list_namespaces`, `omni_delete_namespace`. 12 tools total. **URL / MinIO / S3 / Google Drive input** — every media argument accepts a local path, `file://`, `http(s)://`, a Google Drive share link (converted to a direct download automatically), `s3://bucket/key`, or `gs://bucket/object`. For MinIO or any S3-compatible store, point `S3_ENDPOINT_URL` at it and use the normal AWS credential env vars. Remote objects are downloaded to a temp file, ingested, then deleted; there is a configurable size cap (`OMNI_MCP_MAX_DOWNLOAD_BYTES`) and a kill switch (`OMNI_MCP_DISABLE_REMOTE`). **Isolated memory clusters per agent** — every tool takes an optional `namespace`. Each namespace is a genuinely separate cluster: its own storage dir, MAU store, vector index and event store. Namespace names are validated so they can't traverse the filesystem, and orchestrators load lazily behind an LRU cache so many agents can share one server without loading everything at once. ```jsonc {"name": "omni_add_image", "arguments": {"image": "s3://media/frame_001.png", "namespace": "agent_vision"}} ``` ### What did *not* ship — being straight with you The retrieval-side research features you listed are **not** implemented, and I don't want to imply otherwise: - SAM-style object/character segmentation across multi-frame images - graph search for "the same object/person" across images - art-style classification (retro, anime, …) - audio/video search beyond what the current processors produce (transcription for audio; entropy-sampled frame captions for video) Those need vision models and an indexing design well beyond wiring up MCP, so they're a separate piece of work rather than something I could honestly fold into this. Video ingestion samples visually significant frames and captions them; it does not segment or track objects. ### Notes - Text/document ingestion, stats and namespace management run fully offline. Image/audio/video captioning and `omni_query`/`omni_answer` call an LLM, so they need `OPENAI_API_KEY` (or an OpenAI-compatible gateway via `OPENAI_API_BASE`); without a key you get a clear error rather than a silent failure. - Ships with 50 offline tests (`OmniSimpleMem/tests/test_omni_mcp.py`) covering the protocol layer, namespace isolation and path-traversal rejection, media resolution including a real HTTP download, document extraction, persistence across restarts, and error handling. Please give it a try and open a new issue if you hit problems or if the sampling work you mentioned suggests changes. If you want to take a run at the segmentation/style-search side, I'd be glad to look at a design proposal or PR.

iamvince

Thanks for the update I will certainly try and rais the issue pr. Meanwhile for features not shared I think conditional sampling check could help it along with full api support. Such that if anyone using open model through coding agent plans like alibaba, codex, Claude etc then they same tool call can process through sampling so it does not have to go api route. I belive these features can make the memory as full multimodal memory system where it can compare the image on those params and list us matching file names. Because a image frame has art style, objects, characters etc and if our advanced image search can does that from agents memory in a domain names paced cluster it would be amazing. Refeence- resourcespace digital asset manager does through its clip plug in but not sure if it's advance hybrid search that I purposed. https://www.resourcespace.com/knowledge-base/plugins/clip-ai-smart-search