Hi, Omini-SimpleMem Team:
The paper claims support for text, image, audio and video memory. However, the reported benchmarks (LoCoMo and Mem-Gallery) appear to evaluate only text and image modalities. How were the audio and video components optimized and validated during the autoresearch process?
Thanks!
Good question, and a fair one to ask.
To be precise about what is and isn't benchmarked:
- **Architecturally**, the system defines all four modalities (`ModalityType.TEXT/IMAGE/AUDIO/VIDEO`) and has the corresponding ingestion machinery — e.g. entropy-based triggers for audio and visual streams (`triggers/audio_trigger.py`, `triggers/visual_trigger.py`) and a Whisper-based transcription path in the config. So the pipeline can ingest and store audio/video-derived memory units.
- **Quantitatively**, the reported benchmarks (LoCoMo and Mem-Gallery) exercise the **text and image** modalities. There isn't a corresponding audio/video benchmark in the repo, and I don't want to claim quantitative validation we didn't run — the audio/video support is architectural/qualitative at this point, not something with head-to-head benchmark numbers behind it.
I think the honest fix here is documentation: I'll make the README clear about which modalities have quantitative benchmark results versus which are supported at the architecture/ingestion level, so the "text, image, audio, video" capability claim isn't read as "all four were benchmarked." Thanks for keeping us precise about it.