Question about the "Omni" claim: how were audio/video modules actually discovered and evaluated without corresponding benchmarks?

#69 · closed · 2 comments

View on GitHub ↗

springCozyRock

Hi, Omini-SimpleMem Team: The paper claims support for text, image, audio and video memory. However, the reported benchmarks (LoCoMo and Mem-Gallery) appear to evaluate only text and image modalities. How were the audio and video components optimized and validated during the autoresearch process? Thanks!

Comments

Jiaaqiliu

Good question, and a fair one to ask. To be precise about what is and isn't benchmarked: - **Architecturally**, the system defines all four modalities (`ModalityType.TEXT/IMAGE/AUDIO/VIDEO`) and has the corresponding ingestion machinery — e.g. entropy-based triggers for audio and visual streams (`triggers/audio_trigger.py`, `triggers/visual_trigger.py`) and a Whisper-based transcription path in the config. So the pipeline can ingest and store audio/video-derived memory units. - **Quantitatively**, the reported benchmarks (LoCoMo and Mem-Gallery) exercise the **text and image** modalities. There isn't a corresponding audio/video benchmark in the repo, and I don't want to claim quantitative validation we didn't run — the audio/video support is architectural/qualitative at this point, not something with head-to-head benchmark numbers behind it. I think the honest fix here is documentation: I'll make the README clear about which modalities have quantitative benchmark results versus which are supported at the architecture/ingestion level, so the "text, image, audio, video" capability claim isn't read as "all four were benchmarked." Thanks for keeping us precise about it.

springCozyRock

Thanks for your clarification!