Hi, thank you for your inspiring work!
While trying to reproduce the results on the MemGallery benchmark, I noticed that the provided adapter appears to use only textual information rather than the original images:
https://github.com/aiming-lab/SimpleMem/blob/main/OmniSimpleMem/benchmarks/memgallery/adapter.py#L144-L173
In particular, the _build_image_catalog function seems to construct the image catalog using image captions instead of the image content itself:
https://github.com/aiming-lab/SimpleMem/blob/main/OmniSimpleMem/benchmarks/memgallery/adapter.py#L202-L212
Could you please clarify whether the MemGallery results reported in the paper were obtained using a text-only evaluation protocol based on image captions?
If not, could you point me to the code or configuration used for evaluating the model with the original image modality?
Thank you!
Thanks — you're reading the adapter correctly, and it's a fair question.
In the current `OmniSimpleMem/benchmarks/memgallery/adapter.py`, the visual-question path (`_build_image_catalog` / `_image_search`) builds its catalog from the **textual `image_caption:` content** attached to image-bearing memory units and retrieves over that caption text with BM25. So, as written, the provided adapter's visual retrieval is caption/text-mediated rather than operating directly on raw image embeddings at query time.
That is a property of the adapter, and I don't want to overstate it as if it were the only possible protocol — but you've identified the actual behavior of the code in this repo accurately. I think the honest and useful thing to do is to **document the MemGallery evaluation protocol explicitly in the benchmark README** (what is indexed, what the visual path retrieves over, and where captions vs. raw image content enter the pipeline), so results are reproducible and not ambiguous. I'll add that. Thanks for pushing on this.