Doubt about Omini-SimpleMem SOTA adapters cheating Mem-Gallery with category-conditioned retrieval routing

#68 · closed · 1 comments

View on GitHub ↗

springCozyRock

Thanks for your impressive work! Several high-scoring Mem-Gallery reproductions (e.g., Omni-SimpleMem) route retrieval differently per task category. Questions carry oracle labels such as `[FR]`, `[TR]`, `[KR]` (or equivalent category metadata). Adapters parse these labels and apply **hard-coded, category-specific retrieval rules** on top of the memory backend. This is not general-purpose memory retrieval — it is **benchmark-tuned routing that requires knowing the task type at test time**. Reported F1 numbers may therefore overstate deployable memory capability. The adapter applies the following rules (example: Omni-SimpleMem `benchmarks/memgallery/adapter.py`): | Category | Heuristic behavior | |----------|-------------------| | **FR** | Larger `top_k` (30; up to 40 for “list all / how many” questions) | | **KR** | extra BM25 merge to surface both old and updated facts | | **TR / CD** | Post-retrieval **chronological reordering** by `session_id` + `dialogue_id` | | **VS / VR** | Separate visual path using an **image catalog BM25** over caption text | | **TTL** | Additional image-catalog lookup when `image_caption:` appears in the query | Please clarify whether category-conditioned retrieval is part of the intended Omini-SimpleMem design. If not, please: 1. Document this behavior explicitly in the benchmark README 2. Report scores without category-aware retrieval routing Looking forward to your reply! @Jiaaqiliu

Comments

Jiaaqiliu

Thanks for the detailed and fair critique — I want to answer it straight rather than defensively. You've described the current `benchmarks/memgallery/adapter.py` accurately. It does parse the category markers from the questions (`[FR]`, `[TR]`, `[KR]`, `[VS]`, `[VR]`, `[CD]`, `[TTL]`, etc.) and apply category-conditioned retrieval on top of the memory backend, specifically: - category-dependent `top_k` (`_get_dynamic_top_k`, with a larger budget for `FR`, and up to 40 for "list all / how many" questions), - an extra BM25 merge for `KR` (and for `FR` list questions), - a separate image-catalog BM25 path for `VS`/`VR`, - and category-influenced handling for the time/detail categories. So your core observation is correct: as written, that adapter uses the task-category label at retrieval time, which means those specific numbers reflect category-aware routing rather than a single category-agnostic retrieval policy. That's an important distinction for anyone interpreting the scores, and it shouldn't be buried in code. Concretely, I'll do what you asked: 1. **Document this explicitly** in the benchmark README — that the adapter is category-conditioned, and exactly which rules apply per category. 2. **Add/label a category-agnostic configuration** (single retrieval policy, no per-category routing) so the benchmark can be run and reported without the oracle category labels, and the two settings can be compared apples-to-apples. I appreciate you taking the time to lay this out with the specific rules — it's a legitimate reproducibility/transparency issue and I'd rather fix the documentation and reporting than leave it ambiguous. I'll track the README + category-agnostic-mode work as a follow-up.