Thanks for your impressive work!
Several high-scoring Mem-Gallery reproductions (e.g., Omni-SimpleMem) route retrieval differently per task category. Questions carry oracle labels such as `[FR]`, `[TR]`, `[KR]` (or equivalent category metadata). Adapters parse these labels and apply **hard-coded, category-specific retrieval rules** on top of the memory backend.
This is not general-purpose memory retrieval — it is **benchmark-tuned routing that requires knowing the task type at test time**. Reported F1 numbers may therefore overstate deployable memory capability.
The adapter applies the following rules (example: Omni-SimpleMem `benchmarks/memgallery/adapter.py`):
| Category | Heuristic behavior |
|----------|-------------------|
| **FR** | Larger `top_k` (30; up to 40 for “list all / how many” questions) |
| **KR** | extra BM25 merge to surface both old and updated facts |
| **TR / CD** | Post-retrieval **chronological reordering** by `session_id` + `dialogue_id` |
| **VS / VR** | Separate visual path using an **image catalog BM25** over caption text |
| **TTL** | Additional image-catalog lookup when `image_caption:` appears in the query |
Please clarify whether category-conditioned retrieval is part of the intended Omini-SimpleMem design. If not, please:
1. Document this behavior explicitly in the benchmark README
2. Report scores without category-aware retrieval routing
Looking forward to your reply! @Jiaaqiliu
Thanks for the detailed and fair critique — I want to answer it straight rather than defensively.
You've described the current `benchmarks/memgallery/adapter.py` accurately. It does parse the category markers from the questions (`[FR]`, `[TR]`, `[KR]`, `[VS]`, `[VR]`, `[CD]`, `[TTL]`, etc.) and apply category-conditioned retrieval on top of the memory backend, specifically:
- category-dependent `top_k` (`_get_dynamic_top_k`, with a larger budget for `FR`, and up to 40 for "list all / how many" questions),
- an extra BM25 merge for `KR` (and for `FR` list questions),
- a separate image-catalog BM25 path for `VS`/`VR`,
- and category-influenced handling for the time/detail categories.
So your core observation is correct: as written, that adapter uses the task-category label at retrieval time, which means those specific numbers reflect category-aware routing rather than a single category-agnostic retrieval policy. That's an important distinction for anyone interpreting the scores, and it shouldn't be buried in code.
Concretely, I'll do what you asked:
1. **Document this explicitly** in the benchmark README — that the adapter is category-conditioned, and exactly which rules apply per category.
2. **Add/label a category-agnostic configuration** (single retrieval policy, no per-category routing) so the benchmark can be run and reported without the oracle category labels, and the two settings can be compared apples-to-apples.
I appreciate you taking the time to lay this out with the specific rules — it's a legitimate reproducibility/transparency issue and I'd rather fix the documentation and reporting than leave it ambiguous. I'll track the README + category-agnostic-mode work as a follow-up.