videommu: open-ended predictions corrupted by xlsx round-trip, silently zeroed or crash

#1698 · closed · 0 comments

View on GitHub ↗

rrrxxx0510

## Bug `vlmeval/dataset/videommmu.py`, `evaluate` (lines 668-706): ```python storage = get_intermediate_file_path(eval_file, '_score') # defaults to .xlsx ... data['parsed_pred'] = [ans[idx]['parsed_pred'] for idx in data['index']] dump(data, storage) ... score = aggregate_results([row for _, row in load(storage).iterrows()]) ``` `parse_open_response` (line 140) returns a **list** for open-ended items. Writing a list column to xlsx serializes it as the string `"['thirty', 30.0]"`; loading it back yields that string. `eval_open` (line 382) then does `for pred in pred_i` — iterating a string character-by-character, so every open-ended item is scored wrong (single-char substring matches). If the list column fails to write, `dump` (`vlmeval/smp/file.py:183-187`) falls back to `.pkl`, but `load(storage)` still targets the `.xlsx` path and crashes. ## Reproduction ```python import pandas as pd df = pd.DataFrame({'parsed_pred': [['thirty', 30.0], 'A']}) df.to_excel('/tmp/t.xlsx', index=False) print([type(x).__name__ for x in pd.read_excel('/tmp/t.xlsx')['parsed_pred']]) # ['str', 'str'] — list became "['thirty', 30.0]" ``` ## Impact Every full VideoMMMU run (includes the open-ended Adaptation subset) either silently zeroes open-ended accuracy or crashes at the end. MCQ items unaffected (plain-string `parsed_pred`). Silent / crash, no warning. ## Fix Use pkl for the score cache so list-typed `parsed_pred` survives the round-trip: ```python storage = get_intermediate_file_path(eval_file, '_score', 'pkl') ```

Comments