extract_characters_regex: reversed containment test scores empty/missing predictions as 'A'

#1701 · open · 0 comments

View on GitHub ↗

rrrxxx0510

## Bug `vlmeval/dataset/utils/multiple_choice.py`, `extract_characters_regex` (lines 605-609): ```python matches = re.search(r'[ABCDE]', s) if matches is None: for choice in choices: if s.lower() in choice.lower(): # operands reversed return choice[1] return '' ``` The containment test is backwards — it asks "is the prediction a substring of the option label" instead of "is the option label in the prediction". Since the empty string `''` is a substring of every string, any prediction that reduces to empty matches the first choice `'(A)'` and returns `'A'`. ## Reproduction ```python s = '' choices = ['(A)','(B)','(C)','(D)','(E)'] for choice in choices: if s.lower() in choice.lower(): print(choice[1]) # 'A' <- empty prediction scored as A break ``` ## Impact Empty / whitespace-only / stripped-to-empty predictions score as `'A'` — correct on every item whose gold answer is `A`, a silent free point. ~1/(number of options) of failed generations become correct answers, invisible because callers treat `''` as the extraction-failed sentinel and score 0, while `'A'` flows through as a genuine prediction. Affected call sites: - `vlmeval/dataset/image_mcq.py:1083-1087` (MME-RealWorld): `if extract_pred == '': score = 0 else: score = int(extract_pred == ans)` - `vlmeval/dataset/image_mcq.py:2966-2970` (TreeBench): same shape - `vlmeval/dataset/sis_bench.py:132-140`: maps NaN to `''` then hands it to `extract_characters_regex`, which returns `'A'` — the NaN guard is completely defeated; missing predictions score as 'A'. ## Fix Swap the operands so the option label is matched within the prediction: ```python if choice.lower() in s.lower(): return choice[1] ```

Comments