lujangus
`DUDE_acc` writes its extraction table to TSV, re-reads it with pandas, and calls `.lower()` on every prediction. An empty extraction is written as an empty string and returns from pandas as `NaN`, so scoring dies with `AttributeError: 'float' object has no attribute 'lower'` and the run reports `judge_fail_rate 100%`. The reported failure rate points at the judge. The judge was correct, and so was the model. ### The guard that is present, beside the one that is not `vlmeval/dataset/dude.py`, `DUDE_acc`, lines 37 to 41 at commit `2ae7b28b6033c05f211041a2e53a5a51e4a62504`: ```python if isinstance(item['answer'], float) and math.isnan(item['answer']): item['answer'] = 'Not answerable' item['answer'] = item['answer'].lower() item['pred'] = item['pred'].lower() ``` `answer` is guarded against `NaN`. `pred` is not, and `pred` is the field that can be empty, because it holds what the judge extracted rather than what the dataset published. ### Why an empty extraction is not an edge case DUDE contains unanswerable questions by design. When the question asks for a field the document leaves blank and the model correctly answers that it is blank, the judge extracts nothing. An empty extraction is the normal output of a correct answer to an unanswerable question. One row of 6,315 reached this in our run. The whole set scored zero and the reported cause was the judge. ### Reproduction, with no model and no dataset ```python import pandas as pd, math, tempfile, os df = pd.DataFrame([{'answer': 'blank', 'pred': ''}]) p = os.path.join(tempfile.mkdtemp(), 'r.tsv') df.to_csv(p, sep='\t', index=False) v = pd.read_csv(p, sep='\t').iloc[0]['pred'] print(repr(v), isinstance(v, float) and math.isnan(v)) # nan True v.lower() # AttributeError ``` For the full path, run DUDE against any model that answers an unanswerable question correctly and read `judge_fail_rate` against the traceback. ### Proposed change One guard, mirroring the one three lines above: ```python if isinstance(item['pred'], float) and math.isnan(item['pred']): item['pred'] = '' ``` ### Scope Any scorer in the harness that round-trips a prediction through `dump` and `load` and then calls a string method on it has the same exposure. This report claims only the DUDE path, which is the one we measured. ### Versions VLMEvalKit 0.2rc1 at commit `2ae7b28b6033c05f211041a2e53a5a51e4a62504`, dated 2026-09-18, read on 2026-09-21.