ScreenSpot: click points in 0-1000 coordinates (Qwen-VL style) are scored as misses

#1706 · open · 0 comments

View on GitHub ↗

yepapa-nest

ScreenSpot scores come out near zero for models that answer in 0-1000 normalized coordinates, which is the default for the Qwen-VL family. **What happens** `evaluate_rectangle` normalizes the ground-truth bbox to 0-1 (`convert_bbox` divides by the image size), and `parse_bbox_aguvis` takes the `x=`, `y=` values from the response as-is. So the scorer assumes the click point is already in 0-1. The prompt (`SYSTEM_PROMPT` / `USER_INSTRUCTION`) only asks for `pyautogui.click(x=?, y=?)` and never says which coordinate system to use, so each model answers in its native one. A Qwen3.8-based model served through the OpenAI-compatible API (GPT4V class) answers like `pyautogui.click(x=220, y=228)`, on a 0-1000 grid. Every one of those points lands outside [0, 1], so it is scored as a miss. The Format_Err_Rate stays at 0, which makes it look like a real capability result. **Numbers** (ScreenSpot Mobile / Desktop / Web, 1,272 samples, commit f71d473): | | as scored | same predictions / 1000 | |---|--:|--:| | Mobile | 0.8% | 90-97% | | Desktop | 8.4% | 95-99% | | Web | 1.6% | 87-89% | I got the right-hand column by dividing the parsed x, y by 1000 and checking them against the same normalized boxes (text/icon split within each range). **Possible fixes** 1. State the expected convention in the prompt (for example, "x and y are fractions of the image width and height, between 0 and 1"), or 2. Let the model/API wrapper declare its coordinate scale and normalize before scoring, or at least 3. Warn when parsed points fall outside [0, 1], since that almost always means a unit mismatch rather than a wrong answer. Happy to send a PR for whichever you prefer. Option 3 is the least invasive.

Comments