[BUG] Comparison operators keep nulls while `str` predicates fill them, on the same string dtype

#24393 · open · 0 comments

View on GitHub ↗

a-hirota

## Environment cudf 26.08.01 with pandas 3.0.6, where the default string dtype is `str` and `_PANDAS_NA_VALUE` is `np.nan`. ## Reproducer ```python import cudf import pandas as pd c = cudf.Series(["1", None, ""]) p = pd.Series(["1", None, ""]) print("cudf == '' ", (c == "").to_arrow().to_pylist()) print("pandas == '' ", (p == "").tolist()) print("cudf contains('1')", c.str.contains("1").to_arrow().to_pylist()) print("pandas contains('1')", p.str.contains("1").tolist()) ``` ## Actual ``` cudf == '' [False, None, True] pandas == '' [False, False, True] cudf contains('1') [True, False, False] pandas contains('1') [True, False, False] ``` `str.contains` / `startswith` / `endswith` / `match` / `isalpha` fill nulls with `False` on this dtype; the comparison operators (`==`, `!=`, `<`, `<=`, `>`, `>=`) return nulls. ## Expected The two should agree on the same column. If the NaN-semantics string dtype is meant to follow pandas, the comparison operators should also produce `False` for null rows. ## A second inconsistency in the same area ```python cudf.Series([None, None], dtype="str") == "" # [False, False] null_count 0 cudf.Series([None, None, "x"], dtype="str") == "" # [None, None, False] null_count 2 cudf.Series([None, None], dtype="str") != "" # [True, True] cudf.Series(["a", None], dtype="str") != ["a", None] # [False, None] ``` An all-null column compares without nulls; adding one valid element makes the same rows null. pandas returns `False` / `True` in both cases. This also reproduces on cudf 26.06.01. A consequence: a single-row frame never produces a null from a comparison, because a single-row column is necessarily either all-null or null-free and can never be mixed. ## Note Numeric comparison also propagates nulls (`cudf.Series([1, None, 3], dtype="int64") > 1` gives `[False, None, True]`, pandas float64 gives `[False, False, True]`). That appears to be the established null model rather than a new divergence, so this report is limited to the string dtype, where one predicate was aligned with pandas and the comparison operators were not.

Comments