Afloat16
## Problem filter_k_core accepts col_user and col_item and reads initial statistics using them, but its calls to min_rating_filter_pandas use the default names. Valid custom-schema rating data therefore raises KeyError during iterative user/item filtering. ## Reproduction Against the inspected source revision: ```python import pandas as pd from recommenders.datasets.split_utils import filter_k_core frame = pd.DataFrame({"uid": [1, 1, 2, 2], "iid": [10, 20, 10, 20]}) try: result = filter_k_core(frame, core_num=2, col_user="uid", col_item="iid") print("retained rows:", len(result)) print("columns:", list(result.columns)) except KeyError as error: print(type(error).__name__ + ": " + str(error)) ``` Observed before the candidate change: ```text KeyError: 'itemID' ``` ## Proposed change Forward both column names through both minimum-rating filters. Keep the graph-peeling algorithm, default names and core_num=0 behavior unchanged. ## Verification Before: 10 failed, 2 passed. Candidate: 12 passed, no skipped cases. Independent bipartite-graph peeling reference for three naming schemes and three core sizes; default schema; zero-core and empty-data controls; input immutability. Complete split_utils and constants were imported with real pandas/NumPy. Tests cover the pandas graph filter, not Spark or the full recommendation suite. Source: `recommenders/datasets/split_utils.py`, Git blob `fce197cf897ebbe0ce4d4133d99bf5c61565ef36`; branch `staging`. Python 3.13.5, Linux; exact dependency versions accompany the test logs. Searches: PR: filter_k_core; Issue: filter_k_core col_user. Old PRs #1678 and #1621 concern notebooks rather than custom column forwarding. ## Candidate source patch ```diff diff --git a/recommenders/datasets/split_utils.py b/recommenders/datasets/split_utils.py --- a/recommenders/datasets/split_utils.py +++ b/recommenders/datasets/split_utils.py @@ -181,10 +181,12 @@ if core_num > 0: while True: df_inp = min_rating_filter_pandas( - df_inp, min_rating=core_num, filter_by="item" + df_inp, min_rating=core_num, filter_by="item", + col_user=col_user, col_item=col_item ) df_inp = min_rating_filter_pandas( - df_inp, min_rating=core_num, filter_by="user" + df_inp, min_rating=core_num, filter_by="user", + col_user=col_user, col_item=col_item ) count_u = df_inp.groupby(col_user)[col_item].count() count_i = df_inp.groupby(col_item)[col_user].count() ``` ### Willingness to contribute - [x] Yes, I can contribute for this issue with guidance from Recommenders community.