reza1615
My dataset has 4500 features and for selecting features the initial part before xgboost takes around 1 hour to run. I realized most part of the featurwize is not parallel is it possible to make them parallel?
#128 · closed · 3 comments
My dataset has 4500 features and for selecting features the initial part before xgboost takes around 1 hour to run. I realized most part of the featurwize is not parallel is it possible to make them parallel?
Wherever I could use n_jobs=-1 I have used it. Other than that, I have not used multithreading which is now available in XGBoost. This could be something to think about. Ram
we should inspect each part of the pipeline regarding to time to see how we can reduce the timing by vectorize or parallelize with Joblib or similar libraries
yes that is doable but I don't have the time. If you or anyone is interested I can help out. Ram