making parallel most part of the featurwize

#128 · closed · 3 comments

View on GitHub ↗

reza1615

My dataset has 4500 features and for selecting features the initial part before xgboost takes around 1 hour to run. I realized most part of the featurwize is not parallel is it possible to make them parallel?

Comments

AutoViML

Wherever I could use n_jobs=-1 I have used it. Other than that, I have not used multithreading which is now available in XGBoost. This could be something to think about. Ram

reza1615

we should inspect each part of the pipeline regarding to time to see how we can reduce the timing by vectorize or parallelize with Joblib or similar libraries

AutoViML

yes that is doable but I don't have the time. If you or anyone is interested I can help out. Ram