ElenaNKn/data_projects

★ 0Forks 0HTMLGitHub ↗Compare

README

Data Analytics Projects

This repository contains data analytics projects, completed by Alena Kniazeva.

Data Analysis and Visualization Projects

Project name Description Tools & Skills
A/B-test analysis End-to-end A/B-test: from hypothesis formulation to result interpretation and product decision. The test revealed a statistically significant drop in 7-Day Retention. The product decision: roll back the feature. Design of experiment, homogeneity check, statistical significance of difference between 2 variables, analysis of A/B test results for binary metric. Tools: python libraries (pandas, numpy, math, scipy.stats, statsmodels.stats)
User Flow Sankey Diagram A user flow map was built for users who ordered a product on a b2c e-commerce resource and subsequently returned it. Problem points were identified, and recommendations for reducing returns were provided Database connection from Jupyter Notebook. EDA and Sankey diagram creation. Tools: SQL, Python (pandas, SQLAlchemy, requests, tqdm, plotly)

Machine Learning Projects

Project name Description Tools & Skills
Multiclass classification with imbalanced classes Сomprehensive data analysis, followed by the construction and tuning of four classification models: Logistic Regression, SVM, KNN, and Decision Tree. Model performance is evaluated using the weighted f1-score metric to take into account the class imbalance Multiclass classification (one-vs-rest and one-vs-one strategies), models: logistic regression, SVM, KNN, decision tree. Cross-validation for evaluation and hyperparameter tuning, Grid search, confusion matrix. Tools: python libraries (pandas, scikit-learn, matplotlib)
Feature Selection. Data clusterization Application of various methods for improving regression model performance through feature engineering (log transformation, polynomial features, standardization, feature selection, and categorical encoding) using a decision tree as an example. Application of clustering for data analysis and for estimating statistical characteristics of clusters Hyperparameter tuning via cross-validation, feature importance evaluation, log transformation of the target variable, pipeline construction, feature selection, data clustering. Tools: Python (pandas, scikit-learn, matplotlib)

Contacts

LinkedIn: https://www.linkedin.com/in/alena-n-kniazeva

Email: [email protected]

Telegram: https://t.me/KniazevaAlena (@KniazevaAlena)

Contributors

ElenaNKn

Issues