This repository contains data analytics projects, completed by Alena Kniazeva.
| Project name | Description | Tools & Skills |
|---|---|---|
| A/B-test analysis | End-to-end A/B-test: from hypothesis formulation to result interpretation and product decision. The test revealed a statistically significant drop in 7-Day Retention. The product decision: roll back the feature. | Design of experiment, homogeneity check, statistical significance of difference between 2 variables, analysis of A/B test results for binary metric. Tools: python libraries (pandas, numpy, math, scipy.stats, statsmodels.stats) |
| User Flow Sankey Diagram | A user flow map was built for users who ordered a product on a b2c e-commerce resource and subsequently returned it. Problem points were identified, and recommendations for reducing returns were provided | Database connection from Jupyter Notebook. EDA and Sankey diagram creation. Tools: SQL, Python (pandas, SQLAlchemy, requests, tqdm, plotly) |
| Project name | Description | Tools & Skills |
|---|---|---|
| Multiclass classification with imbalanced classes | Сomprehensive data analysis, followed by the construction and tuning of four classification models: Logistic Regression, SVM, KNN, and Decision Tree. Model performance is evaluated using the weighted f1-score metric to take into account the class imbalance | Multiclass classification (one-vs-rest and one-vs-one strategies), models: logistic regression, SVM, KNN, decision tree. Cross-validation for evaluation and hyperparameter tuning, Grid search, confusion matrix. Tools: python libraries (pandas, scikit-learn, matplotlib) |
| Feature Selection. Data clusterization | Application of various methods for improving regression model performance through feature engineering (log transformation, polynomial features, standardization, feature selection, and categorical encoding) using a decision tree as an example. Application of clustering for data analysis and for estimating statistical characteristics of clusters | Hyperparameter tuning via cross-validation, feature importance evaluation, log transformation of the target variable, pipeline construction, feature selection, data clustering. Tools: Python (pandas, scikit-learn, matplotlib) |
LinkedIn: https://www.linkedin.com/in/alena-n-kniazeva
Email: [email protected]
Telegram: https://t.me/KniazevaAlena (@KniazevaAlena)