niting9881/free-data-science-books
Free resources for learning data science
Free resources for learning data science
"The question of whether a computer can think is no more interesting than the question of whether a submarine can swim." ― Edsger W. Dijkstra
"One person's data is another person's noise." ― K.C. Cole
Source codes for the book titled "Practical MLflow for Generative AI on Databricks"
Unlocking the Power of Health Data With a Modern Data Lakehouse
A curated list of safety-related papers, articles, and resources focused on Large Language Models (LLMs). This repository aims to provide researchers, practitioners, and enthusiasts with insights into the safety implications, challenges, and advancements surrounding these powerful models.
Awesome series for Large Language Model(LLM)s
Awesome-LLM: a curated list of Large Language Model
Capstone project -2
A retrieval-augmented assistant for identifying techniques from incident narratives using the MITRE ATT&CK knowledge base.
Delta Lake Optimization Project: Hands‑on lab to explore partitioning, Z‑Ordering, compaction (manual & auto), Liquid Clustering, and VACUUM using a synthetic sales dataset in Databricks. Includes a step‑by‑step notebook to measure file scans, bytes read, and query performance for each optimization.
Repository for Mlops zoomcamp Project
This day is all about context engineering
This is a repo with links to everything you'd ever want to learn about data engineering
Databricks PySpark Certification Prep Lab: Build an e-commerce analytics pipeline covering Spark DataFrame API, Structured Streaming, data skew handling with salting, broadcast joins, and Pandas UDFs. Designed for the Databricks Certified Associate Developer for Apache Spark exam.
Databricks Data Engineer Associate Certification Lab: End-to-end hands-on project covering Auto Loader, Medallion Architecture, SCD Type 2, Unity Catalog governance, and Databricks Jobs orchestration. Build a production-grade pipeline on Databricks Free Edition.
Databricks Real-Time Fintech Monitoring Pipeline: Hands-on lab to build a streaming fraud detection system using Auto Loader, watermarked deduplication, stream-static joins, and windowed rules engines in Databricks. Covers dual-SLA architecture for real-time alerts and batch compliance reporting.
Databricks DLT Apparel Pipeline Project: Learn medallion architecture, streaming, and data engineering with Delta Live Tables. Includes synthetic data, step-by-step guide, and certification prep.
100 essential Databricks concepts for data engineers, organized by category with difficulty levels and self-assessment scoring
Practice Databricks coding skills with hands-on exercises. Import into Databricks Free Edition, write code, run assertions, check pass/fail. Covers Delta Lake, Spark SQL, PySpark, Auto Loader, medallion architecture, window functions, and more.
Final capstone project