Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.
★ 29,359PythonForks 8,013
大数据入门指南 :star:
★ 16,979JavaForks 4,264
Enterprise job scheduling middleware with distributed computing ability.
★ 7,794JavaForks 1,381
Python clone of Spark, a MapReduce alike framework in Python
★ 2,662PythonForks 516
大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。
★ 1,827PythonForks 393
:dart: :star2:[大数据面试题]分享自己在网络上收集的大数据相关的面试题以及自己的答案总结.目前包含Hadoop/Hive/Spark/Flink/Hbase/Kafka/Zookeeper框架的面试题知识总结
★ 1,662Forks 444
MapReduce, Spark, Java, and Scala for Data Algorithms Book
★ 1,080JavaForks 652
C# and F# language binding and extensions to Apache Spark
★ 948C#Forks 209
distributed_computing include mapreduce kvstore etc.
★ 843GoForks 208
🐎 A serverless MapReduce framework written for AWS Lambda
★ 694GoForks 40
💎🔥大数据学习笔记
★ 677JavaForks 225
A serverless cluster computing system for the Go programming language
★ 555GoForks 35
Uniffle is a high performance, general purpose Remote Shuffle Service.
★ 454JavaForks 174
t-Digest data structure in Python. Useful for percentiles and quantiles, including distributed enviroments like PySpark
★ 408PythonForks 52
Compass is a task diagnosis platform for bigdata
★ 406JavaForks 150
Dynamic execution framework for your Redis data
★ 379RustForks 66
🎉🎉🐳 Datawhale大数据处理导论教程 | 大数据技术方向的开篇课程🎉🎉
★ 362PythonForks 48
Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows locally or on a cluster.
★ 355JavaForks 219
Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.
★ 282JavaForks 59
Firestorm is a Remote Shuffle Service, and provides the capability for Apache Spark and Apache Hadoop MapReduce applications to store shuffle data on remote servers
★ 255JavaForks 69
O'Reilly Book: [Data Algorithms with Spark] by Mahmoud Parsian
★ 229PythonForks 97
An easy-to-use Map Reduce Go parallel-computing framework inspired by 2021 6.824 lab1. It supports multiple workers threads on a single machine and multiple processes on a single machine right now.
★ 225GoForks 12
:zap: 6.824: Distributed Systems (Spring 2017). A course which present abstractions and implementation techniques for engineering distributed systems.
★ 213Forks 76
Companion to Learning Hadoop and Learning Spark courses on Linked In Learning
★ 203HTMLForks 166
A in-process MapReduce library to help you optimizing service response time or concurrent task processing.
★ 174GoForks 24
Big Data Modeling, MapReduce, Spark, PySpark @ Santa Clara University
★ 165HTMLForks 141
Java 实现的分布式系统课程(MIT6.824)
★ 155JavaForks 28
Tangseng search engine including full text search and vector search base on golang. 基于go语言的搜索引擎,信息检索系统
★ 139GoForks 40
Use the MapReduce's Java interface to distributed crawle the data of Chinese universities and learn basic knowledge of hdfs.
★ 133JavaForks 1
DTail is a distributed DevOps tool for tailing, grepping, catting logs and other text files on many remote machines at once.
★ 130GoForks 10