Topic: mapreduce

1,564 repositories

donnemartin/data-science-ipython-notebooks

Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.

★ 29,359PythonForks 8,013

PowerJob/PowerJob

Enterprise job scheduling middleware with distributed computing ability.

★ 7,794JavaForks 1,381

douban/dpark

Python clone of Spark, a MapReduce alike framework in Python

★ 2,662PythonForks 516

collabH/bigdata-growth

大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。

★ 1,827PythonForks 393

water8394/BigData-Interview

:dart: :star2:[大数据面试题]分享自己在网络上收集的大数据相关的面试题以及自己的答案总结.目前包含Hadoop/Hive/Spark/Flink/Hbase/Kafka/Zookeeper框架的面试题知识总结

★ 1,662Forks 444

microsoft/Mobius

C# and F# language binding and extensions to Apache Spark

★ 948C#Forks 209

bcongdon/corral

🐎 A serverless MapReduce framework written for AWS Lambda

★ 694GoForks 40

grailbio/bigslice

A serverless cluster computing system for the Go programming language

★ 555GoForks 35

apache/uniffle

Uniffle is a high performance, general purpose Remote Shuffle Service.

★ 454JavaForks 174

CamDavidsonPilon/tdigest

t-Digest data structure in Python. Useful for percentiles and quantiles, including distributed enviroments like PySpark

★ 408PythonForks 52

cubefs/compass

Compass is a task diagnosis platform for bigdata

★ 406JavaForks 150

datawhalechina/juicy-bigdata

🎉🎉🐳 Datawhale大数据处理导论教程 | 大数据技术方向的开篇课程🎉🎉

★ 362PythonForks 48

cwensel/cascading

Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows locally or on a cluster.

★ 355JavaForks 219

DigitalPebble/behemoth

Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.

★ 282JavaForks 59

Tencent/Firestorm

Firestorm is a Remote Shuffle Service, and provides the capability for Apache Spark and Apache Hadoop MapReduce applications to store shuffle data on remote servers

★ 255JavaForks 69

BWbwchen/MapReduce

An easy-to-use Map Reduce Go parallel-computing framework inspired by 2021 6.824 lab1. It supports multiple workers threads on a single machine and multiple processes on a single machine right now.

★ 225GoForks 12

xingdl2007/6.824-2017

:zap: 6.824: Distributed Systems (Spring 2017). A course which present abstractions and implementation techniques for engineering distributed systems.

★ 213Forks 76

kevwan/mapreduce

A in-process MapReduce library to help you optimizing service response time or concurrent task processing.

★ 174GoForks 24

CocaineCong/tangseng

Tangseng search engine including full text search and vector search base on golang. 基于go语言的搜索引擎,信息检索系统

★ 139GoForks 40

touero/ctenopharyngodon-idella

Use the MapReduce's Java interface to distributed crawle the data of Chinese universities and learn basic knowledge of hdfs.

★ 133JavaForks 1

mimecast/dtail

DTail is a distributed DevOps tool for tailing, grepping, catting logs and other text files on many remote machines at once.

★ 130GoForks 10