Hexiaoqiao/ClassicPapers
Papers about Distributed System
Papers about Distributed System
Ongoing research training transformer models at scale
Mirror of Apache Spark
HDFS file read access for ClickHouse
Apache Curator
Apache ZooKeeper
Mirror of Apache Hadoop
blog
Mirror of Apache Ratis (Incubating)
Kvrocks is a distributed key value NoSQL database that uses RocksDB as storage engine and is compatible with Redis protocol.
The ASF Website
Apache Hadoop Site
Kudu, a native column store for the Hadoop ecosystem. Fast analytics on fast data.
Apache Iceberg
A library that provides an embeddable, persistent key-value store for fast storage.
Apache Hadoop Ozone
Zipkin is a distributed tracing system
Apache Arrow is a cross-language development platform for in-memory data. It specifies a standardized language-independent columnar memory format for flat and hierarchical data, organized for efficient analytic operations on modern hardware. It also provides computational libraries and zero-copy streaming messaging and interprocess communication. Languages currently supported include C, C++, Java, JavaScript, Python, and Ruby.
An Open Source Machine Learning Framework for Everyone
asynchronously synchronise local NSS databases with remote directory services
Hops Hadoop is a distribution of Apache Hadoop with distributed metadata.
A tool for scale and performance testing of HDFS with a specific focus on the NameNode.
Refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20
Mirror of Apache Hive
Mirror of Apache HBase
Mirror of Apache Flink
Google core libraries for Java
Ceph is a distributed object, block, and file storage platform
Mirror of Apache HTrace (Incubating)