ASiegeLion/blaze
Blazing-fast query execution engine speaks Apache Spark language and has Arrow-DataFusion at its core.
Blazing-fast query execution engine speaks Apache Spark language and has Arrow-DataFusion at its core.
2020年最新总结,阿里,腾讯,百度,美团,头条等技术面试题目,以及答案,专家出题人分析汇总。
A course of building an LSM-Tree storage engine (database) in a week.
List of Computer Science courses with video lectures.
Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.
电子书《跟顶级项目学编程-仿Hadoop从0到1完整实现高性能、可扩展的RPC框架》的源代码
Apache Spark - A unified analytics engine for large-scale data processing
Apache Hadoop
Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.
Gluten: Plugin to Double SparkSQL's Performance
FlinkSQL数据脱敏和行级权限解决方案及源码,支持面向用户级别的数据脱敏和行级数据访问控制,即特定用户只能访问到脱敏后的数据或授权过的行。此方案是实时领域Flink的解决方案,类似于离线数仓Hive Ranger中的Row-level Filter和Column Masking方案。
Apache Flink
Coral is a translation, analysis, and query rewrite engine for SQL and other relational languages.
A C++ vectorized database acceleration library aimed to optimizing query engines and data processing systems.
基于antlr4 解析器,支持spark sql, tidb sql, flink sql, Spark/flink jar 运行命令解析器
Apache Calcite
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
专注大数据学习面试,大数据成神之路开启。Flink/Spark/Hadoop/Hbase/Hive...
A search engine which can hold 100 trillion lines of log data.
Demo for service oriented application hosted on Hadoop YARN cluster for HA and scheduling
状态机通用实现
Real time data processing system based on flink and CEP
轻量级,微内核加插件机制,基于Java的RPC框架。可看成是mini版的Dubbo。提供服务注册,发现,负载均衡,支持API调用,Spring集成和Spring Boot starter使用。
使用scala的Akka编写的一个简易rpc框架