LuciferYang/spark
Mirror of Apache Spark
Apache Spark Committer&PMC Member/ Apache uniffle PMC Member/Lance Maintainer
Mirror of Apache Spark
Apache Iceberg
Apache Parquet Java
Apache Doris is an easy-to-use, high performance and unified analytics database.
Modern columnar data format for ML and LLMs implemented in Rust. Convert from parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..
OpenKB: Open LLM Knowledge Base
Agent Skills-compatible LLM wiki for Claude Code, Cursor, and Codex. Build a Karpathy-style knowledge base from raw sources, citations, and linting.
LLM Wiki is a cross-platform desktop application that turns your documents into an organized, interlinked knowledge base — automatically. Instead of traditional RAG (retrieve-and-answer from scratch every time), the LLM incrementally builds and maintains a persistent wiki from your sources。
A toolkit to run Ray applications on Kubernetes
Cosmos Curator is a powerful video curation system that processes, analyzes, and organizes video content using advanced AI models and distributed computing.
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
Distributed pushdown cache for DataFusion
DuckDB is an analytical in-process SQL database management system
AI-driven development workflow with /prd, /goal, /review-it and /ship-it skills
Skills for Real Engineers. Straight from my .agents directory.
A composable and fully extensible C++ execution engine library for data management systems.
Apache Iggy: Hyper-Efficient Message Streaming at Laser Speed
Apache Maka (Incubating) is a local-first AI agent workspace. Model messages, tool calls, tool results, permission decisions, and termination events are recorded as an append-only log.
Kyuubi is a unified multi-tenant JDBC interface for large-scale data processing and analytics, built on top of Apache Spark
Apache Flink Connector for Lance Vector Database
The lance extensions for DuckDB enable reading and writing of lance tables.
The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.
Apache GeaFlow: A Streaming Graph Computing Engine.
Java SDK for Milvus.
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
SGLang is a high-performance serving framework for large language models and multimodal models.
Apache Ossie, industry wide specification effort to standardize how we exchange semantic metadata across analytics, AI and BI platforms, providing a vendor neutral, single source of truth for semantic data
PyIceberg
Apache Iceberg - Go