LuQQiu/ElasticSearch
A movie search engine based on ElasticSearch using Python
software engineer / machine learning engineer [email protected]
A movie search engine based on ElasticSearch using Python
Real time stock data pipeline --play with Kafka, Cassandra, Spark, Redis, Node.js, Zookeeper
Library for bringing distributed capabilities to Apache DataFusion
Benchmark for vector databases.
Modern columnar data format for ML and LLMs implemented in Rust. Convert from parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, with more integrations coming..
Spark integrations for working with Lance datasets
Lance Namespace Specification is an open specification on top of the storage-based Lance data format to standardize access to a collection of Lance tables
Developer-friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps!
Apache DataFusion SQL Query Engine
Qdrant - High-performance, massive-scale Vector Database for the next generation of AI. Also available in the cloud https://cloud.qdrant.io/
Alluxio filesystem spec implementation
Alluxio Python client - Access Any Data Source with Python
A specification that python filesystems should adhere to.
AI reference architecture with Alluxio
Ray is a unified framework for scaling AI and Python applications. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
A distributed compute and storage engine
Alluxio, formerly Tachyon, Unify Data at Memory Speed
File system test suite.
State-of-the-art 2D and 3D Face Analysis Project