GithubHelp home page GithubHelp logo

ciusji / hurricanedb Goto Github PK

View Code? Open in Web Editor NEW

This project forked from guinsoolab/hurricanedb

0.0 0.0 0.0 88.17 MB

a real-time distributed OLAP engine

Home Page: https://guinsoolab.github.io/glab/#/app/home

License: Apache License 2.0

Shell 0.22% JavaScript 0.01% Python 0.02% Java 97.07% Scala 0.34% TypeScript 2.12% CSS 0.01% Thrift 0.03% HTML 0.06% Smarty 0.06% FreeMarker 0.01% Dockerfile 0.05% NASL 0.01%

hurricanedb's Introduction

badge
logo

HurricaneDB

HurricaneDB is a real-time distributed OLAP datastore, built to deliver scalable real-time analytics with low latency. It can ingest from batch data sources (such as Hadoop HDFS, Amazon S3, Azure ADLS, Google Cloud Storage) as well as stream data sources (such as Apache Kafka).

HurricaneDB was built by engineers at LinkedIn and Uber and is designed to scale up and out with no upper bound. Performance always remains constant based on the size of your cluster and an expected query per second (QPS) threshold.

For getting started guides, deployment recipes, tutorials, and more, please visit our project documentation at HurricaneDB.

Features

HurricaneDB was originally built at LinkedIn to power rich interactive real-time analytic applications such as Who Viewed Profile, Company Analytics, Talent Insights, and many more. UberEats Restaurant Manager is another example of a customer facing Analytics App. At LinkedIn, HurricaneDB powers 50+ user-facing products, ingesting millions of events per second and serving 100k+ queries per second at millisecond latency.

  • Column-oriented: a column-oriented database with various compression schemes such as Run Length, Fixed Bit Length.

  • Pluggable indexing: pluggable indexing technologies Sorted Index, Bitmap Index, Inverted Index.

  • Query optimization: ability to optimize query/execution plan based on query and segment metadata.

  • Stream and batch ingest: near real time ingestion from streams and batch ingestion from Hadoop.

  • Query with SQL: SQL-like language that supports selection, aggregation, filtering, group by, order by, distinct queries on data.

  • Upsert during real-time ingestion: update the data at-scale with consistency

  • Multi-valued fields: support for multi-valued fields, allowing you to query fields as comma separated values.

When should I use HurricaneDB?

HurricaneDB is designed to execute real-time OLAP queries with low latency on massive amounts of data and events. In addition to real-time stream ingestion, HurricaneDB also supports batch use cases with the same low latency guarantees. It is suited in contexts where fast analytics, such as aggregations, are needed on immutable data, possibly, with real-time data ingestion. HurricaneDB works very well for querying time series data with lots of dimensions and metrics.

Example query:

SELECT sum(clicks), sum(impressions) FROM AdAnalyticsTable
  WHERE
       ((daysSinceEpoch >= 17849 AND daysSinceEpoch <= 17856)) AND
       accountId IN (123456789)
  GROUP BY
       daysSinceEpoch TOP 100

HurricaneDB is not a replacement for database i.e it cannot be used as source of truth store, cannot mutate data. While HurricaneDB supports text search, it's not a replacement for a search engine. Also, HurricaneDB queries cannot span across multiple tables by default. You can use the Trino-HurricaneDB Connector to achieve table joins and other features.

Building HurricaneDB

More detailed instructions can be found at Quick Demo section in the documentation.

# Clone a repo
$ git clone [email protected]:GuinsooLab/hurricanedb.git
$ cd hurricandb

# Build HurricaneDB
$ mvn clean install -DskipTests -Pbin-dist

# Run the Quick Demo
$ cd build/
$ bin/quick-start-batch.sh

Join the Community

License

HurricaneDB is under Apache License, Version 2.0

hurricanedb's People

Contributors

ciusji avatar jeccy0779 avatar 306labs avatar

Recommend Projects

  • React photo React

    A declarative, efficient, and flexible JavaScript library for building user interfaces.

  • Vue.js photo Vue.js

    ๐Ÿ–– Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.

  • Typescript photo Typescript

    TypeScript is a superset of JavaScript that compiles to clean JavaScript output.

  • TensorFlow photo TensorFlow

    An Open Source Machine Learning Framework for Everyone

  • Django photo Django

    The Web framework for perfectionists with deadlines.

  • D3 photo D3

    Bring data to life with SVG, Canvas and HTML. ๐Ÿ“Š๐Ÿ“ˆ๐ŸŽ‰

Recommend Topics

  • javascript

    JavaScript (JS) is a lightweight interpreted programming language with first-class functions.

  • web

    Some thing interesting about web. New door for the world.

  • server

    A server is a program made to process requests and deliver data to clients.

  • Machine learning

    Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.

  • Game

    Some thing interesting about game, make everyone happy.

Recommend Org

  • Facebook photo Facebook

    We are working to build community through open source technology. NB: members must have two-factor auth.

  • Microsoft photo Microsoft

    Open source projects and samples from Microsoft.

  • Google photo Google

    Google โค๏ธ Open Source for everyone.

  • D3 photo D3

    Data-Driven Documents codes.