Results of the first Apache Druid (incubating) community survey

Jun 25, 2019
Gian Merlino

We recently conducted our first Druid community survey. Every so often we’ll be asking our community a short set of questions to understand how they use the Druid database, and how they would like to see it improved. Thank you to everyone that participated in the survey. The responses are extremely helpful as we think about the Druid roadmap and the future of the project.

Here’s a summary of the survey results:

Use Cases

Over half of respondents indicated use cases that broadly fall into the realm of business and clickstream analytics. The top use cases for Druid include digital advertising/marketing analytics, user behavior analytics, and web/mobile event analytics. There was also a breadth of use cases with all different forms of event data, including application performance monitoring (APM), network telemetry, manufacturing analytics, security analytics, and IoT. It will be interesting to see how the mix of use cases evolves over time.


The large majority (72%) of respondents were running their Druid cluster in the cloud. Not surprisingly, AWS led (44%) the pack, followed by Google Cloud Platform (14%), Azure (6%) and OpenStack deployments (1%). Fewer than a third of respondents are running Druid in their own data center.


Manual cluster deployment was the most popular method (28%), but the answers were somewhat fragmented beyond that, led by Kubernetes (19%), and followed by Ansible (13%), Terraform (11%), and Docker Swarm (7%).


For data ingestion, Kafka (38%) and Hadoop (25%) ingestion methods accounted for nearly two thirds of the responses, with Tranquility, Kinesis, native batch and legacy real-time nodes rounding out the methods used. Sixty-five percent of respondents were using some form of streaming ingestion to load data into Druid. We anticipate this percentage to grow with time as streaming becoming more widespread.


There are a variety of front-end tools being used to query and visualize Druid data, led by Apache Superset (28%), followed by Imply Pivot and others such as Tableau, Looker, Metabase and Grafana. Roughly a quarter of the respondents have created their own custom UI. We should note that the majority of people were running more than one UI, with some running almost every option.


Lastly, we asked about preferred methods of community interaction. The tried and true channels of meetups (23%) and mailing lists (33%) were the majority, with Github and Slack following.


Our final question was a request for general feedback. Although there was a wide variety of things people wanted to see on the Druid roadmap, the most requested features were:

  • Joins
  • Better Kubernetes support
  • Simpler configuration

The good news is that we’ve been actively thinking about and working on these features and much more. Stay tuned over the next few weeks as we present more information on Imply is thinking about the Druid roadmap.

Other blogs you might find interesting

No records found...
Jun 17, 2024

Community Spotlight: Using Netflix’s Spectator Histogram and Kong’s DDSketch in Apache Druid for Advanced Statistical Analysis

In Apache Druid, sketches can be built from raw data at ingestion time or at query time. Apache Druid 29.0.0 included two community extensions that enhance data accuracy at the extremes of statistical distributions...

Learn More
Jun 17, 2024

Introducing Apache Druid® 30.0

We are excited to announce the release of Apache Druid 30.0. This release contains over 409 commits from 50 contributors. Druid 30 continues the investment across the following three key pillars: Ecosystem...

Learn More
Jun 12, 2024

Why I Joined Imply

After reviewing the high-level technical overview video of Apache Druid and learning about how the world's leading companies use Apache Druid, I immediately saw the immense potential in the product. Data is...

Learn More

Let us help with your analytics apps

Request a Demo