Imply Lumi Observability Warehouse Demo
In this 30-minute session, you'll see a demo of Imply Lumi — the observability data layer built to help you store more, search faster, and reduce cost without changing your existing tools.
Watch nowHey, everyone. Thanks for joining us today. My name is Matt Morrissey from Imply, and a little later in the session, I'll be joined by my colleague, Peter Marshall, for a product demo. Now before we get started, just a quick housekeeping note. If questions come up along the way, please feel free to drop them into the chat. We will keep an eye out for them throughout the session and answer them as we go. Now what we want to talk about today is something that we are hearing pretty consistently from almost every large security team that we speak with, and that's the need to access more and more data. The security teams, the folks that we're talking to, they're collecting more logs. They are keeping them for longer and increasingly using AI agents and automation to help investigations move faster. But all of those trends point in the same direction, the need for more context, more history, and more data to really reason over. Now at the same time, the economics of today's SIEM architectures, well, they're becoming harder and harder to justify. Bottom line, organizations are starting to look for new ways to store, to retain, and access security data. Now before I get too far ahead of myself, I did want to provide a little bit of context in terms of who we are. We are the company founded by the creators of Apache Druid, which is one of the leading real time analytics databases. And for the last decade, we've been helping organizations process and analyze just massive volumes of event data all in real time. Now that's given us a front row seat to how observability and security architectures behave at scale, and more importantly, where they start to break down as data volumes continue to grow. Loomi is a direct result of those experiences. Now today's SIEM architectures, they were built for a very different world. Organizations are collecting more data than ever before. You've got more cloud infrastructure, more applications, more identities, and and more endpoints. And now AI is coming into the picture, and it's just pouring gasoline on all of that because AI investigations, well, they, you know, they work better with more context and more history. Now the problem is that the economics of today's SIM providers, they they don't really work when data volumes start growing this quickly. So what we're seeing is that organizations end up making trade offs. Maybe they filter data before it ever reaches the SIM, or maybe they reduce retention from a year down to ninety days or even thirty days, or maybe they just accept that their SIM bill is gonna keep growing year over year. Now most of the teams that we talk to, well, they don't love any of those options. One of the biggest shifts that we're seeing right now is the move toward security lakes. Now the idea is pretty simple. Instead of putting all of your logs into expensive indexed infrastructure, you can start moving more of that data into object storage and separate storage from compute. Now this allows you to keep your your hot, your operational data where you need fast access for things like real time monitoring. And now you can move more of your historical data onto a separate compute layer that can query data directly into object storage. And when you think about it, those investigative workloads, well, they tend to be much more infrequent and bursty by nature anyway. Now you can see this shift happening everywhere. Splunk, for example, has federated search for s three and they also introduced Splunk Machine Data Lake. And then you've got folks like Databricks and Snowflake who are increasingly becoming part of that architecture as well. And overall, from Imply's perspective, we like this move. We think this is the right direction for the industry. When you think about it, the economics are super compelling. But if you look closely at the architecture here, what you're going to see is now you've got two different ways to access and interact with your data. You've got the traditional SIM workflow for your hot index data, and now there's a separated federated workflow for your historical data that's sitting out in object storage. And that's where some of these trade offs really start to show up. So what do I mean by that? Well, the big problem here is that those two query paths don't really come back together again. Your enterprise security dashboards and your detections that are powered by Splunk, for example, continue to run against your index data, but now the data lake becomes a second workflow that's primarily used for manual investigations. So ultimately, instead of having one place to answer questions, analysts end up bouncing between tools trying to piece the story back together. Now, big challenge that we're hearing about is that most of these systems were originally designed for structured analytical data, not operational log data. And again, when you think about logs, they are messy by nature. They are semi structured. They're constantly changing. New fields show up, existing fields can disappear. In different applications log the same event in completely different ways. But many of these architectures, they they still expect schemas. They require catalogs or some amount of data preparation before you can really start asking questions. And that's great for BI workloads, but it's much harder for investigations where the whole point is that you don't know what you're you're actually looking for yet. So organizations with this emerging security lake, they're solving the economics problem, but often end up creating a workflow problem. And that's really the gap Loomi was designed to address. So if that's if that's the problem, let's talk about what the ideal security lake architecture should actually look like. So for starters, we think it should provide a single search and a single access layer across both your hot and your historical data. Because when you think about it, analysts shouldn't have to know where the data lives before they know how to actually search it. We also think it should search logs like logs. I know it seems simple, but that means no catalogs, no schema managements, no data preparation before an investigation can actually begin. And we also don't think it should force you to change how you work. We want you to preserve and and really build on the workflows that your teams have already spent years creating. So that means keeping your dashboards, keeping your detections. And the last thing, we think it should provide a common access layer across your entire organization. So we're talking about security teams, your observability teams, even AI systems. We want that all working from the same underlying data. And that's really the architecture Loomi was designed to enable. So hopefully that gives you a sense of the problem we're trying to solve and what we think the ideal architecture looks like. And I think the net you know, the natural next question is, okay. So where does Imply Loomi fit into all of this? Now since a large percentage of the organizations that we work with are Splunk customers, let's look at this through the lens of a Splunk deployment. And honestly, I I think the easiest way to think about Loomi is as a replacement for the traditional Splunk indexer. So instead of sending data to a Splunk indexer, what you do is you send it to Loomy instead. And Loomy stores the logs, Loomy runs the SPL searches, and Loomy returns native Splunk events back into Splunk and applications like enterprise security. So from the perspective of the analyst that's sitting in Splunk, when you think about it, not much changes. They they're still using SPL. They're still using the same dashboards and the same detections. But what does change are the economics and the architecture underneath all of that. Now all of a sudden, it's it's much smaller, the infrastructure that you have to run, and it's much faster. And most importantly, you can do something traditional indexers simply can't do, which is query logs directly in object storage. So the way we think about it is pretty simple. Loomi is an indexer when you need indexing, and Lumi is also a federation layer when you don't. And the Splunk both look absolutely identical. So one of the reasons customers are able to save so much money with Lumi comes down to storage efficiency. We built Loomi. We designed it specifically for logs. And as a result, it stores machine data much more efficiently than today's sim architectures. And in our testing, a terabyte of raw log data that might consume close to a terabyte in a traditional platform ends up being closer to a third of that size in Lumi. And now once you start talking about tens or hundreds of terabytes a day, those differences add up very, very quickly. And storage efficiency is just one part of the story. The other part is performance. Now, of the trade offs organizations often make when they move data into object storage is that they just generally accept the fact that searches are going to be painfully slow. You save money? Yes, absolutely. But your security investigations now get a lot harder to execute and implement. Our view, again, is that you shouldn't have to make that trade off. In practice, we typically see Loomi searches run anywhere from two times to more than twenty times faster than traditional approaches. So what does that mean? That means analysts spend a lot less time waiting for searches to complete and a lot more time actually investigating problems and finding a root cause. And honestly, that's one of the biggest surprises for customers after deployment. They expect the cost savings, but they didn't necessarily expect searches to get that much faster too. Okay, so so far, we've talked about costs and storage efficiency. We've talked about the accelerated performance that you get with Lumi. But neither of those are actually the thing that customers get most excited about. The thing that tends to get people's attention the most is that Loomi can query logs directly where they live. Now, one of the things we've learned over the past few years is that organizations, they already have huge amounts of security data sitting out in object storage, in S3, in Delta Lake, in Apache Iceberg. The problem isn't storing the data. The problem is making that data searchable and making it actually usable during a security investigation. And historically, if you wanted to investigate that data, you had to build pipelines, you had to create catalogs, define schemas, move data around, or in some cases, rehydrate it back into expensive index infrastructure before you could actually search it. What Loomi does is change that entire workflow. With a feature called LogLake, Loomi can query unstructured logs directly where they live in object storage. So that means there's no pre indexing. There's no rehydration. There's no waiting for data movement jobs to complete. All you have to do is simply point Loomi at the data, and you can start searching it. And from the analyst perspective, it looks exactly like another Splunk search. Now the other trend that's changing security architectures right now is is obviously AI. Whether we're talking about copilots, AI agents, or automated investigations, they all have one thing in common. They work better with more context, more history, more evidence, more data to reason over. And, again, the challenge is that today's sims or the economics of today's sims don't really support that model. Nobody wants an AI agent kicking off hundreds of expensive searches against their most expensive infrastructure. But once you separate storage from compute, the economics, they change pretty dramatically. AI systems can now search years of data sitting out in object storage. They can pull context from a lot more sources. And because the compute is decoupled from the storage, those investigations, they don't impact your production monitoring environment. And connecting AI systems to that data is actually pretty straightforward. You can stand up a MCP server in front of Lumi and connect it to the AI tools your teams are already experimenting with today. So from the perspective of the analysts, again, not much changes. Your teams continue to use Splunk. The workflows stay the same. You're simply giving analysts and increasingly AI systems access to a much larger body of evidence that was ever practical before. Okay. At this point, the question we usually get is, how hard is this to actually deploy? The good news is that the deployment model is super simple. In most cases, it's really just some configuration changes. Forwarders that were previously sending data to Splunk can send data to Lumi instead. You also configure federation between Splunk and Lumi. And from that point forward, your analysts continue working in the same tools they're already using today. Now one of the reasons customers are able to move so quickly is that the operational model doesn't change for them. Workflows stay the same. But, again, the economics underneath the covers is what changes dramatically. So, I think at this point, I think it's a good idea to let's take a look at what this looks like from an analyst perspective. I've talked about the architecture, you know, the economics, the workflows, but I think the the best way to to make this more real is to actually show it. So what I'm gonna do at this point is hand things over to my colleague, Peter, to walk through a quick demo of Implyr. Peter, over to you. Hey, everybody. Imagine that we need to do a security investigation. We need to interact with some logs. These logs have been accumulated in s three. So here's my Databricks workbook. Imagine I'm midway through my investigation into some suspicious activity on the network. I think it's probably credential dumping. Behind this is a Delta Lake tape. I've got Windows Event logs in s three, about a terabyte's worth, and I need to see what happened. I need to find the pattern. I need to understand how far this threat has spread. Now the first thing that any analyst does, he says, let me see the raw events. What actually happened? The last time this ran, it took a really long time, just for a hundred rows. So I'm not going to run this right now. This isn't a resources problem. It's not even down to my fantastic SQL. It's because we spent months doing ETL to optimize my log data for one job, analytics. But now as a security investigator, I need the raw rows, and reconstituting and accessing and interacting with that raw log data, that's the real challenge for systems that are built for nice, clean, structured data. But we got there. So now we know there's something to investigate. Let's move on and find the behavior. Are there any processes touching LSAS, Windows Credential Store? And notice the search. It's a wildcard either side of LSAS. We can't be more precise than that. We don't know how the attacker named the path. This cell took a long time last time. You see, there's just no shortcut for the database. Interacting with logs without a good index means we need to look at every row every time. Let's move on. We've got a suspect process, Sophos OS query dot a x e. This is not a typo. It looks almost like a legitimate Sophos security agent, but actually, we now know that's the attacker. They named their malware really cleverly to blend in with security tooling. And we only caught it because we were searching for patterns, because we were doing ad hoc investigation. So three queries, we've just about scoped the incident. So SQL is the right language here, but without the right indexing layer underneath, our engine has no choice, every row, every time. And that's with a workbook I've already spent time crafting. You're not seeing anybody working out. None of the ad hoc stuff that an investor would gator would do to get to this point. So let me show you now how Lumi takes this to the next level where we can truly interact with this log data at speed. So this is where you get Lumi to build an index on your S3 buckets. We point Lumi at the bucket and away it goes, creating indexes, optimizing instruction, the log data ready for use. There are copy and paste instructions that just here in the setup tab. And if you wanna start consolidating more of your log sources into different workbooks, fine, just add another store, that's it. So let's see what Lumi does with the same questions. It's the same investigation, I'm gonna use the same wildcards. The first cell, it just contains credentials. I'm just gonna kick this off here. The first cell contains credentials. I'm gonna keep that collapsed for obvious security reasons. So what it is doing in that cell, it's connecting to Lumi and it's setting a time window just to give us exactly the same data that we just looked at in the other workbook. And here's where it gets interesting. We're defining a schema, but not to restructure the data. We're telling Spark which fields we want to work with. Time, some other source fields, and raw. Right? We are gonna look at the raw data. Lumi is exposing these directly. There's no casting. There's no JSON extraction. The data is already queryable as log data. And here we go. That cell is completed. Here's the raw data. Same query. Same date. Same question. Any process touching LSAAS? Let's see. Brilliant. Already done. Now the wildcard is still there. We still don't know exactly what the attacker named the path, but now Lumi's indexes are doing the heavy lifting everywhere they can. That's the difference. Sophos OS query dot exe, same fuzzy search, same pattern matching. We still caught the attacker's typo. Eight minutes on Delta Lake. Well, how long do you think it took me on Loomi? Hardly any time at all. Now we can actually investigate real ad hoc, undefined stuff, really dig in using SQL, just taking advantage of this really cool, very fast index layer that Lumi is providing direct to the s three data. The data never moved. It's still in s three. It's still in data lake. It's still in the catalog. It's still accessible to every other tool in the stack, unstructured or unstructured, Parquet, Delta Lake, Iceberg, whatever you've got. Now this is virtual tier. This is Lumi's compute layer. If you're used to modern data platforms, this is a no brainer, it's completely familiar to you. This is where we set up the on demand scalable compute sized to the workloads we're going to run. If you're coming to this from a traditional security tooling background, this might feel a bit alien. Those tools typically require you to pre provision. You have to pay for the headroom, whether you use it or not. That's not how it works in Lumi. Here, you spin up what you need when you need it, and it goes away when you're done. And we could spin up different tiers for different workloads, for monitoring, for investigations, for reporting. The investigation we just ran used a different tier to an always on tier that might use monitoring. There's no contention. Each workload gets what it needs. And here's what those tiers can serve. Databricks, Grafana, Tableau, Splunk, MCP, the tools your team already use or are thinking about using are all connected to the same data in S3 through this really fast efficient index underneath. The flexibility is the point. You decide how to resource each workload when you need it for as long as you need it. I haven't got to duplicate the data. I'm not creating another silo. I'm not rearchitecting my entire stack. So let's take Splunk as an example. I'm going to Splunk search here, and I connected my Splunk Cloud instance here to Lumi via federated search. And that's gonna give me exactly the same s three data we just investigated in the notebook. So I'm gonna ask the same question we asked in SQL. Have I got any processes touching LSAS, same data, same index? It's just a different language. And that's the interoperability point. Your security team works in their tool. Your data team works in theirs. Neither team has to compromise. And just to prove a point, let's look at the raw events. Right? We're gonna prove that this is looking back at exactly the same data in s three. What I haven't had to do is rehydrate anything, create a new pipeline. It's all just flowing in here out of s three directly. And that performance differential that you saw in the notebook, that is not specific to Spark. Every integration, every language, it's the same story, you get a speed improvement. So your Splunk searches, Grafana dashboards, Databricks notebooks, all faster, operating on the same data, minimal changes to your stack. And with that, back to you. Awesome. Thanks, Peter. That was great. What Peter just showed is really the pattern we're seeing over and over again with customers. One example is a customer of ours called b t BTG Pactual, one of the largest investment banks out in Latin America. Like a lot of organizations that we talked to, they weren't necessarily looking to replace Splunk. Their analysts knew SPL very well. They had years of dashboards, detections, workflows built around that platform. So the issue wasn't the workflow. The issue was the economics. They wanted to keep more data. They needed to retain it for longer. They were also looking to bring additional data sources into their Splunk environment or into their security and observability environment. Unfortunately, the cost curve was moving in a completely different direction. Historically, organizations in that position feel like they have to make two choices. One is keep their workflows or two, fix the economics somehow. What Lumi changed was that trade off. BTG was actually able to lower their cost per gigabyte by more than sixty percent. Now at the same time, they increased ingestion from fourteen terabytes a day to twenty five terabytes a day. So they're moving more data into their observability and their security platform. So they're getting even more coverage. And they also went from ninety days of retention to a full year. So they have access to all the context that they need. And they did it without changing the workflows their teams already relied on. So what we're seeing is similar patterns across universities, telecom providers and other financial services as well. That's really the outcome that we're trying to enable, is keeping more data, keeping it longer, and spending less money at the same time, and doing it without asking security teams to completely change how they work today. Okay, so Lumi as well as Imply, we work with customers across a pretty broad range of industries and company sizes today, telecom providers, banks like I was just talking about, manufacturers, retailers across different industries, different environments. But honestly, what we're seeing is a pretty similar architectural shift play out everywhere. More data moving into object storage, more interest in security like architectures, and more demand for more historical context to support those AI and investigation use cases. That's really where we think the industry is headed, and that's the future that we're building Lumi for. So if I had to boil all of this down to, you know, one thing, I think it's the fact that security teams shouldn't have to choose between keeping more data versus controlling their costs. Now, if you if you're thinking about implementing a security like architecture or looking for ways to reduce your Splunk costs without changing how your teams work, we'd love to talk. We're always happy to share what we're seeing across the market and walk through a more a more detailed demo of how customers are approaching these challenges today. And with that, I'd like to say thanks, everyone, for spending some time with us today. Appreciate it.
In this demo, you’ll see exactly how teams are using Lumi to do more with Splunk and modern observability stacks — without replatforming. We’ll also give you a first look at Lumi Loglake — our newest feature — which queries unstructured logs directly where they already live — in S3, Delta Lake, Apache Iceberg, and other open storage — bringing a true lakehouse (separated compute and storage) architecture to your observability and SIEM data.
Imply Lumi Observability Warehouse Demo
In this 30-minute session, you'll see a demo of Imply Lumi — the observability data layer built to help you store more, search faster, and reduce cost without changing your existing tools.
Watch nowLunch & Learn: Imply Lumi Observability Warehouse Demo
In this 30-minute session, you'll see a demo of Imply Lumi — the observability data layer built to help you store more, search faster, and reduce cost without changing your existing tools.
Watch nowImply Lumi: What’s New, What’s Next — and How to Unlock More Observability Value Today
Observability teams must retain more data, investigate faster, and control costs without disrupting existing tools. This live Imply Lumi update shows new ingestion, retention, search, Splunk interoperability,...
Watch now