Why You Shouldn’t Have to Delete Your VPC Flow Logs

Jul 24, 2026
Dave Klein

When a security incident happens, investigators almost always start with the same questions:

  • Which systems communicated?
  • Where did the traffic originate?
  • What changed before the incident?
  • Was data exfiltrated?

AWS VPC Flow Logs are often the fastest way to answer those questions. They provide the network-level visibility needed to reconstruct events, identify suspicious communication patterns, and understand how an attacker moved through an environment.

Unfortunately, they’re also one of the largest and fastest growing sources of security telemetry.

As cloud environments expand, VPC Flow Logs quickly become one of the most expensive datasets to retain inside a traditional SIEM. Many organizations eventually face the same difficult choice:

  • Shorten retention from years to months, or months to weeks.
  • Filter out lower-priority traffic.
  • Archive raw logs into object storage where they become difficult or impossible to search.

None of those decisions improve security.

They’re simply attempts to manage storage costs.

The problem, of course, is that investigators rarely know which data they’ll need until an incident occurs. By then, the logs may already be gone.

This is exactly why security data lake architectures are gaining momentum.

Rather than forcing organizations to store every log inside expensive SIEM infrastructure, a security data lake separates inexpensive storage from interactive investigation. High-volume datasets like VPC Flow Logs remain in object storage while continuing to be searchable from familiar investigation tools.

That’s where Imply Lumi fits.

Lumi serves as the shared data layer beneath your existing security tools. Instead of replacing Splunk or requiring analysts to learn new workflows, Lumi continuously ingests VPC Flow Logs from Amazon S3, automatically parses and enriches them, and makes them immediately searchable from Splunk, Grafana, Databricks, and other supported tools.

In this walkthrough we’ll configure that integration from start to finish.

Why VPC Flow Logs Are the Perfect Security Data Lake Use Case

Not every security dataset has the same characteristics.

Authentication logs tend to be relatively small.

Application logs vary significantly depending on workload.

VPC Flow Logs are different.

They combine three characteristics that make them ideal candidates for a security data lake:

  • Extremely high volume
  • High investigative value
  • Relatively infrequent day-to-day access

Most days, they simply accumulate.

On the day an incident occurs, however, they often become one of the most important datasets investigators have.

That’s exactly the type of data organizations want to retain for years, not weeks, without paying premium SIEM storage costs.

Choose the Ingestion Method That Fits Your Environment

Lumi supports multiple ways to ingest VPC Flow Logs depending on how your environment is already configured.

If your VPC Flow Logs already land in Amazon S3—which is the most common deployment—we recommend using S3 Pull, which we’ll use throughout this guide.

Already sending flow logs into Splunk?

You don’t need to rebuild your ingestion pipeline.

Lumi can receive the same data through existing Splunk Universal Forwarders, Heavy Forwarders, or HEC endpoints, allowing you to preserve existing operational workflows while dramatically reducing long-term storage costs.

Decide Whether to Import Historical Data

Before connecting AWS, decide whether you want to:

  • Backfill existing VPC Flow Logs
  • Continuously ingest new logs

Most organizations begin by enabling continuous ingestion and add historical backfill later if needed. That’s the approach we’ll follow here.

Give Lumi Secure Access to Your Existing Data

One of the advantages of Lumi’s architecture is that your data doesn’t need to move into proprietary storage before it becomes searchable.

Instead, Lumi securely reads directly from the Amazon S3 bucket where your VPC Flow Logs already reside.

Creating the required IAM policy takes only a few minutes.

Once attached to an IAM role, Lumi can begin discovering and ingesting new objects automatically.

Why this matters

Traditional SIEM architectures often require data to be copied into proprietary storage before analysts can search it.

Lumi leaves the data where it already lives while providing fast interactive search through a shared data layer.

Create an IAM Key

Next, create an IAM key inside Lumi.

Beyond authentication, the key also enriches incoming events with useful metadata such as:

  • Environment
  • Team
  • Source Type

This additional context becomes valuable later when correlating investigations across multiple data sources.

Enable Continuous Ingestion

The final AWS configuration uses Amazon SNS notifications to tell Lumi whenever new VPC Flow Logs arrive.

From that point forward, ingestion becomes automatic.

Every new object written into your S3 bucket is detected, parsed, transformed, and indexed without additional operational work.

From Raw Flow Logs to Investigation-Ready Data

Once connected, newly arriving VPC Flow Logs appear immediately inside Lumi.

Behind the scenes, the built-in pipeline automatically:

  • Parses raw records
  • Extracts important fields
  • Normalizes values
  • Maps data into Splunk’s Common Information Model (CIM)

Instead of manually transforming cloud telemetry before it’s useful, analysts receive investigation-ready events immediately.

Continue Investigating from Splunk

One of the biggest advantages of a shared data layer is that analysts don’t have to change how they work.

Using Splunk Federated Search, analysts can continue searching VPC Flow Logs directly from Splunk while the underlying data remains stored in low-cost object storage.

Existing investigation workflows stay intact.

The same shared dataset can also be accessed from Grafana, Databricks, and other Lumi integrations without creating additional copies of the data.

The Integration Takes Minutes. The Pipeline Does the Heavy Lifting.

While the setup itself is straightforward, the pipeline is what creates lasting value.

Every incoming VPC Flow Log automatically flows through a predefined transformation pipeline that:

  • Parses incoming records
  • Extracts structured fields
  • Normalizes data
  • Maps events into CIM
  • Produces investigation-ready events

Instead of spending time preparing cloud telemetry for analysis, security teams receive normalized, searchable data automatically.

The Bottom Line

VPC Flow Logs perfectly illustrate why organizations are adopting security data lake architectures.

They’re incredibly valuable during investigations.

They’re incredibly expensive to retain inside traditional SIEMs.

For years, security teams have been forced to choose between keeping more data or controlling costs.

A shared data layer removes that tradeoff.

With Lumi, VPC Flow Logs remain in inexpensive object storage while staying continuously searchable from Splunk and other familiar investigation tools. Analysts keep the workflows they already know. Security teams keep the evidence they need. Organizations stop letting storage costs determine how much history they can investigate.

If you want to see what Lumi can do for your organization, request a demo.

Other blogs you might find interesting

No records found...
Jun 11, 2026

Supercharging Schema-On-Read: Logs in Object Storage Don’t Need a Data Catalog

Machine data architectures are rapidly changing. As telemetry volumes continue to grow and as costs rise, organizations are increasingly moving logs and other machine data into object stores such as AWS S3....

Learn More

Ready to decouple your observability stack?
No workflow changes. No migrations. More data, less spend.

Request a Demo