---
title: "Solving SIEM Data Gravity: How to Cut Ingest Costs Without Losing a Single Security Event"
publish_date: 2026-10-06
author: "Rohit Dharampal"
channel:
  - name: "Product"
    url: "/ja/resources/blog/channel/product.md"
tags:
  - name: "Apache Kafka"
    url: "/resources/blog/tag/apache-kafka.md"
  - name: "architecture"
    url: "/resources/blog/tag/architecture.md"
  - name: "data streaming"
    url: "/resources/blog/tag/data-streaming.md"
  - name: "Distributed"
    url: "/resources/blog/tag/distributed.md"
  - name: "Logs"
    url: "/resources/blog/tag/logs.md"
  - name: "Observability"
    url: "/resources/blog/tag/observability.md"
  - name: "Optimization"
    url: "/resources/blog/tag/optimization.md"
  - name: "Real-time"
    url: "/resources/blog/tag/real-time.md"
  - name: "Security"
    url: "/resources/blog/tag/security.md"
---

# Solving SIEM Data Gravity: How to Cut Ingest Costs Without Losing a Single Security Event

## Key takeaways

- SIEM platforms such as Splunk and Elastic are priced largely on ingest volume, and a large share of that volume is repetitive, low-value telemetry.
- MariaDB GridGain is a distributed cache and compute layer within the MariaDB Enterprise Platform that summarizes benign events before ingest, reducing SIEM-bound volume by about 40% while preserving 100% of security-relevant events in a reference-implementation run.
- This solves a data gravity problem: instead of shipping every raw event to the SIEM to be indexed and then analyzed, MariaDB GridGain moves compute to the data and reduces the amount of data at the streaming layer.
- The approach complements — it does not replace — your SIEM. An open reference implementation on GitHub demonstrates it using Apache Kafka and a three-node MariaDB GridGain cluster.
 
 

Your SIEM bill climbs every year — but most of what you pay to ingest and index is repetitive, benign noise. The usual fixes, sampling or dropping data, trade away the security visibility you bought the SIEM for in the first place. There is a better option: reduce the noise **in memory, before it ever reaches the SIEM** — while preserving 100% of security-relevant events.

![]()*An in-memory reduction and compute layer sits between your telemetry pipeline and your SIEM. Benign events are summarized; security-relevant events pass through unchanged.*

## Why SIEM costs keep climbing — and why it isn’t really the SIEM’s fault 

Security information and event management (SIEM) platforms such as Splunk and Elastic are priced and sized largely on how much data you ingest and index. Meanwhile, enterprise telemetry keeps growing — often 25% or more per year. Every new log source, endpoint, and cloud service adds volume, and each additional gigabyte adds ingest fees, index load, and storage pressure.

Here is the uncomfortable part: much of that data is repetitive and benign. Firewall allows, routine DNS lookups, successful authentications, and repeated cloud API polls can dominate a stream while carrying little investigative value. You end up paying premium rates to index noise. The common responses — sampling events or dropping fields at the agent — cut the bill, but they also throw away fidelity you may need during an investigation.

### What is data gravity? 

Data gravity is the tendency of a large data store, and the services and compute built around it, to pull everything toward itself. In security operations, that pull looks like this: every raw event is shipped to the SIEM to be indexed first, and only then analyzed. The heavier your central store becomes, the more expensive it is to move new workloads to it — and the more your ingest, storage, and license costs grow.

Data gravity, not the SIEM itself, is what inflates the bill.

## What is MariaDB GridGain? 

First, some context on the history of the MariaDB GridGain. It’s based on the acquisition of GridGain,the pioneer of in-memory computing on [March 24, 2026](https://mariadb.com/newsroom/press-releases/mariadb-completes-gridgain-acquisition-to-power-the-next-generation-of-agentic-ai/).

MariaDB GridGain, integrated with MariaDB Enterprise Platform, is an in-memory distributed cache and compute solution that instead of querying data on disk, distributes data across the RAM of a cluster of nodes and runs logic directly on the node that owns the data, an approach known as colocated compute, or moving compute to the data.

That design reverses data gravity. Rather than pulling every event back to a central store to be processed, you process it where it already is, in memory, at microsecond speed. For a telemetry pipeline, that means you can filter, group, deduplicate, and summarize high-volume streams before they ever reach the SIEM — and do it with a distributed state that is shared across the whole cluster, not trapped inside a single process.

## The pattern: an in-memory reduction layer in front of your SIEM 

The architecture separates three responsibilities cleanly. Apache Kafka transports telemetry. MariaDB GridGain holds the distributed reduction state and performs the in-memory computation. Your SIEM remains the system of record for indexing, alerting, investigation, and long-term retention.

### How reduction works 

Incoming events are assigned to short, time-bounded processing windows. Within a window, benign events that share a reduction key — for example, the same source, destination, and action — are grouped, counted, and summarized into a compact record. The SIEM then receives one summarized event instead of thousands of near-identical lines. Because the state that tracks those keys lives in the distributed cache, reduction stays consistent across every worker and scales as you add nodes.

### How security-relevant events stay intact 

Reduction never touches security-relevant events. Anything flagged as security-relevant — an authentication failure, a policy violation, anomalous or blocked traffic — bypasses the reduction step entirely and is forwarded to the SIEM unchanged and immediately. Preservation is a hard rule, not a tuning knob, which is what makes it safe to place this layer in front of a SIEM.

### Why distributed in-memory state matters 

You could try to deduplicate events in a single-node cache, but its state is trapped in one process and capped by one machine’s memory — so reduction can’t be shared or scaled. A distributed in-memory cache and data grid keeps reduction state partitioned across the cluster, with backup replicas for resilience, so decisions are shared fleet-wide and grow horizontally. The table below compares the common approaches.

| **Approach** | **What it does** | **Trade-off** |
|---|---|---|
| **Drop or sample at the agent** | Cuts volume by discarding events or fields | Loses security fidelity you may need later |
| **Single-node in-memory cache** | Deduplicates within one process | State capped by one machine; not shared or scalable |
| **In-memory distributed cache and compute (MariaDB GridGain)** | Distributed, windowed reduction with shared state; security events bypass | Preserves security events; scales horizontally |
| **Ingest everything into the SIEM** | Full fidelity, no pre-processing | Highest ingest, index, and storage cost |

## What a reference implementation shows 

To validate the pattern, we built an open reference implementation using synthetic firewall, DNS, Windows Active Directory, and cloud telemetry, a real Apache Kafka pipeline, and a three-node embedded MariaDB GridGain cluster. A representative run produced the following results.

| **Metric** | **Value** |
|---|---|
| **Raw events ingested** | 10,000 |
| **Events after reduction** | 6,000 |
| **Volume reduction** | ~40% |
| **Security-relevant events preserved** | 170 / 170 |
| **Preservation rate** | 100% |

In this reference-implementation run, MariaDB reduced SIEM-bound volume by roughly 40% while preserving every security-relevant event. Exact numbers vary with telemetry mix, window size, and configuration — but the invariant is full preservation of security events. The same run also proved distribution: reduction state was partitioned across all three nodes, each carrying both primary and backup ownership.

## Beyond SIEM: the same platform solves data gravity elsewhere 

SIEM reduction is one instance of a broader pattern — placing an in-memory data and compute layer in front of an expensive, volume-priced system to do the cheap work early. The same MariaDB Cache can:

- **Cut observability cost —** filter and aggregate high-volume log streams before they hit Splunk or Elastic, cutting ingest volume by up to about 50% depending on telemetry mix;
- **Relieve legacy systems —** cache frequently read data from a mainframe or relational core in memory, offloading read queries so you avoid costly compute and license expansion;
- **Enable real-time analytics —** run enrichment, scoring, and correlation in memory as data streams in, for real-time decisions instead of after-the-fact batch analysis.

## When an in-memory reduction layer is the right fit 

This approach fits best when you have high-volume, repetitive telemetry, an ingest-priced SIEM, and a hard requirement to keep every security-relevant event. It is a complement to Splunk or Elastic, not a replacement — your SIEM stays the system of record. And the reference implementation is a starting point to learn the pattern and size the savings, not a turnkey, production-hardened product.









## Frequently Asked Questions

 ### What is SIEM data reduction?

  

SIEM data reduction is the practice of filtering, deduplicating, and summarizing telemetry before it is ingested and indexed by a SIEM, so you lower ingest volume and cost. Done well, it removes repetitive benign events while preserving all security-relevant ones.

 

 

 



 

 ### Will reducing SIEM ingest cause me to lose security events?

  

It should not. In this pattern, security-relevant events bypass reduction and are forwarded unchanged; only benign, repetitive events are summarized. In the reference implementation, 100% of security-relevant events were preserved.

 

 

 



 

 ### What is data gravity?

  

Data gravity is the tendency of a large data store and its surrounding services to pull workloads and additional data toward it. In security, it shows up as shipping every raw event to the SIEM before analysis, which inflates ingest, storage, and license costs.

 

 

 



 

 ### How is MariaDB GridGain different from Redis or a single-node cache?

  

A single-node cache, including a standalone Redis node, keeps state in one process, capped by one machine’s memory. MariaDB GridGain distributes state and compute across a cluster in RAM, with partitioning, backup replicas, and SQL — so state is shared fleet-wide and scales horizontally.

 

 

 



 

 ### Does MariaDB GridGain replace Splunk or Elastic?

  

No. MariaDB GridGain is an in-memory reduction and compute layer that sits in front of the SIEM. Splunk or Elastic remains the system of record for indexing, alerting, investigation, and retention.

 

 

 



 

 ### How much can I save?

  

Savings track your reduction rate. The reference implementation showed about 40% volume reduction, and filtering high-volume log streams in memory can cut ingest by up to roughly 50% depending on the telemetry mix. Because SIEM cost scales with ingest, that reduction maps directly to lower license and storage spend.

 

 

 



 

 ### How do I try it?

  

Explore the open reference implementation on GitHub, run it against sample telemetry, then request a MariaDB GridGain license or a proof of concept to size the approach against your own data.

 

 

 



 

 

  









## Learn More

- [MariaDB In-Memory Cache](https://mariadb.com/products/in-memory-cache/)
- [MariaDB GridGain SIEM Demo on GitHub](https://github.com/mariadb-rohitdharampal/gridgain-siem-demo)
- [MariaDB GridGain Documentation](https://www.gridgain.com/docs/gridgain8/latest/getting-started/concepts)