Vvanakor
Back to writing
Observability2 min read

Your SRE team already built your security telemetry

ObservabilityOpenTelemetrySRESIEM

Two pipelines, one set of systems

Your SRE team instruments every service for latency, errors and saturation. Your security team instruments the same services for authentication, access and process execution.

Same hosts. Same agents, often. Two collection paths, two storage bills, and two answers to the question of what happened at 14:32.

That last one is the expensive part. Not the storage. During an incident you do not want two teams reconciling timestamps from two systems before anyone can act.

What to share

Collection and transport. One agent on the host, one schema on the wire, one transport. OpenTelemetry for collection, Kafka in the middle, and let each consumer take what it needs downstream.

The events are genuinely the same events. A process execution is a performance signal and a security signal depending on who is reading it. Collecting it twice is a decision nobody made on purpose — it accumulated.

What not to share

Retention and access. Security data has retention obligations that operational data does not, and access restrictions that operational data should not inherit. Merge the pipeline, not the policy.

I have seen the merge go too far exactly once, and the result was an SRE on-call with standing access to authentication logs for the whole estate. Nobody intended it. It fell out of a storage consolidation.

The number I am not going to give you

You will find posts claiming a specific percentage saved by converging these pipelines. I am not going to give you one, because it depends entirely on how much your two pipelines currently overlap, and that varies enormously between estates.

Measure it instead. List what each pipeline collects, by source and by event type. The overlap is your ceiling. Most teams can do this in an afternoon and the result is more convincing than anything I could assert.

The hard part is not technical

It is that these two teams have different on-call rotations, different tooling budgets and different definitions of urgent. The pipeline merge is a week of engineering. The agreement about who owns the schema is the part that takes a quarter.

Start there. The technology is the easy half.

Back to writing