Research / Detection Engineering

Detection Engineering That Survives Production

A detection is a small piece of software running inside an uncertain data environment. Good engineering therefore requires a telemetry contract, explicit logic, validation evidence, false-positive controls, ownership and a way to measure when the rule has gone blind.

Cylatic Research · Engineering guide · Updated September 2026

1. Start with a behavior

Weak detection programs begin with products or keywords: “write a rule for VPN,” “alert on PowerShell,” or “detect this IOC.” Strong programs begin with a behavior and its evidence. Examples include suspicious authentication followed by privilege change, unusual service-account use, unexpected administrative access, or a cloud control-plane change inconsistent with the asset's role.

The behavior determines the telemetry contract. Without that contract, detection authors often compensate for missing data with increasingly complicated rules.

Design principle: a complex rule cannot repair a missing telemetry source. Fix the data contract before adding more logic.

2. Write the detection specification first

Before writing a SIEM query, define the intended behavior, required events, time window, threshold, exclusions, severity, response action and validation method. This turns a detection from an opaque query into an auditable engineering artifact.

Detection: suspicious privileged authentication
Inputs: authentication, MFA, privilege-change, asset context
Window: 15 minutes
Signal: successful privileged login after anomalous failures
Severity: high
Exceptions: approved jump hosts / break-glass procedure
Validation: controlled test + historical replay

3. Prefer combinations of weak signals

Most individual security events are ambiguous. A failed login can be a typo; a successful login can be normal; a privileged action can be routine. Their sequence and context can be much stronger. Correlation should therefore combine signals that are individually common but jointly unusual.

Be careful with correlation windows. A window that is too short misses slow behavior; one that is too long creates accidental relationships. Choose it from the actual behavior being modeled.

4. Engineer for missing data

Every detection has blind spots. If endpoint telemetry is missing, an endpoint-based rule may silently become useless. If cloud audit events are delayed, time-based correlation can misfire. Make data dependencies explicit and monitor them.

if required_source_freshness > threshold:
    mark_detection_health("degraded")
    suppress_false_confidence()
    notify_telemetry_owner()

A mature SOC distinguishes “no suspicious activity observed” from “the telemetry required to observe it was unavailable.”

5. Control false positives as part of the design

False positives are not merely an analyst inconvenience. They reduce trust in the detection system and consume scarce investigation capacity. Build exceptions around stable business facts rather than brittle IP lists wherever possible. Examples include approved administrative hosts, service accounts with defined ownership, sanctioned automation identities and known maintenance windows.

Every exception needs an owner and review date. Permanent exceptions become blind spots.

6. Validate with replay and controlled tests

A detection should have evidence that it works. Controlled tests are ideal when permitted. Historical replay can also reveal whether the rule would have fired during known incidents. Test both positive and negative cases: the intended behavior should trigger, while common legitimate behavior should not.

Record the test date, dataset, expected result and observed result. Treat detection changes like code changes.

7. Measure detection quality

Counting rules is a poor coverage metric. Better measures include the number of techniques or behaviors with validated coverage, alert precision, time-to-triage, telemetry freshness and the percentage of detections with recent validation evidence.

Track rules that have not generated a signal for a long period separately from rules that have generated many alerts. Silence can mean success, low prevalence or broken telemetry.

8. Detection lifecycle

  1. Hypothesis: define the behavior and attacker objective.
  2. Telemetry contract: identify exact fields and sources required.
  3. Prototype: implement the simplest explainable logic.
  4. Validate: test positive and negative scenarios.
  5. Deploy: assign owner, severity and response workflow.
  6. Tune: reduce noise using evidence, not arbitrary suppression.
  7. Revalidate: test after parser, product or infrastructure changes.
  8. Retire: remove detections whose assumptions are no longer valid.

9. Make alerts useful to the investigator

An alert should provide enough context to start an investigation: actor, asset, source and destination, timestamps, related events, business context and why the behavior is suspicious. A detection that merely says “possible attack” transfers the engineering work to the analyst.

Where possible, include the query or correlation rationale in internal documentation, not just a severity label.

10. Build detection dependencies into platform health

Detection health and platform health belong together. If a parser changes a field name, if an API quota is exhausted, or if an endpoint connector stops sending data, the SOC should know before an incident occurs.

What defenders should do now

  1. Create a one-page specification for every high-value detection.
  2. Document required fields and sources as a telemetry contract.
  3. Separate behavior logic from exception logic.
  4. Give every exception an owner and expiry/review date.
  5. Validate detections with controlled or replayed evidence.
  6. Track telemetry freshness and parser health as detection dependencies.
  7. Measure validated behavior coverage, not rule count.
  8. Retire stale detections instead of accumulating permanent noise.

Conclusion

Detection engineering becomes durable when rules are treated as production software. The query is only the implementation. The real engineering product is the combination of behavior model, telemetry contract, validation evidence, operational ownership and measurable coverage.