Research / SIEM Operations

SIEM Log Onboarding Methodology: A Production Checklist

A log source is not “onboarded” when packets arrive. It is onboarded when the organization can prove that events are complete, parsed, searchable, useful to detections and operationally supportable.

Cylatic Research · Engineering guide · Updated September 2026

1. Define the source contract

Every onboarding effort should begin with a written source contract. Record the product and version, deployment topology, log categories, protocol, sender addresses, destination listener, expected rate, authentication method, time zone, sample events, retention requirement and security use cases.

Do not accept “send all logs” as a sufficient requirement. Explicit categories make volume measurable and allow the SOC to identify blind spots.

Definition of done: another engineer should be able to reproduce the integration from the document without asking what “normal” means.

2. Verify transport before parsing

Start with packet-level or object-level evidence. Confirm that the source is actually sending, the receiver is listening, firewalls permit the path and the expected protocol is being used. For UDP, validate source IP and port and understand that transport itself provides no delivery guarantee. For TCP/TLS, verify connection establishment, certificate expectations and reconnect behavior.

Source → network path → listener → raw capture → parser

Check independently:
1. Is traffic arriving?
2. Is the payload intact?
3. Is the timestamp present?
4. Is the source identity trustworthy?
5. Does the collector acknowledge or buffer the event?

3. Capture representative samples

A single event is not enough. Collect samples covering successful and failed actions, different user roles, multiple devices, common and uncommon event types, and peak-volume periods. For cloud sources, include events from multiple accounts or subscriptions when the schema varies by tenant.

Store the original sample with the onboarding record. It becomes the regression fixture when the vendor changes its format.

4. Build parsing around stable semantics

Good parsers extract meaning rather than simply splitting strings. Identify timestamps, host, actor, source and destination, action, outcome, object, authentication method, process, rule/policy, bytes and vendor-specific identifiers where available.

Keep raw fields available. Normalized fields should make cross-source detections possible without destroying vendor context.

5. Normalize with a documented schema

Normalization should be intentional. Define canonical field names, types and semantics. A field called user should not mean a username in one parser and an email address in another without documentation. Distinguish event time from ingestion time. Preserve the original time zone or UTC conversion method.

event.time       = source event time
event.ingested   = collector arrival time
actor.user       = authenticated principal
source.ip        = originating network address
target.ip        = destination address
action           = normalized operation
outcome          = success | failure | unknown
vendor.*         = source-specific extensions

6. Validate with counts and field population

Validation needs quantitative checks. Compare source-side counts with SIEM counts over the same window. Measure parser success, event-type distribution, timestamp freshness and mandatory-field population.

A parser that accepts 99.9% of messages but silently maps the wrong field is more dangerous than one that rejects malformed events loudly. Prefer explicit failures and alerts over silent degradation.

7. Test detections before calling it complete

Onboarding should map directly to at least one useful investigation or detection. If the source is Palo Alto firewall telemetry, validate that network direction, action, source/destination and rule context support the intended use case. If the source is an identity provider, test authentication success/failure, MFA events, privilege changes and session context.

Use known-good events and controlled test events where permitted. Record expected results. A green parser status is not evidence that a detection can see the behavior it was designed for.

8. Build operational dashboards

Every production source should have health visibility. Useful panels include events per minute, last event time, parser failures, unknown event types, top event categories, ingestion latency and source-specific error codes. These are not decorative dashboards; they are the early warning system for telemetry loss.

9. Document ownership and change control

Record the source owner, SIEM owner, escalation path, credentials or certificate owner, parser version, change date and known limitations. Vendor upgrades should trigger a regression test against stored samples. Do not assume a vendor's “backward compatible” statement means the security fields are unchanged.

10. The production checklist

  1. Source owner confirmed.
  2. Security use cases documented.
  3. Protocol and network path validated.
  4. Representative raw samples retained.
  5. Parser tested against multiple event types.
  6. Normalized fields documented.
  7. Source-to-SIEM count comparison completed.
  8. Mandatory field population measured.
  9. At least one detection or investigation validated.
  10. Source health dashboard and alerting enabled.
  11. Retention and cost expectations approved.
  12. Change ownership and regression process documented.

What defenders should do now

Turn onboarding into an engineering product rather than a ticket-closing exercise. Require the checklist above for every high-value source. Maintain a small library of real samples for parser regression. Track “connected,” “parsed,” and “detection-ready” as separate states. This prevents organizations from reporting impressive onboarding numbers while still carrying major detection blind spots.

Conclusion

The most useful SIEM integration is not the one that produces the most events. It is the one whose data can be trusted during an incident. A disciplined onboarding methodology creates that trust by making transport, semantics, completeness and operational health measurable.