Cortex XSIAM data modeling normalizes logs from different vendors and products into a shared XSIAM Data Model (XDM) so analysts, correlation rules, dashboards, and detections can query consistent fields instead of memorizing every source schema. Current XSIAM provides a published XDM schema with typed fields, constants, and aliases such as IP, user, file hash, domain, hostname, cloud provider, and resource. Data Modeling Rules map raw datasets into this consolidated schema.
Within Palo Alto Security Operations, the data model is the foundation for scalable detection engineering. Normalization lets one detection reason about a source user, target host, file hash, or cloud resource across several data sources without rewriting the rule for every vendor.
Cloud-Native SIEM Architecture provides the broader reason normalization matters: correlation is only reliable when field meaning remains stable across ingestion sources.
XDM creates a consolidated schema
Current Cortex XSIAM data modeling rules map raw logs into predefined data types and XDM fields.
Instead of querying one vendor’s src_ip, another’s clientAddress, and a third’s nested JSON field separately, a normalized query can use the shared XDM representation.
This lowers detection-maintenance cost as the SOC adds or replaces data sources.
Aliases simplify cross-field hunting
XDM aliases group semantically equivalent fields. Current schema examples include aliases for IPv4/IPv6/IP, user, file/file hash, domain, hostname, country, resource, and cloud context.
A file-hash alias can represent hashes from source/target processes, modules, files, or email attachments.
Use aliases when the analytic intent is broad; use explicit fields when the distinction between source, target, or intermediate object matters.
Constants give normalized values consistent meaning
XDM constants define canonical values for concepts such as event tags, outcomes, privilege levels, user types, protocols, cloud providers, operating-system families, and agent types.
Normalization should map raw vendor-specific values into these shared constants.
This lets dashboards and detections compare events without building custom string-normalization logic into every query.
Data Modeling Rules map source data into XDM
Current Cortex XSIAM Developer Guide describes Data Modeling Rules as the mechanism for mapping logs from source datasets into the consolidated schema.
Rules should preserve the raw source dataset while populating normalized fields needed by content.
Test mappings against representative successful, failed, missing-field, and version-changed events so normalization does not silently collapse important distinctions.
Do not normalize away source fidelity
A shared schema cannot represent every vendor-specific detail perfectly.
Keep raw fields or source references available for investigation when analysts need the exact original event.
Use XDM for common analytics and raw data for deep source-specific troubleshooting instead of forcing every nuance into a generic field.
Field type consistency affects XQL and correlation
An IP stored as a string in one mapping and an IP type in another can break filters, comparisons, and enrichment.
Likewise timestamps, arrays, enums, booleans, and numeric fields should map to the documented XDM type.
Validate schema types before enabling detections that depend on those fields.
Normalization should be monitored for data-source drift
SaaS APIs, firewall firmware, cloud audit schemas, and endpoint products change field names and nested structures over time.
A source can continue ingesting while the mapping stops populating critical XDM fields.
Monitor mapping completeness and use health/correlation rules to detect sudden drops in normalized user, host, IP, or outcome values.
Detection content should depend on stable XDM fields where possible
Correlation and BIOC rules become more portable when they use normalized fields instead of one vendor’s raw schema.
Cortex XSIAM Detection Engineering should therefore treat data-model coverage as a prerequisite for cross-source analytics.
When a detection requires a source-specific raw field, document that dependency so teams know the rule will not generalize automatically.
Third-party content packs can include data-model rules
Current Cortex Marketplace content can include data-model rules alongside parsers, correlation rules, dashboards, and other content.
Review imported mappings before production because a content pack’s assumptions may not match your source version or organization-specific parsing.
Version-control local changes so marketplace updates do not overwrite important field semantics silently.
Data model changes are schema changes for the SOC
Renaming, repurposing, or changing the type of a normalized field can break searches, dashboards, automation, and detections.
Treat data-model edits like database/API schema changes: test downstream dependencies, stage rollout, and preserve backward compatibility where possible.
Document deprecations so detection engineers can migrate before the old mapping disappears.
Cortex XSIAM data modeling succeeds when one field means the same thing everywhere
The mature SOC uses XDM aliases and constants for cross-source analytics, keeps raw evidence accessible, validates types and mapping completeness, monitors schema drift, and treats data-model changes as controlled dependencies.
Normalization creates leverage only when analysts can trust that a user, host, hash, IP, or resource field carries consistent meaning across the telemetry estate.
Data-source onboarding should include a mapping acceptance checklist. Identify required XDM fields for planned detections, run representative success/failure events through the parser/model, confirm types and aliases, and record gaps. This prevents teams from declaring a connector ‘onboarded’ when events arrive but cannot support the analytics the SOC expected to build.
Normalization should preserve directionality. Source, target, and intermediate identities, hosts, processes, and files can have very different meanings in an attack chain. Use broad aliases for generic hunts, but do not collapse source and target into one custom field just because both contain an IP or user name.
Null handling deserves design. Some products omit a field when unknown; others send empty strings, zero values, `N/A`, or vendor-specific placeholders. Mapping rules should normalize these consistently so detection filters do not treat `unknown` as a real username or convert missing data into misleading evidence.
Timestamp modeling should distinguish event occurrence, collection, ingestion, and processing time where available. Delayed SaaS logs or offline endpoints can arrive hours after the event. Detections that correlate by ingestion time instead of event time can miss sequences or create false ordering, so choose the correct field intentionally.
Identity normalization should account for email addresses, UPNs, SAM account names, cloud principals, service accounts, device identities, and aliases. Use stable identity fields plus domain/tenant context, and avoid merging two principals only because a display name matches. Cross-source identity analytics are powerful only when entity resolution is trustworthy.
Cloud-resource normalization should preserve provider, account/project/subscription, region/zone, resource type, and canonical resource ID. A resource name alone may not be globally unique. Current XDM includes cloud provider/project/zone-style fields and aliases that should be populated so multi-cloud correlation can distinguish similar resource names.
Parser and model rules should be monitored independently. A parser can successfully extract raw fields while a data-model rule fails to map them, or a mapping can keep running after a parser changes semantics. Create health checks at both layers and sample normalized output after connector upgrades.
Content-pack model rules should be compared against local custom mappings before updates. Two rules mapping the same source field differently can create inconsistent analytics. Maintain a change log for overridden vendor content and retest detections when adopting a newer marketplace pack.
Schema governance should include deprecation periods. If a custom normalized field is replaced by a standard XDM field, populate both during a transition, migrate queries and dashboards, then remove the custom field after dependency inventory reaches zero. Abrupt schema cleanup saves little and can silently disable detections.
Data-model quality metrics can include percentage of events with required core fields, invalid-type rate, unknown outcome values, mapping-rule failures, lag between source version and parser update, and detections blocked by missing fields. These metrics turn normalization from a one-time integration task into an observable service.
Schema testing should include rare/error events, not only normal successful activity. Authentication failure, denied network traffic, missing usernames, malformed URLs, cloud API errors, and partial telemetry often populate fields differently. Detections focused on failures will be weakest precisely where mapping was never validated.
Entity enrichment should be kept conceptually separate from normalization. XDM should map what the source event says; enrichment can add asset criticality, identity risk, threat intelligence, or CMDB context afterward. Mixing enrichment assumptions into base mappings makes source truth harder to reconstruct.
Data-model ownership should be assigned per source. When a SaaS vendor changes schema, someone should own validating parser/mapping updates and notifying detection engineers. A shared SOC data platform without source owners tends to accumulate silent mapping drift that only becomes visible after a missed detection.
Search templates and dashboards should favor documented XDM fields over ad hoc aliases invented by individual analysts. Standardization creates compounding value: new analysts learn one schema, marketplace detections integrate more easily, and source migration requires fewer dashboard rewrites.
Data-model documentation should include example raw-to-XDM mappings for important sources. A short mapping sample helps analysts understand where normalized values came from and gives integration engineers a regression fixture when parsers change.
Quality review should include high-value fields used by automation. If an automation disables the user in `xdm.source.user`, confirm the mapping is truly the actor and not a target or service account. Incorrect semantics can turn normalization errors into automated response errors.
Protect semantic consistency.
Data-model governance should include field ownership and change control. When a source renames, retypes, or stops populating a field, detections and investigations depending on that semantic contract should fail visibly rather than drifting into partial or misleading results.