Data Quality and Observability in Fabric: One Operational System

Data quality and observability are often treated as separate disciplines. Quality teams define rules about completeness, validity, and consistency, while platform teams monitor runs, capacity, and errors. In production those signals belong together. A pipeline can be technically healthy while publishing incorrect data, and a quality rule can fail because the platform never delivered the expected partition.

Within DP-700, the stronger operational model treats quality and observability as one system: telemetry explains how the pipeline behaved, while validation explains whether the resulting data is trustworthy. A reliable data product needs both views connected by shared run and lineage context.

Quality should be defined as a contract, not a feeling

“This data looks wrong” is a useful incident report but a weak control. A data product should state which properties matter: required fields, unique keys, acceptable value ranges, relationship integrity, timeliness, or reconciliation against a source of record. Different datasets need different rules because quality is always tied to intended use.

The ownership question is central, which is why accountability for data quality matters as much as the checks themselves. Engineering can implement controls, but domain owners must define which failures are meaningful.

Contracts should also specify who can change the rules. If engineers quietly relax a null threshold to stop alerts, the metric may turn green while the business risk remains. Rule changes should be reviewed like code because they redefine what the organization accepts as trustworthy data.

Contracts also need tolerance rules. A unique key might allow zero violations, while an optional marketing attribute may tolerate a small null rate. Stating tolerance explicitly prevents teams from treating every metric as either perfect or useless and makes escalation consistent.

The contract should also identify the system of record for each field. When two sources disagree, a quality rule cannot decide correctness unless ownership and precedence are defined. Data observability becomes far more actionable when the team knows which upstream authority should be trusted.

Observability should connect runs to data outcomes

A pipeline run ID, notebook application ID, source partition, destination table, and business date should be connected wherever possible. That lineage lets an operator move from a bad number in a report back to the exact processing run that produced it.

Without that connection, teams maintain two disconnected histories: operational telemetry and data validation results. Incidents take longer because no one can quickly answer which technical event created the bad data.

Lineage context can be captured with batch identifiers and source versions even when full automated lineage is unavailable. A simple consistent identifier propagated through pipeline logs and validation results can dramatically reduce the time required to connect a bad output to its origin.

Correlation also enables automated incident grouping. If several failed quality rules and pipeline warnings share the same batch identifier, the platform can present them as one event instead of many independent alerts. That reduces noise and points responders toward a common cause.

Freshness is both a platform metric and a data-quality dimension

A pipeline may succeed on schedule while the source itself is stale. Conversely, a source can be current while a failed downstream job leaves consumers one day behind. Measuring only run completion misses both cases.

Useful freshness controls compare expected business time with observed data time. Maximum event timestamp, source extract time, target arrival time, and report refresh time can form a chain that shows exactly where delay entered the system.

Freshness objectives should reflect business deadlines. A dataset used for hourly fraud monitoring has a different acceptable delay from a monthly finance extract. The metric should encode the decision window that the data supports rather than an arbitrary universal threshold.

Freshness should be measured at multiple points when latency matters: source creation, ingestion, transformation completion, publication, and consumer refresh. That decomposition shows whether delay belongs to the producer, Fabric workflow, or reporting layer and keeps teams from optimizing the wrong stage.

Volume anomalies need business context

A 40 percent drop in records can indicate an ingestion failure, or it can reflect a real holiday, outage, or business event. Observability should flag the deviation; domain context decides whether it is an incident. Static thresholds are helpful, but historical and seasonal baselines are often stronger.

This is where data analytics becomes operational rather than descriptive. The same techniques used to understand trends can help define expected ranges for data production itself.

Volume baselines should be segmented when necessary. Overall row count can look normal while one region or product disappears completely. Critical dimensions should have their own coverage checks so aggregate volume does not hide localized loss.

Anomaly detection should be explainable enough for operators to trust it. A sophisticated model that flags volume changes without showing the baseline or contributing dimensions can create alert fatigue. Simple seasonal ranges with clear evidence may be more useful in many data products.

Schema checks protect structure before semantics

Unexpected columns, missing columns, type changes, or nullability drift can break consumers before a business-quality rule even runs. Schema validation should therefore be an early control, especially at ingestion boundaries where upstream systems change independently.

Schema success is still not semantic success. A field can retain the same type while changing units, codes, or meaning. That is why contracts need both technical shape and business definition.

Semantic checks can include referential integrity, domain codes, and reconciliation equations that express real business invariants. These rules often catch defects that generic profiling misses because the values are individually legal but inconsistent when considered together.

Referential checks should account for expected sequencing. A fact row can legitimately arrive before its dimension in some streaming or incremental systems. The rule may need a grace period or quarantine strategy rather than immediate failure. Quality logic should reflect how the pipeline actually delivers data.

Quality checks should fail at the right severity

Not every defect should stop the platform. A missing critical key might justify quarantine or pipeline failure, while a small increase in optional nulls may deserve an alert and continued processing. Treating every rule as fatal creates fragile systems; treating every rule as informational creates untrusted ones.

Rules should therefore have severity, ownership, and response. The team should know which failures block publication, which isolate bad records, and which create follow-up work without interrupting the main flow.

Severity policy should also determine notification routing. A warning about an optional attribute belongs with the data owner, while a failed financial reconciliation may need immediate operational escalation. Alerts are effective when they reach someone who can actually decide what happens next.

Quarantine paths need their own monitoring. Moving bad records aside is not resolution if no one reviews them. Teams should track quarantine volume, age, owner, and disposition so exceptions do not accumulate into an invisible secondary dataset.

Observability needs enough retention to explain delayed incidents

Data defects are not always discovered immediately. A financial reconciliation problem may surface days later, after default run-history windows or transient logs have become harder to access. Retention should reflect how long the business may need to investigate a published dataset.

General logging and monitoring principles apply: telemetry is useful only while it remains searchable, attributable, and connected to the incident timeline.

Retention needs should be tested against audit and incident requirements. If quality results are kept for a year but the pipeline logs disappear after a few weeks, older defects may be impossible to reconstruct. Related evidence should survive for compatible periods.

Audit investigations benefit from immutable or protected evidence. If operators can edit or delete the same logs used to prove what happened, the observability system has weak forensic value. Sensitive workloads may need centralized retention and restricted log administration.

Ownership should follow the boundary where the defect can be fixed

A source-system team owns invalid upstream values, a platform team owns a failed ingestion path, and a domain model owner may own a transformation that misclassifies records. Escalating every quality issue to one central data team creates a queue without improving accountability.

Quality rules should therefore point to the boundary that can actually correct the defect. The platform can route and surface evidence, but responsibility should stay close to the system that controls the cause.

Ownership can be encoded in catalog metadata, runbook references, alert routes, or item descriptions. The exact mechanism matters less than making it easy to determine who owns the source, transformation, and published data product when an incident begins.

Escalation ownership should include backups. A quality rule tied to one individual can become ineffective during leave or organizational change. Group-based ownership and documented service boundaries make response more resilient.

The end state is confidence with evidence

The Fabric data engineering discipline is mature when a team can prove not only that a run succeeded but that the published data met defined expectations. That requires operational telemetry, data validation, lineage, retention, and clear ownership working together.

The goal is not a dashboard with hundreds of green indicators. It is the ability to answer, quickly and defensibly, whether the data is complete, current, structurally valid, semantically plausible, and traceable to the processing that produced it.

Confidence should also include trend evidence. A dataset that passes every rule today but has steadily worsening lateness or null rates may still deserve attention. Observability is strongest when it shows direction, not just current status.

Trend review can be tied to service objectives. If null rates, lateness, or reconciliation differences worsen steadily while staying inside current thresholds, the team can act before a hard failure. Observability is valuable not only for incidents but for recognizing when a data product is slowly losing reliability.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!