Cloud troubleshooting requires one habit that on-premises administrators can forget: before changing local configuration, check whether the provider is already reporting a service problem. Microsoft 365 Service Health exists for exactly that reason. The current MS-102 study guide explicitly includes monitoring Microsoft 365 services by using Service Health, because tenant operations sit at the boundary between customer-managed configuration and Microsoft-managed infrastructure.
The interesting skill is diagnosis, not dashboard navigation. A user reports that Teams messages are delayed. Is the issue tenant-wide, regional, network-specific, identity-related, client-specific, or an active Microsoft incident? The first useful action is not to restart services you do not control. It is to establish scope and gather evidence.
A good investigation compares three views: what users experience, what the tenant telemetry shows, and what Microsoft reports about the service. Differences between those views narrow the fault domain. If Service Health shows an Exchange incident but only one department is affected, there may be an additional local condition. If Microsoft reports healthy service but multiple workloads fail after a Conditional Access change, the tenant likely owns the problem.
Define the symptom in user terms first
“Microsoft 365 is down” is not a diagnosis. Ask which workload, action, geography, user population, client, and time window are affected. Determine whether the problem is total failure, latency, intermittent errors, missing content, authentication, or degraded functionality.
Precise symptom definition prevents unrelated incidents from being grouped together. It also creates a baseline for verifying recovery later.
Symptom collection should include screenshots or exact error codes only when they add information. A generic screenshot of a failure dialog is less useful than knowing that the same action succeeds in Outlook on the web but fails in the desktop client. Capture evidence that separates layers rather than evidence that merely proves users are frustrated.
Check Service Health before making disruptive changes
The Microsoft 365 admin center reports active incidents and advisories affecting cloud services. If a known issue matches the symptom, local changes may be unnecessary and risky. Administrators can focus on communication, workarounds, and impact tracking instead.
This step should be built into runbooks. It is easy to skip during pressure, especially when users demand immediate action. A disciplined sequence protects the tenant from well-intentioned configuration churn during provider incidents.
Provider status pages can lag the first customer reports. A green Service Health page does not prove Microsoft is healthy, especially early in an incident. Treat it as one signal and continue comparing independent tenant telemetry, network tests, and user scope while avoiding disruptive changes until the evidence is stronger.
Compare provider scope with tenant scope
A Microsoft incident may affect only certain regions, features, or populations. Tenant symptoms may also be narrower because of licensing, configuration, routing, or client differences. Do not assume a provider advisory explains every observed problem.
Map affected users to workload, region, network path, device type, and policy scope. If unaffected users share most of those characteristics, the distinguishing factor becomes a valuable diagnostic clue.
Regional and licensing differences can create partial impact that looks inconsistent. Two users in the same office may hit different service features or policies. Build an affected-user matrix and look for the smallest characteristic shared by failures; that is often more useful than broad assumptions about geography.
Identity problems can masquerade as service outages
Authentication failures, token issues, Conditional Access changes, and directory problems can prevent users from reaching otherwise healthy services. When several applications fail at once, identity is a plausible shared dependency.
The related MD-102 path also highlights the endpoint dimension. A client update, certificate issue, device compliance change, or local network condition can produce symptoms that resemble a Microsoft 365 service incident.
Identity investigation should include token timing. A policy or role change may not appear immediately in every session, and cached credentials can create mixed outcomes. Compare fresh sign-ins with existing sessions before concluding that a control is applied inconsistently.
Network connectivity is part of tenant health
Cloud applications rely on DNS, internet routing, proxies, inspection, VPNs, and endpoint connectivity. A corporate network issue can affect many Microsoft 365 workloads simultaneously while Service Health remains green.
Compare affected and unaffected networks, test direct paths where policy permits, review proxy and inspection changes, and use available connectivity insights. Do not equate “many users” with “provider outage” if those users share the same network dependency.
Network isolation is stronger when tested from a known-good alternate path. If permitted, compare corporate routing with an external connection for a controlled account. A difference does not automatically prove the corporate network is at fault, but it can quickly narrow where further evidence should be collected.
Change history often explains sudden tenant-specific failures
If Service Health is normal, ask what changed. New security policy, DNS configuration, license assignment, app consent rule, endpoint baseline, or network inspection can create broad symptoms. Recent change does not prove causation, but it narrows the hypothesis space.
Maintain records for high-impact configuration. During an incident, operators should be able to identify candidate changes quickly and know whether rollback is safe.
Change review should consider vendor-side feature rollout as well as local configuration. Microsoft may introduce behavior gradually across tenants. When symptoms appear without a local change, check message-center communications and release notes in addition to formal service incidents.
Communication is an operational control
Users and leaders need accurate status even when root cause is still uncertain. State what is known, what is affected, which workaround exists, and when the next update will occur. Avoid claiming a Microsoft outage until evidence supports it.
Consistent communication reduces duplicate tickets and prevents local teams from improvising conflicting workarounds. It also creates a record that can be reviewed after recovery.
Communication templates should distinguish confirmed cause from working hypothesis. Saying “Microsoft is investigating” when that is verified is different from saying “we are investigating reports that may be related to Microsoft 365.” Precise language preserves trust during uncertain incidents.
Recovery needs verification across representative workflows
A provider status change to “resolved” is not the end of the incident. Confirm sign-in, message delivery, file access, meetings, or other representative user actions. Check whether backlogs remain and whether local workarounds should be reversed.
Measure recovery from the user perspective. Some cloud incidents have delayed effects, caches, queued messages, or client state that take additional time to normalize.
Recovery validation should sample different user populations and clients. A service may recover first for web access while desktop clients still hold stale state, or one region may normalize before another. Close the incident only after the representative business workflows meet the expected baseline.
Service Health remains relevant after MS-102
MS-102 retires on November 30, 2026, but AB-650 continues to include Microsoft 365 service health and extends monitoring toward AI services and Copilot. The operating principle survives the exam transition: administrators need to distinguish service-provider failure from tenant, identity, endpoint, and network failure.
That distinction is part of mature administration across the Microsoft ecosystem. The best operator changes the system only after evidence has made the fault domain small enough that the change is justified.
Operational maturity improves when post-incident reviews compare internal detection time with provider notification time. If users reported impact long before the team noticed it, strengthen telemetry and ownership. If Service Health was available but ignored, update the runbook and training rather than blaming individual responders.
Service-health diagnosis should also include dependency awareness. A Teams symptom may ultimately depend on Exchange, Entra, SharePoint, networking, or client services. When Microsoft publishes an incident for one component, map which user workflows inherit that dependency before assuming unrelated workloads should fail in the same way.
Escalation quality improves when the team preserves a concise evidence package: affected users, timestamps, correlation IDs when available, network path, policy state, screenshots of relevant health notices, and results from known-good tests. That package reduces repeated data gathering when a case moves between help desk, internal engineering, and Microsoft support.
Operational teams should periodically rehearse cloud-service incidents. A tabletop exercise can test who checks Service Health, who communicates with executives, which critical business processes have alternatives, and how recovery is validated. Rehearsal exposes ownership gaps without waiting for a real outage to reveal them.
After major incidents, update monitoring thresholds and runbooks with what actually proved useful. If user reports consistently beat technical alerts, improve telemetry. If network tests were misleading, refine them. The troubleshooting process should learn from evidence just as the configuration does.
Service incidents can also expose hidden single points in business process. If one Microsoft 365 workload becomes unavailable, teams should know which critical activities depend on it and which temporary alternatives are approved. Recording those dependencies turns outage response into continuity planning rather than improvised technical troubleshooting.
Finally, incident records should preserve the exact start and recovery windows observed by the business. Provider timestamps, user reports, and local telemetry may differ. Reconciling them helps calculate real impact, improve alerting thresholds, and distinguish a short provider incident from a longer tenant-specific recovery problem.