Load-balancing incidents are difficult because the visible symptom—timeouts, uneven traffic, failed health checks, intermittent errors—can be caused by the load-balancing service, the backend, the network path, DNS, TLS, routing, or the application itself. Changing the load balancer before identifying the failing layer often makes the incident harder to understand.
The current AZ-700 blueprint includes application delivery services such as Azure Load Balancer, Application Gateway, Front Door, and Traffic Manager. A network engineer therefore needs a diagnostic model that distinguishes layer 4 distribution, layer 7 proxying, global routing, health evaluation, and backend behavior rather than treating every traffic-distribution problem as the same thing.
A clean troubleshooting path begins with expected behavior. What hostname should the client use? Which frontend should answer? Which rule should match? Which backend is eligible? What does healthy mean? What response should the client receive? Once those expectations are explicit, each observation can eliminate entire fault domains.
Identify the delivery layer before debugging it
Azure offers several services that influence how traffic reaches an application. A public load balancer forwards transport traffic, Application Gateway makes layer 7 decisions, Front Door operates globally at the application layer, and Traffic Manager influences DNS-based endpoint selection. The same user symptom can involve more than one service.
Start by drawing the actual request path from client to application. Include DNS, any global entry point, regional load balancer or proxy, network security controls, backend service, and return path. Troubleshooting becomes faster when every hop has a reason to exist.
Health probes are evidence about one path, not the whole application
A healthy probe proves that the probe request received an acceptable response under its own method, path, port, and source behavior. It does not prove that real users can authenticate, resolve dependencies, or complete business transactions. Conversely, a failed probe may reflect an intentionally protected endpoint rather than a dead application.
Verify the exact probe configuration and compare it with backend logs. If the probe path redirects, requires a host header, depends on a database, or is blocked by a security rule, the load balancer may correctly remove an otherwise functioning server. Treat probe status as a clue that must be explained.
Backend eligibility should be checked before routing changes
Determine which instances or endpoints are actually eligible to receive traffic. A backend may exist in the pool but be unhealthy, administratively drained, unable to listen on the expected port, or blocked from returning traffic. Load distribution cannot fix a backend that is not ready to serve requests.
The basic mechanics described in Exam-Labs’ load-balancing foundation become operationally useful when paired with backend evidence. Confirm listeners, health, local resource pressure, application logs, and dependency status before changing frontend rules.
Network security can create selective failure
NSGs, firewalls, user-defined routes, private endpoints, and asymmetric routing can allow the probe but block the client—or allow inbound traffic but break the return path. Selective failures are especially common when different subnets or source addresses take different policy paths.
Compare a successful flow and a failed flow. Source, destination, protocol, port, next hop, and security rule should be checked side by side. This narrows the problem more quickly than scanning a long configuration for anything that looks unusual.
DNS determines which load balancer users actually reach
A perfect regional configuration cannot help clients that resolve an old address, a private endpoint they cannot reach, or a global endpoint that sends them somewhere else. DNS caching can also keep a fixed problem looking broken for longer than expected.
The adjacent AZ-104 administration matters because application delivery depends on resource configuration outside the load-balancing service itself. Always record the hostname, resolved address, TTL, and DNS source during an incident.
TLS failures should be isolated from routing failures
When a proxy or gateway terminates TLS, certificate name, chain, protocol, backend encryption, and host-header behavior can all produce errors that look like connectivity problems. If the client reaches the frontend but the secure session or backend validation fails, changing routes will not help.
Capture where TLS is terminated and whether encryption is re-established to the backend. Check which certificate is presented, which hostname is used for validation, and whether the backend trusts the proxy path. The evidence should show whether the failure occurs before or after routing.
Distribution complaints need a traffic model
Users may report that one backend receives more traffic than another. That is not automatically a defect. Persistence, connection reuse, long-lived flows, health state, client distribution, and service-specific hashing behavior can all create uneven counts.
Before trying to force perfect balance, identify the outcome that matters: response time, capacity utilization, fault tolerance, or session correctness. A statistically uneven distribution can still meet the application objective, while a visually even distribution can hide a single slow backend.
Network Watcher and logs should test a hypothesis
Diagnostic tools are most valuable when used to answer a specific question. Connection troubleshoot, NSG diagnostics, flow logs, route inspection, backend health, proxy access logs, and application telemetry can generate large amounts of data; collecting all of it without a hypothesis slows analysis.
Use the broader Microsoft networking platform to compare expected and actual behavior at the suspected layer. If the hypothesis is asymmetric routing, inspect route and flow evidence. If the hypothesis is failed health evaluation, correlate probe results with backend logs.
Recovery is not complete until the reason is understood
A restart, rule change, or backend removal may restore service without proving the root cause. Close the incident only after the team can explain the failure sequence: which condition changed, which component reacted, why the symptom appeared, and why the remediation corrected it.
Then verify normal distribution, health, logs, security policy, and application behavior over enough time to catch recurrence. A clean troubleshooting path is not just about reaching green status quickly; it is about leaving the system better understood and less likely to fail the same way again.
A realistic incident often spans more than one delivery service. A global endpoint may direct traffic to a regional proxy, which then forwards to an internal load balancer and finally to application instances. The right troubleshooting strategy is to test each boundary with the smallest useful request. This avoids blaming the component that happens to be visible in the user-facing error while a failure several hops deeper is actually causing retries, resets, or unhealthy backend state.
Capacity problems should be distinguished from configuration problems. A system can be correctly configured yet fail during a traffic spike because backend connection limits, ephemeral ports, CPU, memory, or downstream services are saturated. Compare the failure window with resource metrics and request volume. If the error rate rises with load and disappears when demand falls, the architecture may need scaling or throttling changes rather than a new routing rule.
Deployment processes can also create transient imbalance. New instances may enter a backend pool before dependencies are warm, or health checks may remain green while caches and connection pools are cold. Safe rollout design should define when a backend becomes eligible, how long it is observed before receiving full traffic, and how it is drained during removal. These lifecycle details reduce the number of incidents that look like random load-balancer behavior.
The strongest post-incident report identifies the earliest observable signal that distinguished the failing state from normal operation. That might be probe latency, a route change, backend reset rate, TLS failure, or a specific application dependency. Turning that signal into monitoring or deployment validation is what prevents the troubleshooting exercise from becoming a one-time success with no lasting improvement.
Client diversity can expose delivery-layer assumptions that synthetic probes miss. Browsers, mobile clients, APIs, and internal services may use different protocols, connection reuse, headers, or DNS behavior. When only one client class fails, compare its request characteristics with a successful client before changing shared infrastructure. This can reveal issues such as host-header expectations, TLS negotiation differences, or persistence behavior that would otherwise be misdiagnosed as general load-balancer instability.
Stateful applications need an additional check: whether session affinity or connection persistence is part of correctness. A backend rotation that looks healthy at the network layer can disrupt users if requests in one workflow land on different instances without shared session state. The troubleshooting path should therefore ask whether persistence is intentional, where session data lives, and whether a scaling or deployment event changed that assumption.
The diagnostic model also helps during architecture reviews before incidents occur. For every critical application path, teams can document the expected frontend, probe behavior, backend eligibility, security path, DNS answer, and success response. That creates a baseline that operators can compare with live evidence when something breaks. It also exposes design gaps early, such as a probe that cannot distinguish partial failure or a backend pool with no safe draining process during maintenance.
This same model should be used in design reviews, not only outages. If teams cannot state how a request is selected, checked for health, routed, secured, and returned before production, they will have to learn that behavior under pressure when users are already affected.