Cisco 350-401: IP SLA Tracking

Cisco IP SLA performs active measurements from the network device and can feed enhanced object tracking so routing, first-hop redundancy, event automation, or monitoring reacts to measured service state. The core design decision is what to probe. Pinging the next-hop interface proves local reachability; probing a remote application dependency can prove much more about the path users actually need.

Within Cisco Network Engineering, IP SLA tracking connects observability to control. The existing IP SLA alerts article focuses on alerting; this page focuses on object tracking and routing decisions.

Current IOS XE enhanced object tracking distinguishes IP SLA state from reachability. An operation returning OverThreshold can still count as reachable even though performance is degraded, while state tracking may report down because the return code is not OK.

Choose the probe that represents the failure you want to detect

ICMP echo is simple and widely supported, but IP SLA can measure other operation types depending on platform/software, including UDP jitter, TCP connect, HTTP, DNS, and application-related probes.

Use a destination beyond the local gateway when the failover objective is internet or service reachability rather than interface status.

At the same time, avoid probing an unstable third-party endpoint whose maintenance would trigger unnecessary route changes.

Source interface and VRF determine what path is actually tested

An IP SLA operation should originate from the same routing context and, where relevant, source interface/address as the traffic the track object is intended to protect.

In multi-VRF devices, a probe in the global routing table can remain healthy while the customer VRF is broken.

Document source, destination, VRF, frequency, and expected path so operators know what the measurement proves.

State and reachability tracking have different semantics

IOS XE object tracking can track an IP SLA operation by state or reachability.

State typically expects an OK return code. Reachability can remain up for OK or OverThreshold outcomes, which is useful when performance is poor but the route should not be withdrawn automatically.

Select the mode based on the control objective; do not let a latency threshold accidentally become a hard routing failure if that was not intended.

Delay timers prevent transient flaps from becoming routing churn

Enhanced object tracking supports up/down delay timers so one missed probe does not immediately remove a route or change HSRP priority.

Set timers according to probe frequency, application tolerance, and expected transient loss.

Too little delay creates oscillation; too much delay leaves users on a dead path after the failure is already obvious.

Floating static routes are a common control client

A tracked primary static route can be removed when the IP SLA object goes down, allowing a higher-administrative-distance backup route to become active.

This is useful at small branches or internet edges where dynamic routing with the provider is unavailable.

The backup path should be tested for NAT, firewall policy, DNS, MTU, and capacity; route installation alone does not prove the application works.

FHRP can consume tracked state

HSRP and related first-hop redundancy designs can decrement priority when a tracked object fails, causing the peer with better upstream reachability to become active.

This solves the classic problem where a gateway interface is locally up but its upstream path is broken.

Coordinate tracking on both peers so they do not oscillate or both prefer themselves under asymmetric failures.

IP SLA can drive EEM or operational automation

Tracked-state transitions can feed Embedded Event Manager or monitoring systems to run diagnostics, send alerts, or perform bounded remediation.

Automation should remain conservative when the same IP SLA also controls routing; a probe failure can already change forwarding, and a second automation action may amplify the effect.

Use event correlation and clear ownership for any remediation beyond the routing change itself.

Probe targets should be redundant and meaningful

One probe to one DNS server or public IP creates a false dependency on that target.

For high-value decisions, use track lists/boolean logic or multiple independent measurements where supported so one target failure does not withdraw an otherwise healthy path.

The logic should still be simple enough to troubleshoot. Complex boolean combinations can obscure why the object changed state.

Monitor the measurement system itself

Use show ip sla, show track, operation history/statistics, timestamps, and logging to understand last result, return code, threshold behavior, and transition time.

A route that disappeared because the probe never started is different from a route removed because the destination became unreachable.

Configuration and scheduling state are therefore part of IP SLA health.

Test failover and failback independently

Remove or impair the primary path and measure detection, track transition, routing/FHRP change, application recovery, and log evidence.

Then restore the path and measure how long the network waits before returning. A quick failover paired with aggressive failback can create repeated disruption when a provider circuit flaps.

Delay and hysteresis decisions should reflect actual user recovery, not only router timers.

IP SLA tracking is successful when the measured condition matches business reachability

The mature design can state exactly what the probe proves, how the track interprets the result, which client reacts, what delay applies, which backup path takes over, and how operators verify the state.

Active measurement becomes valuable when it removes traffic from a path users would actually consider failed—not merely when one interface happens to be down.

Probe frequency, timeout, and threshold must be coherent. A probe every five seconds with a ten-second timeout can overlap and create confusing transitions; a one-second probe across hundreds of branches can create unnecessary CPU and WAN traffic. Design cadence from the detection objective and scale, then confirm the platform can schedule the full set reliably.

DNS-based application checks introduce another dependency. If the probe targets a hostname and DNS fails, the SLA may report the application unavailable even when the remote service IP is reachable. That may be correct if users also require DNS, but the runbook should state which dependency the probe represents so teams do not “fix routing” for a DNS incident.

Return-path asymmetry should be considered. An ICMP echo may leave over the primary link and return over another provider or firewall path, producing a measurement different from the application flow. Source address, routing, NAT, and policy should mimic the protected service closely enough that the SLA result is meaningful.

Track lists can express composite health such as “two of three upstream targets reachable” or “interface up AND remote probe reachable.” Keep the logic small and document it. Complex boolean objects can become difficult to interpret during failure if operators cannot quickly identify which child object drove the result.

IP SLA responder features can improve accuracy for UDP jitter or other measurements between Cisco devices by timestamping/responding in a controlled way. This is useful for voice/path-quality baselines when raw ICMP is insufficient. Treat responders as production services with ACL, VRF, and availability requirements.

Static-route tracking should avoid recursive surprises. The tracked probe’s own destination route must not depend on the route being removed in a way that creates oscillation or sends the probe over the backup path and immediately declares the primary healthy again. Use source/interface/host routes or topology design that preserves correct test semantics.

Monitoring systems should receive both the raw IP SLA result and the resulting track/client action. This lets operators distinguish “latency exceeded threshold,” “track went down,” and “floating route installed” as separate timeline events. Correlation shortens root cause analysis and supports tuning of delays.

Failover exercises should include partial impairment such as high latency or loss, not only link shutdown. The selected state/reachability tracking mode should react exactly as intended. A route designed to fail only on complete loss should not flap because the probe crossed a performance threshold during temporary congestion.

Thresholds should not be confused with reachability unless the service objective truly requires it. A DNS response rising from 40 ms to 200 ms may violate a performance target but still be preferable to a backup path with 500 ms latency. Separate alert thresholds from routing-withdrawal conditions when service degradation and service failure need different responses.

VRRP/HSRP and static-route consumers should have clear precedence during multi-failure scenarios. If an IP SLA changes first-hop role while a separate dynamic-routing event also removes prefixes, the combined outcome can differ from either mechanism tested alone. Failure testing should therefore include compound events such as upstream loss on the current active gateway.

Use a neutral target when measuring provider reachability. Probing only the provider next hop can miss upstream internet failure; probing only one SaaS service can withdraw an otherwise healthy provider because that SaaS had its own outage. Multiple carefully chosen targets or a provider-operated SLA endpoint can produce a more defensible signal.

Object-tracking state should be logged with the route or FHRP action it controls. A downstream monitoring platform that sees only the track transition may not know which static route was removed or which HSRP priority changed. Correlated event records make failover explainable and simplify threshold tuning after a real incident.

Tracking should follow an objective that represents the service, not merely an interface. Probes to a meaningful remote dependency, combined with realistic source and routing context, provide a stronger basis for failover than checking whether the next-hop device still answers.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!