gNMI—gRPC Network Management Interface—is a standards-based API used by modern network platforms to retrieve, subscribe to, and in some implementations modify YANG-modeled data. In current Cisco IOS XE model-driven telemetry, applications can subscribe to specific YANG paths through supported programmable interfaces, including gNMI, and receive structured data periodically or on-change rather than repeatedly scraping CLI output.
Within Cisco Network Engineering, gNMI telemetry is the observability interface between device state and collectors, assurance platforms, or automation systems. The existing network assurance and telemetry article provides the signal-quality perspective; this page focuses on subscription engineering.
Current IOS XE documentation distinguishes dynamic dial-in subscriptions from configured dial-out subscriptions and notes that exact model, subscription, receiver, and scaling support is platform dependent.
YANG paths define the data contract
gNMI requests reference structured YANG-modeled paths rather than screen-scraped CLI strings.
This makes field hierarchy and data type explicit, which is valuable for automation and analytics.
Collectors should record the model/module and software version because paths can be added, deprecated, or behave differently across IOS XE trains and platform families.
OpenConfig and Cisco-native models have different portability trade-offs
OpenConfig models aim for cross-vendor consistency, while Cisco-native models often expose deeper platform-specific state or features.
Choose OpenConfig when portability covers the operational requirement; use native models where important device data is not represented adequately.
A telemetry architecture can support both, but dashboards should not assume fields from different models have identical semantics.
Dial-in subscriptions are client initiated
In a dynamic subscription, the collector connects to the device and requests a path, mode, and sampling behavior.
This centralizes subscription lifecycle in the collector and is convenient for ad hoc or collector-driven monitoring.
Network reachability, authentication, TLS, and per-device connection scaling become collector responsibilities.
Dial-out subscriptions are device initiated
Configured subscriptions cause the IOS XE publisher to initiate the telemetry connection toward a receiver.
This can fit environments where collectors should not open management sessions inbound to every device.
Device configuration must include receiver details and subscription parameters, so source-of-truth/automation should manage them consistently across the fleet.
Periodic and on-change modes solve different signal problems
Periodic sampling is appropriate for counters, utilization, environmental metrics, and data where regular time series are useful.
On-change is efficient for state that changes infrequently, such as interface operational status or certain configuration/state leaves, when the platform/model supports on-change semantics.
Not every data node supports meaningful on-change behavior. Validate the platform documentation rather than applying one subscription mode to all paths.
Sampling interval should match decision latency
Faster telemetry is not automatically better. A 100-millisecond stream across thousands of devices can overwhelm devices, network, brokers, and time-series databases while providing little operational value for slowly changing metrics.
Start from the question: how quickly must the monitoring/automation detect this condition?
Then select the coarsest interval that meets the SLO and test CPU/bandwidth impact at realistic scale.
TLS and authentication are part of the telemetry design
gNMI normally runs over gRPC and can use TLS with server/client authentication depending on platform configuration.
Certificate lifecycle, trust roots, username/credential or certificate authorization, and management-VRF reachability should be automated like any other production service.
A collector outage caused by expired telemetry certificates can remove visibility across a fleet without affecting forwarding, which makes monitoring of the monitoring path important.
Timestamp quality and device clock matter
Telemetry pipelines need consistent event/sample timestamps for correlation across devices and other systems.
Maintain NTP/PTP/time health and understand whether a metric timestamp is generated at the device, transport layer, collector, or database.
Late or out-of-order samples should be distinguishable from actual network-state changes.
Collector architecture should expect bursts and disconnects
Large fleets can reconnect simultaneously after a collector, network, or certificate outage.
Use load-balanced collectors, buffering/message brokers where appropriate, backpressure, and capacity sized for reconnection bursts rather than only steady-state sample rate.
Device subscription limits and receiver behavior are platform dependent; do not design one collector fan-out ratio without testing the actual IOS XE platforms.
Telemetry should feed operational questions, not one giant data lake
Useful examples include interface error/drop rates, BGP neighbor changes, route counts, environmental sensors, queue drops, CPU/memory, optics, or policy state.
Collect paths tied to defined dashboards, alerts, baselines, or automation decisions and document ownership.
The existing data-center telemetry troubleshooting article reinforces the value of layer-specific signals.
gNMI telemetry is successful when structured state becomes actionable
The mature deployment can explain which YANG path is collected, from which devices/VRFs, at what cadence/mode, through which authenticated transport, into which collector, with what freshness and retention.
Streaming telemetry should make device state easier to correlate and automate—not simply produce more data than operators can trust or use.
Path discovery should be automated rather than hard-coded from memory. Devices expose supported YANG models and capabilities, and collector pipelines should verify that a requested path exists before rolling a subscription across a mixed software fleet. This reduces silent data gaps after one platform or software train lacks the expected node.
Encoding choice matters. gNMI commonly uses structured protobuf-based messages with typed values, but collectors may translate data into JSON, Prometheus labels, OpenTelemetry, or time-series schemas. Preserve enough original path/type metadata that downstream users can distinguish counters, gauges, enums, and strings accurately.
Counter reset behavior should be modeled. Interface and hardware counters can reset after reload, process restart, clear command, or line-card event. Rate calculations need device boot/session context so a counter reset does not appear as a huge negative or positive traffic spike.
On-change telemetry can generate bursts during topology transitions. A routing flap, switch stack failover, or interface storm can cause many subscribed leaves to change simultaneously. Collectors and message brokers should be sized for event storms, not only quiet-state on-change volume.
Access control deserves careful review because some IOS XE telemetry interfaces/platforms can expose broad operational state once a subscriber is authenticated. Use dedicated read-only identities, management-plane ACLs, TLS, and collector segmentation. Observability credentials should not also have configuration privileges unless the same service explicitly needs them.
Schema normalization should not erase device context. If a cross-vendor dashboard maps several models into one “interface state” metric, retain vendor/platform/software labels so anomalies can be traced back to the original semantics. Standardization is useful only when it does not make unsupported comparisons look identical.
Telemetry pipelines should include freshness and completeness metrics for themselves: active subscriptions, last sample time, dropped messages, collector queue depth, decode errors, and device connection failures. A dashboard showing flat CPU because the subscription died is more dangerous than an obvious “no data” state.
Configuration automation and telemetry can close the loop when used carefully. Ansible or controllers can make a change, gNMI can stream the resulting operational state, and validation logic can compare that state with the expected outcome. Keep remediation bounded; a noisy telemetry threshold should not automatically push configuration across a fleet without evidence and guardrails.
Subscription ownership should be centralized enough to avoid duplicate collection. Several teams independently subscribing to the same high-frequency path on every device can multiply CPU, bandwidth, and collector load without adding information. A telemetry catalog can record path, cadence, consumers, retention, and business purpose so data is reused rather than recollected.
High-cardinality labels can make telemetry storage expensive. Interface names are manageable; per-route, per-MAC, per-client, or per-flow dimensions can explode series counts in time-series databases. Estimate cardinality before streaming large tables and consider event/log or on-demand APIs for data that does not belong in a dense metric store.
Backfills and historical queries are a data-platform concern, not a device concern. gNMI streams current state; the collector/database must decide how long to retain it, downsample old data, and preserve events around incidents. Retention should match troubleshooting and capacity-planning needs rather than keeping every 100-ms sample forever.
Change control should include telemetry schema dependencies. A software upgrade can add or change YANG behavior while dashboards and alerts still expect the previous path. Validate critical subscriptions on canary devices before fleet upgrades and alert when a path stops decoding instead of silently dropping the metric.
Telemetry and CLI should be cross-validated during rollout. Pick representative counters/state, compare gNMI values with show commands or other supported interfaces, and verify units and reset behavior. This establishes trust in the pipeline before operations relies on it for automated decisions.
Collector failover should be tested explicitly. If only one receiver can be attempted for a given configured subscription on a platform, resilience may need multiple subscriptions, DNS/load-balancing, or collector architecture rather than assuming the device will automatically try a second destination. Design around documented receiver behavior for the exact IOS XE release.
Telemetry consumers should publish data-quality status beside the metric. A chart should be able to show stale, missing, decoded-with-error, or partially subscribed states instead of drawing a flat line that looks normal. This is especially important when operations uses streamed state as an input to automated remediation.
Keep subscription health visible beside every critical metric.
Telemetry design should also define retention, cardinality, and ownership before collection expands. High-frequency structured data becomes expensive and hard to interpret if teams do not know which paths answer operational questions, which values need alerting, and which can remain on-demand.