Ansible provides a practical way to make Cisco IOS XE configuration repeatable, reviewable, and fleet-oriented. The current Ansible cisco.ios collection includes configuration, interface, routing, ACL, BGP, HSRP, LACP, and other resource modules, while ios_config remains useful for deterministic configuration sections not covered cleanly by a resource model. The collection is separate from ansible-core when it is not included through a broader Ansible distribution.
Within Cisco Network Engineering, the engineering goal is not “replace CLI with YAML.” It is to move desired state, validation, credentials, change review, and rollback into a controlled automation workflow.
The existing network automation article provides the right starting principle: automate high-volume, well-understood, measurable changes before attempting fully autonomous remediation.
Inventory should represent real management boundaries
Ansible inventory can group devices by site, platform, function, software train, environment, or change cohort and attach group/host variables that drive playbook behavior.
A giant flat list with per-host copy/pasted variables scales poorly and hides which settings are intended to be common.
Use stable metadata and dynamic inventory where appropriate, but keep the source of truth clear enough that operators can explain why one device received one variable value.
Use the current Cisco IOS collection explicitly
Modern playbooks should use fully qualified collection names such as cisco.ios.ios_config, cisco.ios.ios_interfaces, or cisco.ios.ios_bgp_global rather than relying on historical module aliases.
The current collection documents supported ansible-core versions and module behavior separately from ansible-core itself.
Pin collection versions for production automation and test upgrades, because module arguments, parsing, and feature coverage can evolve independently of your playbooks.
Resource modules are preferable for structured configuration
Resource modules model IOS configuration as structured data and support states such as merged, replaced, overridden, gathered, parsed, or rendered depending on the module.
This is safer than templating raw CLI when the resource is well supported because the module understands hierarchy and can compare intended versus current state.
Use ios_config for gaps or specialized sections, but avoid making every playbook one large string template when structured modules exist.
Idempotency depends on canonical command form
Network automation is useful when re-running the same playbook does not generate unnecessary changes.
Ansible documentation warns that abbreviated IOS commands and mismatched indentation/command representation can cause configuration modules to report changes repeatedly.
Write intended lines in the form IOS XE returns them, test repeated runs, and use diff mode so reviewers see whether the playbook converges cleanly.
Connection and privilege escalation should be standardized
The current IOS platform guidance commonly uses ansible.netcommon.network_cli with ansible_network_os: cisco.ios.ios for CLI automation and supports enable-mode privilege escalation.
SSH keys, vault-backed passwords, bastion configuration, and service accounts should follow enterprise credential policy.
Do not store production enable passwords or device credentials in inventory files committed to source control.
Check and diff should be part of change review
Use supported check/diff workflows, rendered states, or predeployment parsing to show what will change before pushing to a fleet.
For ios_config, diff options can compare running/startup/intended state according to module capabilities.
Automation should make review better than manual CLI, not remove the review step. A pull request should reveal the intended network behavior and affected device cohort clearly.
Backups are useful but not a rollback strategy by themselves
Ansible can capture running configuration before change.
That backup is valuable evidence, but blindly replacing the full running configuration can be unsafe if device state or unrelated changes moved forward after the snapshot.
Design rollback as the inverse desired-state change where possible, and know which changes require a full configuration restore, reload, or out-of-band recovery.
Validation should test network state, not only task success
A task reporting changed: true means Ansible sent configuration successfully; it does not prove BGP adjacencies remained healthy, interfaces are up, routing converged, or users can reach an application.
Add postchecks using ios_command, facts/resource gathering, APIs, telemetry, IP SLA, or external assurance tests.
The existing network assurance and telemetry article provides a useful framework for choosing evidence.
Roll out by cohort and control concurrency
Do not change every access switch or edge router simultaneously simply because Ansible can.
Use serial/batch execution, maintenance cohorts, canary devices, and failure thresholds so one bad variable or platform-specific command does not create a fleet outage.
Automation increases speed; staged rollout limits the blast radius of wrong intent.
Templates and variables should not hide complex branching
Jinja can generate large device configurations, but deeply nested conditionals become hard to test and review.
Prefer clear data models and role/task composition, with separate playbooks or roles when two device classes genuinely require different behavior.
If one template contains dozens of platform/version exceptions, the abstraction may be obscuring several distinct automation products.
Ansible is successful when automation becomes a safer change process
The mature workflow has source-controlled inventory/data, pinned collections, structured modules, secure credentials, deterministic diffs, staged deployment, state validation, and tested rollback.
The result is not fewer network engineers. It is fewer one-off CLI variations and more time spent designing network intent and verifying that the automated change produced the expected service outcome.
Data modeling should separate business intent from device syntax. A branch definition might say primary WAN, backup WAN, management VLAN, and BGP policy, while roles convert that data into IOS XE resource-module inputs. When inventory stores raw CLI fragments, changing syntax across platform generations becomes much harder because business data and implementation are fused together.
Use facts and gathered resource state carefully. Gathering every possible fact from hundreds of devices can be slow and expensive; collect only the state needed for validation or drift checks. For high-frequency operational telemetry, streaming or controller APIs may be more appropriate than repeatedly launching Ansible jobs just to read counters.
Secrets management should include device credentials, vault passwords, API tokens, and private keys used by jump hosts or PKI. Integrate Ansible Vault or an external secret manager and use short-lived credentials where the network platform supports them. Automation service accounts should have only the CLI/API privileges required for the playbooks they run.
Prechecks should reject incompatible devices before configuration starts. Validate platform family, IOS XE version, license/feature availability, available memory/storage, current HA state, and configuration sync where relevant. A playbook can be syntactically correct and still push a command unsupported on one older switch in the batch.
Network state can change between precheck and apply. For sensitive changes, keep rollout batches small and rerun critical state checks immediately before committing the device. Routing adjacency, interface state, or maintenance status may have changed since the pipeline first built its plan.
Use handlers or explicit save logic deliberately. Some organizations save configuration only after all validation succeeds; others save after each successful change. Understand how save_when, startup configuration, and rollback interact so a reload does not either erase a valid production change or persist a failed experiment unexpectedly.
Test playbooks against lab/CML or representative canary devices before broad deployment. Unit testing Jinja/data transformations catches schema errors, while integration testing catches IOS parser behavior, feature dependencies, and sequencing problems. The same test data should include edge cases such as empty lists, dual-stack interfaces, shutdown ports, and preexisting configuration.
Automation ownership should be split clearly between platform and network-domain teams. A central automation group can maintain CI, credentials, collection versions, linting, and shared roles, while routing/campus engineers own the intent and validation rules for their domain. This avoids a platform team becoming responsible for network policy it cannot judge.
Drift detection should be separated from enforcement. A scheduled playbook can gather key configuration and report deviations without automatically pushing every difference back to intended state. Some drift is an emergency local fix or a legitimate temporary exception. Report first, then decide whether to reconcile automatically or through reviewed change.
Configuration generation should be deterministic across runs. Sort data where order is not semantically meaningful, normalize interface names, avoid timestamps inside generated config, and keep host/group variable precedence understandable. Nondeterministic templates create noisy diffs that hide the few lines reviewers actually need to evaluate.
Observability around automation matters too. Record playbook version, inventory source, collection version, operator/service account, start/end time, affected devices, changed tasks, validation outcome, and any failed hosts. This turns automation execution into an auditable network release instead of a terminal command that disappears after the run.
Finally, define what Ansible should not own. Controller-managed fabrics, intent systems, cloud-managed platforms, or APIs with their own transaction model may be better automated through supported controllers rather than forcing CLI configuration through Ansible. Use Ansible as an orchestrator where it adds control, not as a universal hammer.
Network automation should also have a decommissioning path. When a device leaves service, remove it from dynamic inventory, credentials, monitoring, and automation groups so future playbooks cannot target an address later reused by another asset. Source-of-truth hygiene is part of automation safety because stale inventory can turn a perfectly correct playbook into a change against the wrong device.
Keep inventory current and owned.
Treat every playbook run as a network change with evidence. Diff the intended configuration, limit the device scope, validate after execution, and preserve enough output to explain what changed. Idempotence matters, but it does not replace change control or post-checks.