Network Automation Engineer Roadmap
A path from configuring switches by hand to running a network as code — Python, structured device APIs, Ansible, a source of truth, automated testing, and telemetry that closes the loop.
Network automation is usually sold as "learn Python and Ansible". That framing produces engineers who can push a config to fifty devices and cannot tell you whether they should have.
The order here reflects the real constraint: the hard part is not pushing configuration, it is knowing what the configuration should be. That is why version control comes before scripting, and why the source of truth arrives before the tooling that scales. A fast automation reading bad data is a fast way to break a network.
The phase most people skip is testing. It is also the one that separates a scripter from an engineer, because it is what makes a change reversible before it is made rather than after.
New to Linux and the command line?
This path assumes fundamentals you may not have yet. Our Foundations Pack is out and free — Linux, the shell and Git, with exercises that mark your work and explain why you got it wrong. We're writing an agents pack next; leave your email if you want to hear when it ships.
One email when the pack launches. No spam, unsubscribe any time.
The path, phase by phase
Version Control Before Automation
Put every configuration you own into Git before writing a line of automation. Learn branches, diffs, pull requests and rollback with real device configs as the payload. Done when you can point at a commit and say what changed on which device, when, and who approved it — without opening a ticket system.
3 weeks7 SkillsGit branching and mergingReading and writing diffsPull request review workflowConfiguration backup automationSemantic commit historySecrets handling in repositoriesMarkdown documentationShow details, projects and resourcesSkills you'll master
Git branching and mergingbeginnerReading and writing diffsbeginnerPull request review workflowbeginnerConfiguration backup automationbeginnerSemantic commit historybeginnerSecrets handling in repositoriesintermediateMarkdown documentationbeginnerHands-on projects
- 01Back up the running configuration of every device you own into a Git repository on a nightly schedule
- 02Reconstruct what changed on a device over the last month using only the commit history
- 03Write a pull request template that forces a rollback plan for every network change
- 04Set up pre-commit hooks that reject a commit containing a plaintext password
- 05Document your topology as a file in the same repository as the configs it describes
Python for Network Engineers
Enough Python to parse, transform and generate — not enough to build a web app. Data structures, files, error handling, virtual environments and the standard library. Done when you can turn a directory of show-command output into a CSV that answers a question your manager asked.
6 weeks7 SkillsPython data structures and comprehensionsWorking with JSON and YAMLRegular expressions for text parsingVirtual environments and dependency pinningError handling and retriesWriting and running unit testsReading library documentationShow details, projects and resourcesSkills you'll master
Python data structures and comprehensionsbeginnerWorking with JSON and YAMLbeginnerRegular expressions for text parsingintermediateVirtual environments and dependency pinningbeginnerError handling and retriesintermediateWriting and running unit testsintermediateReading library documentationbeginnerHands-on projects
- 01Parse the output of "show ip interface brief" from twenty devices into a single structured report
- 02Build an inventory script that flags every interface that has been down for more than 30 days
- 03Write a script that compares two configuration files and reports only semantic differences
- 04Generate a per-site VLAN allocation table from a YAML definition file
- 05Package one of your scripts with a pinned requirements file so a colleague can run it unchanged
Talking to Devices Without the CLI
Move from screen-scraping to structured interfaces: NETCONF, RESTCONF, gNMI and the YANG models underneath them. Learn where each is supported and where it is not, because vendor coverage is uneven and that is a planning constraint. Done when you can retrieve interface state as structured data from two different vendors.
5 weeks7 SkillsNETCONF and the XML data it returnsRESTCONF over HTTPYANG data modelsgNMI subscriptionsNetmiko for devices with no APINAPALM multi-vendor abstractionCertificate and credential management for device accessShow details, projects and resourcesSkills you'll master
NETCONF and the XML data it returnsintermediateRESTCONF over HTTPintermediateYANG data modelsintermediategNMI subscriptionsadvancedNetmiko for devices with no APIintermediateNAPALM multi-vendor abstractionintermediateCertificate and credential management for device accessintermediateHands-on projects
- 01Retrieve the full interface state from a device over NETCONF and store it as JSON
- 02Write the same query twice, once with RESTCONF and once by parsing CLI output, and compare reliability
- 03Use NAPALM to collect facts from three devices of different vendors with one script
- 04Subscribe to a gNMI stream and record how quickly a link-down event arrives compared with SNMP polling
- 05Map which of your device models support NETCONF, and which are stuck on SSH, into a table you can plan from
Configuration Management with Ansible
The declarative layer most teams standardise on. Inventories, variables, templates, idempotence and check mode against network modules rather than servers. Done when you can push a change to fifty devices, run it twice, and have the second run report zero changes.
6 weeks7 SkillsAnsible inventories and group variablesJinja2 configuration templatingIdempotence and check modeNetwork collections and connection pluginsAnsible Vault for credentialsRoles and reusable structureLimiting blast radius with serial and limitShow details, projects and resourcesSkills you'll master
Ansible inventories and group variablesintermediateJinja2 configuration templatingintermediateIdempotence and check modeintermediateNetwork collections and connection pluginsintermediateAnsible Vault for credentialsintermediateRoles and reusable structureintermediateLimiting blast radius with serial and limitadvancedHands-on projects
- 01Template the full configuration of one device family from variables and render it for every site
- 02Deploy an NTP and syslog change to your whole estate with check mode first, then for real
- 03Write a role that is safe to run repeatedly and prove it with a second run that reports no changes
- 04Convert one manual change runbook into a playbook and delete the runbook
- 05Roll a change out in batches of five devices with an automatic stop on first failure
A Source of Truth Worth Trusting
Automation is only as good as the data it reads. Model your network in NetBox, define what is authoritative there rather than on the device, and reconcile the difference. Done when a device's configuration is generated from the source of truth and drift is reported rather than discovered.
6 weeks7 SkillsNetBox data model and object relationshipsIPAM and prefix managementDefining authority — device versus databasePopulating a source of truth from existing devicesDrift detection and reconciliationDynamic inventory from an APIData validation and schema enforcementShow details, projects and resourcesSkills you'll master
NetBox data model and object relationshipsintermediateIPAM and prefix managementintermediateDefining authority — device versus databaseadvancedPopulating a source of truth from existing devicesintermediateDrift detection and reconciliationadvancedDynamic inventory from an APIintermediateData validation and schema enforcementintermediateHands-on projects
- 01Model one site completely in NetBox — racks, devices, interfaces, cables, prefixes
- 02Import your existing IP allocations and find the conflicts the spreadsheet was hiding
- 03Drive an Ansible inventory dynamically from NetBox instead of a static file
- 04Generate a device configuration entirely from NetBox data with no per-device file
- 05Build a nightly job that reports every device whose running config differs from the generated one
A Lab You Can Destroy
You cannot test network automation against production, and a physical lab does not scale to every change. Build reproducible virtual topologies with containerlab or GNS3 and treat the topology file as code. Done when you can spin up a copy of a production site from a file in under ten minutes.
4 weeks6 Skillscontainerlab topologiesContainerised network operating systemsGNS3 or EVE-NG for image-based labsReproducible environment definitionTraffic generation and verificationResource planning for lab hostsShow details, projects and resourcesSkills you'll master
containerlab topologiesintermediateContainerised network operating systemsintermediateGNS3 or EVE-NG for image-based labsintermediateReproducible environment definitionintermediateTraffic generation and verificationintermediateResource planning for lab hostsbeginnerHands-on projects
- 01Define a three-tier topology as a containerlab file and bring it up from scratch
- 02Reproduce a production incident in the lab from the configs in your Git repository
- 03Run your Ansible playbooks against the lab before every production change for one month
- 04Automate lab teardown and rebuild so that every test starts from a known state
- 05Generate traffic across the lab and verify the path taken matches what the design claims
Testing Network Changes Before They Ship
The step that separates scripting from engineering. Validate intent with Batfish, assert operational state with pytest, and run both in CI on every pull request. Done when a change that would black-hole a prefix is rejected by a pipeline rather than by a customer.
6 weeks7 SkillsBatfish configuration analysisPre-change validation and what-if analysisPost-change state assertionsCI pipelines for network repositoriesLinting configurations and playbooksTest data and topology fixturesFailing a pipeline safelyShow details, projects and resourcesSkills you'll master
Batfish configuration analysisadvancedPre-change validation and what-if analysisadvancedPost-change state assertionsintermediateCI pipelines for network repositoriesintermediateLinting configurations and playbooksintermediateTest data and topology fixturesintermediateFailing a pipeline safelyintermediateHands-on projects
- 01Run Batfish against your config repository and list every unreachable ACL line it finds
- 02Write assertions that fail if BGP session count drops after a change
- 03Wire a CI pipeline that lints, renders and validates every pull request against the repo
- 04Prove the pipeline works by opening a pull request that would break routing and watching it fail
- 05Add a post-deploy verification stage that rolls back automatically on assertion failure
When Ansible Stops Being Enough
At a few thousand devices, playbook runtime and rigid task structure become the constraint. Nornir keeps Python as the control flow and parallelises properly; know when the swap is justified and when it is résumé-driven. Done when you can state the device count at which your current tooling stops being the right answer.
5 weeks6 SkillsNornir inventory and task modelConcurrency and thread poolsAsync device interactionProfiling and runtime measurementStructured logging for bulk operationsPartial failure handling across many devicesShow details, projects and resourcesSkills you'll master
Nornir inventory and task modeladvancedConcurrency and thread poolsadvancedAsync device interactionadvancedProfiling and runtime measurementintermediateStructured logging for bulk operationsintermediatePartial failure handling across many devicesadvancedHands-on projects
- 01Reimplement your slowest Ansible playbook in Nornir and measure the runtime difference honestly
- 02Collect facts from a thousand simulated devices and record where the bottleneck actually is
- 03Handle a run where 3% of devices are unreachable without losing the results from the other 97%
- 04Emit structured logs from a bulk run that let you answer "which devices failed and why" from a query
- 05Write the decision note that says which tool your team should use, with the numbers behind it
Telemetry and Closing the Loop
Automation that only pushes configuration is half a system. Stream telemetry, define what good looks like, alert on the gap, and let a verified remediation run itself. Done when one recurring manual fix has been replaced by an automation you trust enough to leave unattended.
5 weeks6 SkillsStreaming telemetry with gNMITime-series storage and queryingAlerting on network SLIsAutomated remediation with guardrailsRunbook automation and approval gatesMeasuring automation reliabilityShow details, projects and resourcesSkills you'll master
Streaming telemetry with gNMIadvancedTime-series storage and queryingintermediateAlerting on network SLIsintermediateAutomated remediation with guardrailsadvancedRunbook automation and approval gatesintermediateMeasuring automation reliabilityadvancedHands-on projects
- 01Stream interface counters from ten devices into a time-series database and graph the difference against SNMP
- 02Define three network SLIs and alert on burn rate rather than on raw thresholds
- 03Automate one recurring fix end to end, with a kill switch and an audit trail
- 04Record every time your automation acted and review a month of it for false positives
- 05Write the post-incident note for the first time your automation does the wrong thing
Automation as a Product, Not a Side Project
The last barrier is organisational. Who may run what, how a change is approved, what happens when the person who wrote the tooling leaves. Done when a colleague who did not build the automation can use it to make a production change safely, guided only by its documentation.
5 weeks6 SkillsRole-based access to automationChange approval and audit trailsSelf-service interfaces for other teamsDocumentation that survives the authorMigration strategy for legacy devicesMeasuring adoption rather than coverageShow details, projects and resourcesSkills you'll master
Role-based access to automationadvancedChange approval and audit trailsintermediateSelf-service interfaces for other teamsadvancedDocumentation that survives the authorintermediateMigration strategy for legacy devicesadvancedMeasuring adoption rather than coverageadvancedHands-on projects
- 01Put your automation behind an interface a non-author can use, with permissions that mean something
- 02Write the runbook for what to do when the automation itself is the thing that broke
- 03Hand your tooling to a colleague, watch them use it, and fix only what they got stuck on
- 04Produce a migration plan for the devices that will never speak anything but SSH
- 05Report on how many changes went through automation versus by hand, monthly, and make the number public
What the job is actually like
- Day to day
- Half software engineering and half network engineering, and the balance shifts as the automation matures. Early on the week looks like writing scripts that replace something you used to type, chasing a device that returns almost-valid data, and reconciling a source of truth against what the network actually reports. Later it looks like maintaining a small internal product whose users are your colleagues, and the failures become software failures — a bad template that would have configured two hundred devices identically wrong. There is a change-window culture here that most software engineers find unfamiliar, and it exists for good reasons.
- The interview
- Two halves, often assessed by two different people. The network half is standard and unforgiving: routing, switching, and a troubleshooting scenario where you are expected to reason about layers rather than guess. The automation half asks for Python that parses something ugly, and increasingly for how you would test a network change before it ships — the question that separates candidates who script from candidates who engineer. Expect to be asked about a change that went wrong. Vendor-specific knowledge matters less than teams claim, and structured data handling matters more.
- How people get in
- Overwhelmingly from network engineering rather than from software. The typical arrival is a network engineer who got tired of typing the same configuration and learned Python, and that ordering is the right one — automating a network you do not understand produces outages at scale rather than efficiency. Software engineers do arrive from the other side and have to earn operational credibility, which takes longer than they expect. The reliability and platform paths share this roadmap's testing and source-of-truth phases, and the habits move in both directions. What does not transfer is assuming a rollback is cheap.
- After senior
- The obvious fork is toward platform and infrastructure engineering, where the network becomes one of several things managed as code and the title stops mentioning networks at all. A second fork stays in networking and goes up: network architect, or reliability engineering of the network itself in an organisation large enough to need it. A third is the vendor side — solutions engineering and developer advocacy for the tooling, which pays well and suits people who like explaining. The specialisation is narrow enough that the people who are good at it get known by name, which is unusual at this level.
- Why people leave
- The first is automating on top of a mess: scripts that encode the existing chaos and make it faster, which is how a team ends up unable to change either the network or the automation. A source of truth comes first for exactly this reason. The second is the lone automator, where one person builds tooling nobody else can maintain and the organisation quietly reverts to the command line the moment they leave; write it for a colleague or it will not survive you. The third is a company that says automation and means occasional scripts — no lab budget, no testing, and no appetite for the change.
Frequently asked questions
Related certifications
- Cisco Certified Network Associate Automation (200-901 CCNAAUTO)Cisco's automation associate exam — Python, REST APIs, Cisco platform SDKs, containers and CI/CD, model-driven programmability with YANG, NETCONF and RESTCONF. Renamed from DevNet Associate in February 2026.
- Linux Foundation Certified System Administrator (LFCS)A performance-based Linux administration certification taken entirely from the command line, covering deployment, networking, storage, essential commands and user management on a live system.
- Certified Kubernetes Administrator (CKA)A hands-on, performance-based certification proving you can install, configure, and troubleshoot production Kubernetes clusters from the command line.
Related roadmaps
- Site Reliability Engineer RoadmapA path from DevOps fundamentals into the specialized discipline of site reliability engineering, covering SLOs, observability, incident response, data reliability, and capacity planning.
- Platform Engineer RoadmapThe path DevOps engineers move into — building an internal developer platform as a product, covering Kubernetes as substrate, IaC at scale, GitOps, golden paths, portals, policy, multi-tenancy and adoption.
- DevOps Engineer RoadmapA structured path from Linux fundamentals through cloud infrastructure, automation, containers, and monitoring to a production-ready DevOps engineering career.
- Observability Engineer RoadmapA path into observability as a craft of its own — wide events, signal correlation, telemetry cost, collector pipelines, high-cardinality analysis, continuous profiling, and running observability as a platform other teams consume.