Network Automation & Reliability Engineer
Ontrac Solutions · Chicago, IL
About The Role
Overview
Ontrac Solutions is seeking a
Network Automation & Reliability Engineer
to operate and automate hyperscale data center, backbone, and out-of-band (OOB) networks. This is a
Python-first
role: you will spend more of your week writing code that operates the network than typing on a CLI. You will build the tooling, telemetry pipelines, and self-healing automation that keep a global fleet of data centers and POP sites running, and you will carry the on-call pager for the systems you build.
We are looking for an engineer who is genuinely production-grade on
both
sides of the job — deep routing and switching fundamentals
and
- real, sustained Python software work. Candidates who are strong on one side and thin on the other will not clear screening.
- What your application must clearly show
We screen against the requirements below exactly as written — your resume should make these easy to find
- Python you actually wrote, described as software, not as a skills keyword.
- Name the project, what it did, roughly how large it was, and who used it. "Python (scripting)" in a skills list will not clear this bar.
- A GitHub, GitLab, or public repo link is strongly preferred
- — we look at code.
The
- specific Python network libraries
- you have used in production — for example Netmiko, NAPALM, Nornir, pyATS/Genie, Scrapli, ncclient, or a vendor SDK — and what you built with each.
- BGP in production at scale.
- Name the policy work: local preference, MED, communities, import/export policy, route reflection, ECMP, or multihoming across carriers.
Hands-on with
- at least two
- of Arista EOS, Juniper Junos (QFX / SRX / PTX / MX), or Cisco IOS-XR / NX-OS — named by platform, not just by vendor.
- Telemetry and observability
- you implemented: gNMI/gRPC, OpenConfig, streaming telemetry, flow telemetry, or SNMP-to-TSDB pipelines. Say what you subscribed to and what you did with the data.
Config-as-code
Jinja2 templating, NETCONF/YANG or REST-API-driven provisioning, golden configs, ZTP, and the Git/CI workflow you shipped changes through.
Your
- networking certifications, named, with dates and credential IDs or verification links
- — we verify certifications.
- Whether you have carried
- production on-call
- , and at what scale (sites, devices, or POPs).
- Shortly after you apply you will receive a short role-specific questionnaire — completing it promptly is the fastest way to move into screening.
Required Qualifications
Python — primary requirement
Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts.
Routing & switching depth
Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays.
Multi-vendor hardware
Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS.
Automation frameworks
Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management.
Telemetry & monitoring
gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them.
Linux
Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting.
Version control and CI
Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions).
Production on-call
Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis.
Experience level
Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply.
Location & work authorization
Must be located in the United States and authorized to work in the US.
Preferred Qualifications
Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity.
Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing.
MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC.
Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation.
Building or contributing to a network
source of truth
(NetBox or in-house) aggregating BGP, link-state, and drain-state data.
Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization.
A master's degree in Network Engineering, Telecommunications, or Computer Science.
Key Responsibilities — Network Automation & Tooling (primary focus)
- Design, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet.
- Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift.
- Develop automated remediation for recurring failure classes — route flaps, hardware faults, optical degradation — triggered by syslog and telemetry.
- Reduce operational toil: identify manual runbook steps that recur, and replace them with tested, reviewed code.
Key Responsibilities — Backbone, Data Center & Out-of-Band Operations
Operate and scale data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites.
Design and tune BGP policy — local preference, MED, communities, import/export policy, default-route propagation — to control path selection and eliminate single-carrier points of failure.
Support capacity expansion and site turn-ups: high-level design, port channelization, optics and cabling standards, and hardware qualification.
Execute zero-downtime change management in production, including hitless migrations and staged rollbacks.
Key Responsibilities — Telemetry, Observability & Reliability
- Operationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal quality and management-plane efficiency.
- Build and maintain observability that supports hop-by-hop path tracing, multi-layer fault isolation, and fast root cause analysis.
- Contribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data.
- Participate in a 24x7 on-call rotation; lead incident response, write RCAs, and drive the follow-up automation that prevents recurrence.
Key Responsibilities — Security & Change Governance
- Maintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and performant.
- Support SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines.
- Author and maintain technical documentation — designs, runbooks, and API contracts — for the tooling and networks you own.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring