Skip to content
← Back to job listings

Network Automation & Reliability Engineer

Ontrac Solutions · Chicago, IL

RemoteExternal listingcontract10 days ago

About The Role

Overview

Ontrac Solutions is seeking a

Network Automation & Reliability Engineer

to operate and automate hyperscale data center, backbone, and out-of-band (OOB) networks. This is a

Python-first

role: you will spend more of your week writing code that operates the network than typing on a CLI. You will build the tooling, telemetry pipelines, and self-healing automation that keep a global fleet of data centers and POP sites running, and you will carry the on-call pager for the systems you build.

We are looking for an engineer who is genuinely production-grade on

both

sides of the job — deep routing and switching fundamentals

and

  • real, sustained Python software work. Candidates who are strong on one side and thin on the other will not clear screening.
  • What your application must clearly show

We screen against the requirements below exactly as written — your resume should make these easy to find

  • Python you actually wrote, described as software, not as a skills keyword.
  • Name the project, what it did, roughly how large it was, and who used it. "Python (scripting)" in a skills list will not clear this bar.
  • A GitHub, GitLab, or public repo link is strongly preferred
  • — we look at code.

The

  • specific Python network libraries
  • you have used in production — for example Netmiko, NAPALM, Nornir, pyATS/Genie, Scrapli, ncclient, or a vendor SDK — and what you built with each.
  • BGP in production at scale.
  • Name the policy work: local preference, MED, communities, import/export policy, route reflection, ECMP, or multihoming across carriers.

Hands-on with

  • at least two
  • of Arista EOS, Juniper Junos (QFX / SRX / PTX / MX), or Cisco IOS-XR / NX-OS — named by platform, not just by vendor.
  • Telemetry and observability
  • you implemented: gNMI/gRPC, OpenConfig, streaming telemetry, flow telemetry, or SNMP-to-TSDB pipelines. Say what you subscribed to and what you did with the data.

Config-as-code

Jinja2 templating, NETCONF/YANG or REST-API-driven provisioning, golden configs, ZTP, and the Git/CI workflow you shipped changes through.

Your

  • networking certifications, named, with dates and credential IDs or verification links
  • — we verify certifications.
  • Whether you have carried
  • production on-call
  • , and at what scale (sites, devices, or POPs).
  • Shortly after you apply you will receive a short role-specific questionnaire — completing it promptly is the fastest way to move into screening.

Required Qualifications

Python — primary requirement

Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts.

Routing & switching depth

Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays.

Multi-vendor hardware

Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS.

Automation frameworks

Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management.

Telemetry & monitoring

gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them.

Linux

Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting.

Version control and CI

Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions).

Production on-call

Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis.

Experience level

Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply.

Location & work authorization

Must be located in the United States and authorized to work in the US.

Preferred Qualifications

Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity.

Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing.

MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC.

Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation.

Building or contributing to a network

source of truth

(NetBox or in-house) aggregating BGP, link-state, and drain-state data.

Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization.

A master's degree in Network Engineering, Telecommunications, or Computer Science.

Key Responsibilities — Network Automation & Tooling (primary focus)

  • Design, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet.
  • Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift.
  • Develop automated remediation for recurring failure classes — route flaps, hardware faults, optical degradation — triggered by syslog and telemetry.
  • Reduce operational toil: identify manual runbook steps that recur, and replace them with tested, reviewed code.

Key Responsibilities — Backbone, Data Center & Out-of-Band Operations

Operate and scale data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites.

Design and tune BGP policy — local preference, MED, communities, import/export policy, default-route propagation — to control path selection and eliminate single-carrier points of failure.

Support capacity expansion and site turn-ups: high-level design, port channelization, optics and cabling standards, and hardware qualification.

Execute zero-downtime change management in production, including hitless migrations and staged rollbacks.

Key Responsibilities — Telemetry, Observability & Reliability

  • Operationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal quality and management-plane efficiency.
  • Build and maintain observability that supports hop-by-hop path tracing, multi-layer fault isolation, and fast root cause analysis.
  • Contribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data.
  • Participate in a 24x7 on-call rotation; lead incident response, write RCAs, and drive the follow-up automation that prevents recurrence.

Key Responsibilities — Security & Change Governance

  • Maintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and performant.
  • Support SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines.
  • Author and maintain technical documentation — designs, runbooks, and API contracts — for the tooling and networks you own.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing