Skip to content
← Back to job listings

Senior Lead Infrastructure Engineer — On Prem OpenShift Platform

JPMorgan Chase · Jersey City, NJ, United States

IT - Network / Systems / DB AdminExternal listingfull-time17 minutes ago

About The Role

We're looking for a talented, senior engineering professional ready to take their career to new heights at one of the world's most influential companies.

As a Senior Lead Infrastructure Engineer at JPMorgan Chase within Enterprise Technology Compute Infrastructure Platforms team, you will design, engineer, and operate an on-premises OpenShift-based platform that hosts both containers and virtual machines (e.g., OpenShift Virtualization / KubeVirt). You will own the underlying infrastructure and OpenShift cluster foundations, compute, hypervisor platforms, operating systems, networking, storage, security controls, and reliability enabling product teams to deploy workloads safely, consistently, and efficiently. This is an infrastructure engineering role focused on platform lifecycle and reliability, partnering closely with hardware engineering, networking, storage, security, identity/access, and platform consumers.

Job Responsibilities

  • Design OpenShift cluster architectures (on-premises and/or hybrid) to meet availability, scalability, security, and operability requirements.
  • Build and operate OpenShift clusters end-to-end, including install, upgrade, patching, scaling, resilience testing as well as virtualization capabilities supporting VM lifecycle (images/templates and relevant hardware acceleration concepts such as SR-IOV/DPDK/GPU passthrough, where applicable).
  • Engineer the platform to support both Kubernetes workloads and VM workloads, including capacity planning, performance management, placement strategies, and HA/DR considerations.
  • Own Linux platform fundamentals (e.g., RHEL/CoreOS concepts, kernel/sysctl tuning, certificates, and identity integration basics) required for reliable cluster operations.
  • Implement and troubleshoot core networking capabilities (routing/switching fundamentals, DNS, L4/L7 concepts, load balancing, firewalling, segmentation, and packet-path analysis).
  • Integrate storage services for container and VM workloads (block/file/object concepts, CSI drivers, performance and failure modes, and backup/restore approaches).
  • Drive reliability practices including observability, incident response, root-cause analysis, and preventative improvements across the platform lifecycle.
  • Develop runbooks, standards, and reference architectures; lead operational readiness reviews and post-incident actions.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate analysis of complex infrastructure signals and documentation of mitigation options, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted practices across delivery and automation routines to reduce recurring issues, ensuring changes are validated, traceable and auditable, and aligned to resiliency and security expectations.

Required qualifications, capabilities, and skills

  • Formal training or certification on infrastructure engineering concepts and 5+ years applied experience.
  • Hands-on administration of Red Hat OpenShift and Kubernetes in production environments (e.g., install/upgrade patterns such as IPI/UPI, Operators, SCC/RBAC, cluster operators, ingress).
  • Experience operating Kubernetes/container platforms through day-2 operations (upgrades, scaling, troubleshooting, and platform lifecycle management).
  • Experience with OpenShift Virtualization / KubeVirt (or equivalent VM-on-Kubernetes) and VM/container co-tenancy design considerations.
  • Strong Linux administration and troubleshooting experience (e.g., system performance, certificates, OS configuration, and cluster node operations).
  • Demonstrated troubleshooting depth in at least two of the following domains: networking, storage, virtualization.
  • Experience designing and operating highly available platforms, including capacity management, incident handling, root-cause analysis, and remediation tracking.
  • Experience partnering with security/compliance stakeholders on hardening, RBAC, secrets/certificates handling, and vulnerability management aligned to secure-by-default configurations.
  • Experience implementing infrastructure automation and configuration-as-code practices (e.g., Ansible/Terraform/GitOps patterns) with traceability and auditability.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.

Preferred qualifications, capabilities and Skills

  • Familiarity with OpenShift observability stacks (Prometheus/Grafana/Alertmanager, logging, tracing) and operational maturity practices.
  • Experience with service mesh, ingress controllers, API gateways, or platform networking plugins aligned to your stack.
  • Experience with enterprise storage technologies and their performance/failure characteristics in containerized environments.
  • Familiarity using generative AI tools to accelerate research, troubleshooting, and documentation workflows (with appropriate governance).

This is an external listing. JobSpring does not represent or verify the employer. Report this listing