Site Reliability Engineer III
JPMorgan Chase · Jersey City, NJ, United States
About The Role
There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.
As a Site Reliability Engineer III at JPMorgan Chase within the Asset and Wealth Management team, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform.
Job responsibilities
- Supports and collaborates with other engineers to design, develop, and implement deployment and reliability approaches using automated CI/CD pipelines
- Implements infrastructure, configuration, and network as code for applications and platforms in your remit
- Contributes to observability improvements including white and black box monitoring, service level objective alerting, and telemetry collection
- Assists in identifying and resolving complex problems by using service level indicators and objectives to proactively address issues before they impact customers
- Participates in incident response, triage, and post-incident analysis; helps document findings and remediation actions to prevent recurrence
- Identifies opportunities to eliminate or automate remediation of recurring issues to reduce toil and improve overall operational stability
- Uses enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
- Proactively recognizes roadblocks and identifies improvements to solve operational problems, including exploring new technologies where appropriate
- Documents and shares knowledge within your organization via internal forums and communities of practice
- Supports adoption of site reliability engineering best practices within your team
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Foundational understanding of SRE culture and principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- General observability and monitoring understanding with working experience using industry-standard tooling (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk, CloudWatch)
- Proficiency in at least one scripting or programming language such as Python, Bash, or similar for tool development and operational support
- Experience with incident and response management, including on-call participation and structured post-incident review
- Knowledge of CI/CD pipelines and best practices using tools such as Jenkins, GitLab CI, or similar
- Skills in automating repetitive tasks and managing configurations at scale using tools such as Ansible, Terraform, or similar
- Familiarity with container technologies and container orchestration (e.g., Docker, Kubernetes)
- Working knowledge of using enterprise-authorized AI capabilities to support SRE workflows, with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred qualifications, capabilities, and skills
- Experience operating in cloud environments (AWS and/or Azure), including understanding of resiliency, scalability, and observability patterns
- Familiarity with Kubernetes ingress, networking, and certificate deployment patterns
- Experience improving infrastructure-as-code patterns (e.g., Terraform modules, reusable configurations)
- Understanding of controls-focused operations in regulated environments, including change management discipline and audit support
- Experience with service mesh, load balancing, and DNS troubleshooting (e.g., ALB/NLB, Route 53)
- Drive to self-educate and evaluate emerging technologies in the SRE and cloud-native space
- Strong communication skills with the ability to collaborate across different levels and stakeholder groups
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
