Skip to content
← Back to job listings

VP- SRE - Platform Reliability Engineering

Jefferies · Pune, India

Software DevelopmentExternal listingfull-timeabout 2 hours ago

About The Role

Innovation Hub Overview

Jefferies is creating a Technology Innovation Hub in Pune, a greenfield opportunity to build the systems that power global markets. As our first India technology center, this hub brings together hands on builders who engineer the platforms behind Jefferies’ growth across capital markets, investment banking, and institutional securities. We’re scaling toward an elite team of 500 engineers while maintaining the agility, ownership, and meritocratic spirit that defines Jefferies. From cloud and data to AI, risk, and core business technologies, teams in Pune will lead high impact work with a global mandate.

IT Platform Reliability Engineering & Management

The PREM team is responsible for improving the stability, availability, and performance of Jefferies’ core services and systems. The team provides day-to-day support for business-critical platforms across Equities, Fixed Income, Investment Banking, and Corporate functions, and builds tooling to strengthen observability (metrics/logs/dashboards), capacity, and performance management across distributed and cloud environments.

Vice President, Platform Reliability Engineer (SRE)

Location: Pune

Role Overview

We are seeking an experienced Vice President, Platform Reliability Engineer (SRE) to lead reliability engineering initiatives across critical front-to-back trading, post-trade, and operations platforms. The role combines hands-on technical expertise with strategic leadership, driving platform stability, operational excellence, observability, automation, and resilience across the technology estate.

The successful candidate will partner with Engineering, Infrastructure, Architecture, Operations, and Business stakeholders globally to define reliability standards, drive platform modernization, reduce operational risk, and improve service availability.

Key Responsibilities

Reliability & Platform Engineering Leadership

  • Provide technical leadership for platform reliability, stability, scalability, and resilience across business-critical production systems.
  • Define and drive the strategic roadmap for Platform Reliability Engineering and Site Reliability Engineering practices.
  • Establish reliability objectives, service level indicators (SLIs), service level objectives (SLOs), and operational excellence standards across supported platforms.
  • Act as a senior escalation point during major incidents, driving resolution, recovery, stakeholder communication, and post-incident reviews.
  • Lead root cause analysis initiatives and ensure corrective and preventative actions are implemented effectively.

Operational Excellence

  • Drive reduction of operational toil through automation, self-healing capabilities, and process simplification.
  • Establish best practices for incident management, problem management, change management, and release governance.
  • Identify reliability risks and proactively implement mitigation strategies to improve platform resilience.
  • Define operational KPIs and reliability metrics, leveraging data-driven insights to drive continuous improvement.

Observability & Monitoring

  • Own and enhance enterprise observability capabilities across applications, infrastructure, middleware, and cloud environments.
  • Drive adoption of modern observability frameworks leveraging Datadog, OpenTelemetry, Grafana, Prometheus, Loki, and Jaeger.
  • Ensure effective monitoring, alerting, logging, tracing, and capacity planning practices are implemented across platforms.

Engineering & Automation

  • Partner with development teams to embed reliability principles throughout the software development lifecycle.
  • Lead engineering efforts focused on infrastructure automation, deployment automation, and platform modernization.
  • Champion Infrastructure as Code (IaC), CI/CD, and DevOps best practices.
  • Drive automation initiatives using Python, Terraform, Ansible, Jenkins, Kubernetes, and cloud-native technologies.

Stakeholder & Team Leadership

  • Collaborate closely with senior technology leaders, application owners, infrastructure teams, cybersecurity teams, and business stakeholders.
  • Provide technical mentorship and guidance to SRE, PRE, DevOps, and Production Support engineers.
  • Influence technology strategy and architectural decisions with reliability, scalability, and operational sustainability in mind.
  • Lead cross-functional initiatives spanning multiple regions and technology teams.
  • Represent Platform Reliability Engineering in governance forums, technology reviews, and operational risk discussions.

Financial Services Platform Reliability

  • Ensure operational stability and support of platforms that underpin trading, post-trade processing, settlements, risk management, and regulatory reporting.
  • Maintain high service availability and minimize disruption to revenue-generating and business-critical workflows.
  • Drive regulatory, audit, and operational risk compliance within supported environments.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related discipline.
  • 8+ years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, Production Engineering, or Application Support.
  • Proven experience supporting and operating large-scale, mission-critical production platforms.
  • Strong programming experience in Python, Go, Java, C#, or similar languages.
  • Extensive experience with Linux/Unix environments and distributed systems.
  • Strong understanding of SRE principles, operational excellence frameworks, and reliability engineering practices.
  • Experience leading major incident management and problem management processes.
  • Strong understanding of databases, messaging systems, and middleware technologies.
  • Excellent troubleshooting and analytical problem-solving skills across application, infrastructure, and data layers.
  • Experience working in globally distributed teams and managing senior stakeholder relationships.
  • Strong verbal and written communication skills with both technical and business audiences.

Preferred Qualifications

Observability

  • Datadog
  • OpenTelemetry
  • Grafana
  • Prometheus
  • Loki
  • Jaeger

DevOps & Automation

  • Git
  • Jenkins
  • GitHub Actions
  • Ansible
  • Terraform
  • CI/CD Frameworks

Cloud & Containers

  • Kubernetes
  • Docker
  • OpenShift
  • AWS / Azure / GCP

Data & Messaging Platforms

  • Kafka
  • Redis
  • MongoDB
  • Elasticsearch
  • PostgreSQL
  • SQL Server

Financial Services Experience

  • Investment Banking
  • Capital Markets
  • Equities
  • Fixed Income
  • Prime Brokerage
  • Post-Trade Processing
  • Operations Technology

This is an external listing. JobSpring does not represent or verify the employer. Report this listing