Senior Consultant - Site Reliability Engineer
HCA Healthcare India · Hyderabad, Hyderabad, Telangana, India
About The Role
Senior Consultant - Site Reliability Engineering (SRE) Experience 11+ years About This Role The Senior Consultant - Site Reliability Engineering (SRE) is a hands-on technical leadership and consulting role responsible for driving the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical services. The role serves as a principal technical escalation point and trusted advisor for complex production challenges, shaping SRE strategy and engineering improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness. The Senior Consultant partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, Security, and business stakeholders; leads cross-functional initiatives; and mentors senior engineers. Essential Duties § Serve as a principal technical escalation point and SRE consultant for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting, executive-level technical communication, and service restoration. § Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry. § Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk. § Define, govern, and mature SRE and observability practices, including SLIs/SLOs, error budgets, dashboards, alerting, instrumentation, event correlation, monitoring coverage, and reliability reporting. § Lead enterprise reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop prioritized multi-quarter improvement roadmaps. § Lead strategic toil-reduction programs and design reusable automation, self-healing, and auto-remediation patterns that improve engineering productivity and service reliability at scale. § Provide technical governance and hands-on leadership for Infrastructure-as-Code, configuration management, GitOps, platform engineering, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab. § Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability. § Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers. § Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements. § Own technical direction for disaster recovery, business continuity engineering, resiliency validation, recovery objectives/procedures, failure testing, and rollback readiness. § Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis. § Establish SRE standards, reference architectures, playbooks, runbooks, templates, operating procedures, governance mechanisms, and reusable engineering patterns across teams. § Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, platform engineering, and AI-assisted operations; lead proof-of-concept evaluations, technical recommendations, and adoption roadmaps. § Mentor senior SRE, Production Engineering, DevOps, and Application Support engineers; provide technical coaching in troubleshooting, observability, automation, incident management, root cause analysis, architecture, and reliability practices. § Use incident trends, operational metrics, SLO performance, risk indicators, and engineering data to define reliability priorities, influence stakeholders, and drive measurable continuous improvement. § Provide consultative leadership to application and platform teams on reliability architecture, SRE adoption, production readiness, cloud modernization, and operational risk reduction. § Lead reliability reviews with senior stakeholders, translate technical risk into business impact, and define measurable remediation plans, success criteria, and governance checkpoints. § Participate in an on-call or senior production escalation rotation where required. Position Requirements § 11+ years of progressive experience in Site Reliability Engineering, Production Engineering, Application/Production Support, DevOps, Cloud/Platform Engineering, or a related technical discipline, including significant experience leading complex enterprise reliability initiatives. § Strong hands-on experience supporting enterprise or business-critical production applications and troubleshooting across multiple technology layers. § Advanced experience with monitoring, logging, APM, and observability platforms such as Dynatrace, Splunk, Grafana, or comparable technologies. § Strong working knowledge of Windows and/or Linux/Unix platforms, relational databases, SQL, APIs, integrations, and distributed application dependencies. § Experience with cloud platforms such as Azure, GCP, AWS, or comparable enterprise cloud technologies. § Hands-on experience with CI/CD, source control, GitOps/deployment tools, Infrastructure-as-Code, and automation technologies such as Azure DevOps, GitHub/GitLab, Argo CD, Terraform, or Ansible. § Strong understanding of networking concepts including DNS, firewall rules, ports, load balancing, routing, certificates, and application connectivity. § Strong understanding of identity, privileged access, service accounts, secrets, and enterprise security concepts. § Demonstrated experience leading complex incident troubleshooting, root cause analysis, and implementation of corrective/preventive improvements. § Demonstrated ability to act as a senior technical consultant, mentor experienced engineers, influence architecture and engineering decisions, communicate effectively with technical and executive stakeholders, and drive cross-functional improvements without formal people-management authority. Education § Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience and demonstrated SRE leadership may be considered. § Relevant certifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial. Knowledge and Skills Capability Knowledge & Skill Expectation Application & Production Engineering Advanced troubleshooting across complex applications, integrations, dependencies, and production environments. Observability & Reliability Expert knowledge of SRE principles, SLIs/SLOs, error budgets, logs, metrics, traces, dashboards, alerting, performance, availability, resilience, and reliability governance. Cloud & Infrastructure Strong understanding of cloud platforms, operating systems, databases, networking, infrastructure services, and cross-platform dependencies. DevOps & Infrastructure-as-Code Hands-on experience with CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, and Infrastructure-as-Code practices. Automation & Engineering Ability to design reusable automation, reduce operational toil, and develop self-healing or remediation workflows. Incident & Problem Management Ability to lead complex incident troubleshooting, RCA, corrective actions, and preventive improvements. Performance, Capacity & Security Ability to analyze performance/capacity trends and troubleshoot application-security, identity, access, and configuration issues with specialist teams. Technical Leadership & Improvement Expert ability to operate as a senior SRE consultant, mentor experienced engineers, influence architecture and stakeholders, establish enterprise standards, lead transformation roadmaps, and drive measurable reliability improvements.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
