Skip to content
← Back to job listings

Senior Consultant - Site Reliability Engineer

HCA Healthcare India · Hyderabad, Hyderabad, Telangana, India

Imported listingfull-time11 days ago

About The Role

Senior Consultant - Site Reliability Engineering (SRE) Experience 11+ years About This Role The Senior Consultant - Site Reliability Engineering (SRE) is a hands-on technical leadership and consulting role responsible for driving the reliability, availability, performance, resilience, and operational maturity of enterprise and business-critical services. The role serves as a principal technical escalation point and trusted advisor for complex production challenges, shaping SRE strategy and engineering improvements across observability, automation, Infrastructure-as-Code, incident management, reliability engineering, performance optimization, and operational readiness. The Senior Consultant partners with Development, Architecture, SRE, DevOps, Cloud, Database, Network, Infrastructure, Security, and business stakeholders; leads cross-functional initiatives; and mentors senior engineers. Essential Duties § Serve as a principal technical escalation point and SRE consultant for complex, high-impact, and business-critical incidents; lead cross-functional troubleshooting, executive-level technical communication, and service restoration. § Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry. § Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk. § Define, govern, and mature SRE and observability practices, including SLIs/SLOs, error budgets, dashboards, alerting, instrumentation, event correlation, monitoring coverage, and reliability reporting. § Lead enterprise reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop prioritized multi-quarter improvement roadmaps. § Lead strategic toil-reduction programs and design reusable automation, self-healing, and auto-remediation patterns that improve engineering productivity and service reliability at scale. § Provide technical governance and hands-on leadership for Infrastructure-as-Code, configuration management, GitOps, platform engineering, and CI/CD practices using technologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab. § Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability. § Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers. § Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements. § Own technical direction for disaster recovery, business continuity engineering, resiliency validation, recovery objectives/procedures, failure testing, and rollback readiness. § Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis. § Establish SRE standards, reference architectures, playbooks, runbooks, templates, operating procedures, governance mechanisms, and reusable engineering patterns across teams. § Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, platform engineering, and AI-assisted operations; lead proof-of-concept evaluations, technical recommendations, and adoption roadmaps. § Mentor senior SRE, Production Engineering, DevOps, and Application Support engineers; provide technical coaching in troubleshooting, observability, automation, incident management, root cause analysis, architecture, and reliability practices. § Use incident trends, operational metrics, SLO performance, risk indicators, and engineering data to define reliability priorities, influence stakeholders, and drive measurable continuous improvement. § Provide consultative leadership to application and platform teams on reliability architecture, SRE adoption, production readiness, cloud modernization, and operational risk reduction. § Lead reliability reviews with senior stakeholders, translate technical risk into business impact, and define measurable remediation plans, success criteria, and governance checkpoints. § Participate in an on-call or senior production escalation rotation where required. Position Requirements § 11+ years of progressive experience in Site Reliability Engineering, Production Engineering, Application/Production Support, DevOps, Cloud/Platform Engineering, or a related technical discipline, including significant experience leading complex enterprise reliability initiatives. § Strong hands-on experience supporting enterprise or business-critical production applications and troubleshooting across multiple technology layers. § Advanced experience with monitoring, logging, APM, and observability platforms such as Dynatrace, Splunk, Grafana, or comparable technologies. § Strong working knowledge of Windows and/or Linux/Unix platforms, relational databases, SQL, APIs, integrations, and distributed application dependencies. § Experience with cloud platforms such as Azure, GCP, AWS, or comparable enterprise cloud technologies. § Hands-on experience with CI/CD, source control, GitOps/deployment tools, Infrastructure-as-Code, and automation technologies such as Azure DevOps, GitHub/GitLab, Argo CD, Terraform, or Ansible. § Strong understanding of networking concepts including DNS, firewall rules, ports, load balancing, routing, certificates, and application connectivity. § Strong understanding of identity, privileged access, service accounts, secrets, and enterprise security concepts. § Demonstrated experience leading complex incident troubleshooting, root cause analysis, and implementation of corrective/preventive improvements. § Demonstrated ability to act as a senior technical consultant, mentor experienced engineers, influence architecture and engineering decisions, communicate effectively with technical and executive stakeholders, and drive cross-functional improvements without formal people-management authority. Education § Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience and demonstrated SRE leadership may be considered. § Relevant certifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial. Knowledge and Skills Capability Knowledge & Skill Expectation Application & Production Engineering Advanced troubleshooting across complex applications, integrations, dependencies, and production environments. Observability & Reliability Expert knowledge of SRE principles, SLIs/SLOs, error budgets, logs, metrics, traces, dashboards, alerting, performance, availability, resilience, and reliability governance. Cloud & Infrastructure Strong understanding of cloud platforms, operating systems, databases, networking, infrastructure services, and cross-platform dependencies. DevOps & Infrastructure-as-Code Hands-on experience with CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, and Infrastructure-as-Code practices. Automation & Engineering Ability to design reusable automation, reduce operational toil, and develop self-healing or remediation workflows. Incident & Problem Management Ability to lead complex incident troubleshooting, RCA, corrective actions, and preventive improvements. Performance, Capacity & Security Ability to analyze performance/capacity trends and troubleshoot application-security, identity, access, and configuration issues with specialist teams. Technical Leadership & Improvement Expert ability to operate as a senior SRE consultant, mentor experienced engineers, influence architecture and stakeholders, establish enterprise standards, lead transformation roadmaps, and drive measurable reliability improvements.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing