← Back to job listings
EG
Site Reliability Engineer III
Egug · Bengaluru, KA, India
About The Role
Site Reliability Engineer III advances efforts to enhance system resilience, scalability, and performance through feature development, automation, architectural design, chaos engineering, and disaster recovery planning, while promoting best practices for continuous improvement and reliability.
- Manages the collaboration with Software Engineering teams to design, develop, and implement features that enhance system resilience, scalability, and performance, proactively identifying and resolving system bottlenecks and failure points
- Develops and refines sophisticated automation tools and frameworks, including advanced infrastructure as code (IaC) practices, to streamline operational workflows, deployment processes, and infrastructure management, ensuring high system efficiency
- Engages in architectural design discussions, ensuring that advanced reliability, scalability, and performance considerations are integrated into strategic decision-making processes
- Designs and executes comprehensive chaos engineering experiments and advanced resiliency testing, analyzing results to implement robust improvements that enhance system robustness and recovery capabilities
- Develops, optimizes, and maintains comprehensive disaster recovery plans and business continuity strategies, ensuring systems can recover quickly and effectively from complex and unexpected disruptions
- Advocates for observability practices by promoting and implementing best practices such as error budgeting, service-level objectives (SLOs), and service-level indicators (SLIs), contributing to a culture of continuous improvement and reliability
- Collaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives
Education Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, and/or comparable experience; advance degree preferred
- Knowledge of modern observability stack – Splunk, Elastic Search, Prometheus, Grafana
- Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture
- Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms
- Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud
Work Experience
- Experience in software development, or technology operations, with a focus on Site Reliability Engineering
- Experience in Linux/Unix systems, object-oriented programming languages (e.g., Java), scripting languages (e.g., Python, Bash), and cloud platforms (e.g., AWS, Azure, GCP)
Licenses and Certifications
- Advanced certification in Site Reliability Engineering (SRE) or related is a plus
Similar roles you might like
See all →NI
Manager Software Engineering, ITC
NIKE, Inc.
Salary not disclosedPosted today
NI
Lead Software Engineer, ITC
NIKE India Technology Center Private Limited
Salary not disclosedPosted today
4D
Senior Software Engineer_Golang Developer
451 Discovery Communications India
Salary not disclosedPosted today
MG
AVP Senior Software Engineer
MUFG Global Service Private Limited
Salary not disclosedPosted today
NV
Site Reliability Engineer (II)
NCR Voyix
Salary not disclosedPosted today
NV
Software Engineer I
NCR Voyix
Salary not disclosedPosted today
SI
AVP, Software Engineer III, Servicing Apps (L10) position
Synchrony International Services Private Limited
Salary not disclosedPosted today
RI
Lead Salesforce Developer
RingCentral Innovation India
Salary not disclosedPosted today
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
