Senior Site Reliability Engineer
Eeho · BENGALURU, KARNATAKA, India
About The Role
Our primary objectives are
- Ensure maximum possible service availability and performance
- Deliver premier customer service
- Provide comprehensive support to our Engineering and other Operational and technical teams
These objectives translate into a broad and dynamic scope of responsibilities for the GNOC team. Engineers will have the capability to centrally manage OCI’s networks and implement automated solutions to address common operational challenges efficiently.
NETWORK OPERATIONS
- Fault handling on incident tickets. provide break-fix support and escalation for event remediation, including leading root cause analysis (RCA) efforts
- Work closely with shift lead to ensure tickets are handled in a timely manner
- Using existing procedures and tooling, develop and safely complete network changes
- Mentor junior engineers as needed
- Participate in operational rotations providing break-fix support
- Possess analytical skills in resolving network issues with advanced troubleshooting and coordination with onsite support teams and vendors
- Identifying actionable incidents using a monitoring system, strong analytical problem-solving skills to mitigate network events/incidents, and following up on routine root cause analysis (RCA), coordinating with support teams and vendors
- Join major events/incidents calls, use technical and analytical skills to resolve network issues that impact Oracle customer/service, coordinate with SMEs, and provide RCA document
- Fault handling and escalation (identifying and responding to faults on OCI’s systems and networks, collaborating closely with 3rd party suppliers, handling escalation through to resolution)
- Experience with Incident Response plans and strategies in cloud computing environments
- Have worked with large enterprise network infrastructure and cloud computing environments, supporting 24/7 and willing to work in rotational shifts in a network operations role
- Provide on-call support services as needed, job duties are varied and complex, needing independent judgment
·
AUTOMATIONS/SCRIPTING
- The role includes collaborating with networking automation services to integrate support tooling and frequently developing scripts to automate routine tasks
- Preference for individuals with experience in scripting and network automation - Python, Puppet, SQL, and/or Ansible
- You will use automation to complete work and develop scripts for routine tasks
PROJECT MANAGEMENT
- Ability to act in a project lead role as needed
TECHNICAL QUALIFICATIONS
NETWORKING
- Experience working in a large ISP or cloud provider environment
- Exposure to commodity Ethernet hardware (Broadcom/Mellanox)
- Protocol experience with BGP/OSPF/IS-IS, TCP, IPv4, IPv6, DNS, DHCP, MPLS
- Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL
- Experience supporting network technologies, especially Juniper, Cisco, Arista, firewalls, and switches
- Strong analytical skills and ability to collate and interpret data from various sources
- Ability to diagnose network alerts to assess and prioritize faults and respond or escalate accordingly
- Cisco and Juniper certifications are desired
SOFT SKILLS
- Highly motivated and self-starter
- Bachelor’s degree is preferred with at least 3-5 years of network-related experience
- Strong oral and written communication skills
- Excellent time management and organization skills
- Comfortable and able to deal with a wide range of issues in a fast-paced environment
- Excellent organizational, verbal, and written communication skills
Key Responsibilities
Capacity Ingestion and Management
- -Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.
- -Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.
- -Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.
- -Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).
Incident and Service Lifecycle Management
-Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
-Independently monitors services, maintains up-to-date knowledge of their performance, and documents their condition.
-Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
-Provides health and performance reporting and takes appropriate actions based on trends in data.
-May independently perform provisioning to support infrastructure, applications, and services.
-May perform standard and non-standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
Automation
- -Identifies opportunities for automation and assesses potential benefits.
- -Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
- -Independently conducts testing to ensure automation performs the task correctly and produces expected results.
Technical Communication and Guidance
- -Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.
- -Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.
Troubleshooting and Resolution
- -Provides operational support for technology, escalating incidents and other standard and non-standard issues arising within Oracle services.
- -Participates in on-call shifts to address issues.
- -Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).
- -Documents incidents and performs root cause analyses according to standard reporting methods.
- -Independently performs post-mortem procedures to prevent incident reoccurrence.
Innovation and Improvement
- -Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.
- -Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.
- -Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.
- -Performs standard and non-standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).
Core Responsibilities
Planning & Execution
Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
Collaboration & Partnership
Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving
Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.
Continuous Learning
Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement
Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.
Minimum Job Qualifications
Education and/or Experience
8 years of experience in software engineering, infrastructure management, or related field
OR
Bachelor's Degree in Computer Science, Engineering, or related field AND 4 years of experience in software engineering, infrastructure management, or related field
OR
Master's Degree in Computer Science, Engineering, or related field AND 2 year of experience in software engineering, infrastructure management, or related field.
OR
Doctorate in Computer Science, Engineering, or related field
Job Skills
- Same skills as prior level plus;
- Operating Systems Demonstrated ability in or knowledge of operating systems, including installing, upgrading, and troubleshooting various operating environments.
Automation Experience
3 years of experience in automation.
Programming Experience
3 years of experience in programming and/or scripting.
Preferred Job Qualifications
Education and/or Experience
9 years of experience in software engineering, infrastructure management, or related field
OR
Bachelor's Degree in Computer Science, Engineering, or related field AND 5 years of experience in software engineering, infrastructure management, or related field
OR
Master's Degree in Computer Science, Engineering, or related field AND 3 years of experience in software engineering, infrastructure management, or related field
OR
Doctorate in Computer Science, Engineering, or related field AND 1 year of experience in software engineering, infrastructure management, or related field.
Automation Experience
5 years of experience in automation.
Programming Experience
5 years of experience in programming and/or scripting.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
