← Back to job listings
BY
Senior Site Reliability Engineer - Traffic Infrastructure
ByteDance · San Jose, California, United States of America
About The Role
About the Team
The Global Traffic Infrastructure (GTI) team leverages unified platform capabilities to manage edge infrastructure outside China (both self-built and third-party) providing standardized, compliant, scalable, and cost-effective traffic infrastructure capabilities for edge services. Our vision is to build a global edge traffic infrastructure platform and become the long-term cornerstone of ByteDance’s global edge business in terms of scale, performance, and cost.
Responsibilities
- Responsible for the operation and maintenance as well as stability assurance of ByteDance's "Network-Traffic Infrastructure".
- Responsible for the delivery and operation & maintenance of the production system, including the delivery, change and release of facilities, components and products, and improving the efficiency of both delivery and operation & maintenance.
- Responsible for the design and implementation of the stability assurance system, covering system observability (monitoring/alerts/logging), troubleshooting (root cause analysis/impact assessment), and issue resolution (manual/self-healing).
- Responsible for the design and implementation of the emergency response system, including work such as ticket processing, emergency response, risk governance, and long-term optimization, to enhance the risk emergency response capability.
Minimum Qualifications
- Bachelor's degree or above in computer science or a related field, with at least 3 years of relevant experience in R&D, system operation and maintenance, or SRE.
- Familiar with infrastructure architecture, and have a solid understanding of Kubernetes, edge computing, cloud networking, Load Balance, microservice architecture and other related technologies.
- Possess strong analytical skills, excellent communication abilities, a strong sense of responsibility and team spirit.
Preferred Qualifications
- Practical operation and maintenance and stability assurance experience in Kubernetes, cloud computing, edge computing, and cloud networking.
- Well-versed in high availability, stability/reliability assurance, and emergency response systems for infrastructure or distributed systems, with relevant operation and maintenance experience.
- Hands-on experience in handling risks, potential hazards and failures of infrastructure or distributed systems.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring