← Back to job listings

机器学习平台 SRE 工程师
moonshot · 南山区, 广东, 中国; 海淀区, 北京市, 中国
About The Role
职位描述 负责月之暗面大规模 GPU 训练与推理集群的稳定性保障,支撑业务 7×24 高可用运行; 负责 Kubernetes 及云原生周边系统(监控、日志、镜像分发、存储)的运维保障与平台化工具开发以及疑难问题的排查解决; 负责 GPU 节点硬件故障的自动化巡检、自愈体系与告警治理; 参与 OnCall 值班,响应集群级突发事件(网络拥塞、调度热点、训练任务失败)。 职位要求 3 年以上大规模分布式系统或云原生基础设施 SRE 经验,有 1000+ 节点 Kubernetes 集群的运维或建设经历; 了解 Kubernetes 及生态,理解 Operator、Device Plugin 或 CNI/CSI 插件的工作原理; 熟悉 Prometheus / Grafana / Loki / ELK 等可观测体系,具备基于 eBPF 或 NVIDIA Nsight 进行性能剖析与故障定位的能力; 具备 Go / Python / Rust 至少一门语言的开发能力,能独立完成自动化运维工具、Operator 或监控 Exporter 的开发。 具备良好沟通合作能力和扎实的工程素养。
Similar roles you might like
See all →JJ
Staff Engineer-Chiller SME, APAC
JBE JC Bld Eff Tech (Wuxi) Co Ltd
Salary not disclosedPosted today
C
Senior Software Engineer
checkout.com
Salary not disclosedPosted today
C
Software Engineer
checkout.com
Salary not disclosedPosted today
2B
Principal Systems Software Developer
2000 BlackBerry Limited
Salary not disclosedPosted today
LS
Product Engineer
Littelfuse Semiconductor (Wuxi) Co., Ltd. (China)
Salary not disclosedPosted today
FE
QA Intern
Fa Exhj Saasfaprod1
Salary not disclosedPosted today
BG
Senior Software Architect_HC
Bosch Group
Salary not disclosedPosted 1 day ago
CN
Senior Web Software Engineer, AI Infrastructure
CN05 NVIDIA Shanghai WFOE
Salary not disclosedPosted 4 days ago
This is an external listing. JobSpring does not represent or verify the employer. Report this listing