Senior Technical Program Manager (Fleet Operations & Reliability)
CoreWeave · Sunnyvale, United States
About The Role
Join CoreWeave, the world's largest AI Cloud platform, as a Senior Technical Program Manager. In this role, you will be responsible for fleet-wide reliability and operations programs, driving improvements in fleet operating efficiency and ensuring reliability and stability goals are met. You will work closely with various teams, establish and own fleet reliability metrics, identify systemic reliability gaps, and lead complex cross-functional programs. The ideal candidate will have a bachelor's degree in a related technical field, 7+ years of technical program management experience, and a strong technical aptitude across infrastructure domains.
- Act as the central owner for fleet reliability and operations programs, aligning teams on priorities, defining success metrics, driving execution, and ensuring improvements hold at scale.
- Establish and own fleet reliability metrics and dashboards, driving alignment on fleet reliability OKRs across engineering and operations teams.
- Lead complex, cross-functional programs to improve fleet delivery, readiness, and operational stability that scale with fleet growth, avoiding solutions that only work at current size.
- Bachelor's degree in Computer Engineering, Computer Science, or a related technical field, or equivalent practical experience
- Experience operating in a rapidly growing or fast-scaling infrastructure environment, with programs designed to hold up under 2x, 5x, or greater fleet growth
- Demonstrated ability to use data and metrics to drive prioritization, execution, and decision-making
- Excellent communication and stakeholder management skills, including executive-level reporting
- Strong technical aptitude across infrastructure domains (compute, storage, networking, hardware, or SRE)
- 7+ years of technical program management experience in large-scale compute infrastructure, cloud, or platform environments
- Experience operating at scale in data center, cloud infrastructure, or hyperscale environments balancing reliability investments against capacity goals, with an understanding of how operational decisions affect sellable capacity
- Understanding of hardware failure and processes, including failure diagnosis, vendor coordination, and replacement lifecycle management
- Familiarity with reliability frameworks such as SLIs/SLOs, error budgets, incident management practices, and root cause analysis
- Background in observability, monitoring, or telemetry systems (e.g., Prometheus, Grafana, OpenTelemetry)
- Experience with fleet lifecycle management (provisioning, firmware/OS updates, decommissioning) at scale
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing