Skip to content
← Back to job listings

Manager of Technical Support Engineering (Bare Metal)

CoreWeave · Sunnyvale, United States

RemoteImported listingfull-time19 days ago

About The Role

Join CoreWeave, a leading cloud provider dedicated to powering the AI revolution. As the Manager of Technical Support Engineering, you will lead a skilled team responsible for maintaining and optimizing physical infrastructure across multiple client environments. You will oversee the resolution of infrastructure-related incidents, improve support processes, and work closely with product and infrastructure teams. This role requires a strong background in infrastructure support, data center operations, and Linux system administration.

  • Lead daily support operations, triage incidents, drive escalations, and ensure that hardware is monitored, maintained, and delivered effectively for clients.
  • Oversee the resolution of infrastructure-related incidents, escalation management, and collaborate with internal teams to deliver effective solutions.
  • Build, develop, and lead a dedicated Infrastructure Support team focused on supporting key infrastructure, handling escalations, and ensuring smooth hardware operations.
  • Experience working with high-performance rack-scale hardware, including CPU and GPU-based compute nodes
  • Familiarity with hardware-level diagnostics, troubleshooting, and replacement (servers, power, cabling, etc.)
  • Experience managing ticket-based workflows (Jira, Zendesk, etc.) in a high-urgency technical environment
  • Understanding of GPU infrastructure (e.g., NVIDIA A100/H100s, PCIe/NVLink, liquid cooling) or a demonstrated ability to quickly learn and adapt to HPC environments
  • Proven track record in incident and escalation management, with direct ownership of client or production-impacting issues
  • Travel up to 30% annually
  • 5+ years of experience leading teams responsible for infrastructure support, data center operations, or physical compute environments
  • Hands-on experience with Linux system administration and command-line tools
  • Skilled in managing scheduling, shift coverage, and team logistics in 24/7 or hybrid support models
  • Comfortable interpreting and acting on metrics (MTTR, SLOs, backlog, ticket trends) to drive operational improvements
  • Have a track record of improving infrastructure reliability through clear processes and team accountability
  • Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk
  • Thrive in fast-paced environments where priorities can shift quickly
  • Think critically about how to scale operations without overengineering
  • Are comfortable getting close to the work, but know when to step back and lead
  • Communicate clearly, especially under pressure
  • Care about delivering for customers—but know when to hold the line to protect the team and long-term goals
  • Experience managing infrastructure support teams in high-growth or rapidly evolving environments
  • Proven ability to develop and implement operational processes that scale with business needs
  • Strong familiarity with server and GPU hardware lifecycle management: deployment, maintenance, thermal/power concerns, RMA coordination, and decommissioning
  • Demonstrated success in coaching and growing technical teams through training, mentorship, and performance development
  • Familiarity with AI/ML workloads, cluster utilization patterns, or the infrastructure needs of GPU-heavy clients is a plus
  • Skilled in both developing and interpreting metrics to drive accountability, continuous improvement, and executive visibility

This is an external listing. JobSpring does not represent or verify the employer. Report this listing