Principal Engineer, Service Delivery
Iawmqy · Cyberjaya, Selangor, Malaysia
About The Role
Senior GenAI & HPC Engineer (Service Delivery Principal Engineer)
Dell Technologies customers rely on our products and services to drive progress. So, we take the service we provide extremely seriously. Service Delivery is all about making sure our technical solutions help clients fulfil their priorities, challenges and initiatives. As trusted advisors, we build in-depth knowledge of what each client wants to achieve. Then we make sure the services delivered by Dell Technologies deliver on all our promises. We also work closely with Sales and Global Services colleagues to develop strategic account growth plans, and to identify and pursue sales opportunities.
What you’ll achieve
We’re seeking a Senior GenAI & HPC Engineer with deep experience in GPU accelerated systems, Linux performance tuning, and benchmarking. This role is highly hands on and customer facing, supporting onsite deployments across the South East Asia/APJ for advanced HPC and GenAI solutions. You will work as a part of a team to help build, integrate, and test some of the world’s largest multi GPU systems, benchmark them using industry standard tools, make suggestions on how to optimize performance, and deliver the next generations of AI/HPC infrastructure.
Join us to do the best work of your career and make a profound social impact as a Senior GenAI & HPC Engineer on our Service Delivery Team in Malaysia.
You will
- Design and deliver advanced service solutions.
- Develop automation and monitoring; reduce toil.
- Conduct root-cause analyses and corrective actions.
- Prepare handover and operational documentation.
Take the first step towards your dream career
Every Dell Technologies team member brings something unique to the table. Here’s what we are looking for with this role
Essential Requirements
- 8+ years of related experience
- Experience in deploying GPU accelerated compute clusters for AI with NVIDIA Base Command Manager especially in a NVL72 environment
- Experience in implementing large scale GPU networking (more than 100,000 connections)
- Experience in configuring L2/L3 Leaf/Spine networking using Nvidia Spectrum switches/Cumulus OS
- Experience in InfiniBand switches
- Experience in performing L2/L3 networking using Sonic OS (Dell/Nvidia Switches)
- Strong troubleshooting, problem-solving, and stakeholder management skills
- Network cabling design will be an added experience
- Experience in deployment of Air Cooled and/or Liquid cooled racks will be an added advantage
- Knowledge of Kubernetes, Ubuntu, Openshift will be advantageous
- High amount of travel across SEA
- Flexibility to support project activities outside office hours when required
Desirable Requirements
- Bachelor’s degree in Engineering, Computer Science, or related field
- Knowledge of Kubernetes, Ubuntu, Openshift will be advantageous
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
