Skip to content
← Back to job listings

Principal Engineer, Service Delivery

Iawmqy · Cyberjaya, Selangor, Malaysia

IT - Network / Systems / DB AdminImported listingfull-timeabout 4 hours ago

About The Role

Senior GenAI & HPC Engineer (Service Delivery Principal Engineer)

Dell Technologies customers rely on our products and services to drive progress. So, we take the service we provide extremely seriously. Service Delivery is all about making sure our technical solutions help clients fulfil their priorities, challenges and initiatives. As trusted advisors, we build in-depth knowledge of what each client wants to achieve. Then we make sure the services delivered by Dell Technologies deliver on all our promises. We also work closely with Sales and Global Services colleagues to develop strategic account growth plans, and to identify and pursue sales opportunities.

What you’ll achieve

We’re seeking a Senior GenAI & HPC Engineer with deep experience in GPU accelerated systems, Linux performance tuning, and benchmarking. This role is highly hands on and customer facing, supporting onsite deployments across the South East Asia/APJ for advanced HPC and GenAI solutions. You will work as a part of a team to help build, integrate, and test some of the world’s largest multi GPU systems, benchmark them using industry standard tools, make suggestions on how to optimize performance, and deliver the next generations of AI/HPC infrastructure.

Join us to do the best work of your career and make a profound social impact as a Senior GenAI & HPC Engineer on our Service Delivery Team in Malaysia.

You will

  • Design and deliver advanced service solutions.
  • Develop automation and monitoring; reduce toil.
  • Conduct root-cause analyses and corrective actions.
  • Prepare handover and operational documentation.

Take the first step towards your dream career

Every Dell Technologies team member brings something unique to the table. Here’s what we are looking for with this role

Essential Requirements

  • 8+ years of related experience
  • Experience in deploying GPU accelerated compute clusters for AI with NVIDIA Base Command Manager especially in a NVL72 environment
  • Experience in implementing large scale GPU networking (more than 100,000 connections)
  • Experience in configuring L2/L3 Leaf/Spine networking using Nvidia Spectrum switches/Cumulus OS
  • Experience in InfiniBand switches
  • Experience in performing L2/L3 networking using Sonic OS (Dell/Nvidia Switches)
  • Strong troubleshooting, problem-solving, and stakeholder management skills
  • Network cabling design will be an added experience
  • Experience in deployment of Air Cooled and/or Liquid cooled racks will be an added advantage
  • Knowledge of Kubernetes, Ubuntu, Openshift will be advantageous
  • High amount of travel across SEA
  • Flexibility to support project activities outside office hours when required

Desirable Requirements

  • Bachelor’s degree in Engineering, Computer Science, or related field
  • Knowledge of Kubernetes, Ubuntu, Openshift will be advantageous

This is an external listing. JobSpring does not represent or verify the employer. Report this listing