Skip to content
← Back to job listings

Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Uvation · Romania

IT - Network / Systems / DB AdminRemoteExternal listingcontract1 day ago

About The Role

Job Overview

We are seeking a highly experienced

Senior Linux Infrastructure Engineer

with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation

AI Factory / GPU infrastructure platforms

. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is

not a DevOps-focused role

. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in

Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms

.

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in
  • bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating
  • Bare Metal as a Service (BMaaS)
  • platforms and large-scale infrastructure environments

Strong understanding of server hardware, including

BIOS/UEFI

  • RAID controllers
  • Firmware management
  • iLO/iDRAC/IPMI

NICs and SmartNICs

  • HBA cards
  • Hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

  • Experience deploying and managing
  • GPU-accelerated infrastructure
  • for AI/ML workloads

Understanding of NVIDIA GPU technologies including

  • A100, H100, H200, B200, or equivalent GPU platforms
  • NVIDIA DGX and OEM GPU servers
  • GPU provisioning and lifecycle management
  • GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads

Understanding of

  • GPU resource allocation and scheduling
  • Multi-GPU systems
  • GPU networking requirements
  • High-bandwidth, low-latency infrastructure design

Familiarity with NVIDIA ecosystem technologies such as

CUDA

NCCL

GPUDirect Storage

NVIDIA Fabric Manager

NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

Advanced Linux storage administration

LVM

XFS, EXT4

NFS

iSCSI

Fibre Channel SAN

Multipath I/O

Strong hands-on experience with

Ceph

, including

Cluster architecture

MON, OSD, MDS

RBD, CephFS, RGW

  • Capacity planning
  • Performance tuning
  • Failure recovery

Experience with high-performance AI storage platforms such as

WEKA

VAST Data

Dell PowerScale

Pure Storage FlashBlade

NetApp

Understanding of

NVMe-over-Fabrics (NVMe-oF)

RDMA

GPUDirect Storage

  • Parallel file systems
  • AI data pipelines

Networking & Infrastructure

Strong networking knowledge

Bonding

VLANs

Routing

MTU optimization

DNS

DHCP

Experience with high-performance data center networking

100G/200G/400G Ethernet

RoCE

RDMA

  • Spine-Leaf architectures
  • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

Experience with high availability, clustering, and disaster recovery

Strong troubleshooting skills across

  • Linux operating systems
  • Hardware platforms
  • GPU infrastructure

Networking

  • Enterprise storage
  • Experience supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Experience creating operational documentation, runbooks, and infrastructure standards

Nice to Have

  • Kubernetes infrastructure (especially AI/ML and GPU integration)
  • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
  • Ansible automation

NVIDIA Base Command Manager

  • Slurm or HPC workload schedulers
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
  • Data Center Infrastructure Management (DCIM) tools
  • IPAM solutions
  • AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

  • Candidates whose experience is primarily CI/CD pipeline engineering
  • Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
  • Cloud-only administrators with limited bare metal, storage, or hardware experience
  • Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing