
Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program
Sandisk · Kfar Saba, Center District, Israel
About The Role
Hive is a swarm of hundreds of identical storage nodes (Marvell/XSight DPU + SSD), each running a full Ceph plane: OSD, Monitor (MON), Manager (MGR), MDS, plus SeaStore and the NFS front end. Standing up, re-configuring, and recovering a cluster of this size by hand does not scale. We need an engineer who owns the automated hardware discovery and Ceph role-assignment pipeline : the flow that inventories every node and its hardware, decides which daemons each node should run, and drives the cluster from bare metal to a serving state. The flow must run fully autonomously by default, with a clean manual-override path for lab, bring-up, and failure-injection scenarios.
This is a specialized automation and lab-hardware-discovery discipline that the Hive team does not currently have dedicated ownership for. It is directly on the critical path for every test cluster, every silicon bring-up, and every customer-shaped deployment.
What you'll own
- Automated hardware discovery. Detect nodes as they power on and enumerate their hardware (DPU/SoC model, SSD media, DRAM, RNIC/network ports, BMC) using out-of-band and in-band inventory (BMC/Redfish/IPMI, PXE/DHCP boot, cloud-init/first-boot agents). Produce a single authoritative machine inventory that the rest of the pipeline consumes.
- Role assignment and cluster composition. Given the discovered inventory, decide and apply which Ceph roles each node runs — OSD, MON, MGR, MDS (and the NFS gateway) — encoded as declarative placement specifications . Own MON quorum sizing and placement, MGR redundancy, MDS/gateway placement, and CRUSH map / failure-domain layout so data and metadata land correctly across the swarm.
- Bare-metal → serving bring-up. Drive the end-to-end sequence: node provisioning, OS/image deploy, cluster bootstrap , daemon deployment via the Ceph orchestrator (cephadm-style) , ceph-volume-style OSD provisioning on the SSD media, and health convergence to HEALTH_OK.
- Autonomous and manual modes. Make the default path zero-touch (a rack powers on and self-assembles into a healthy cluster), while exposing deterministic manual controls to pin roles, hold a node out, force a specific topology, or reproduce a customer/lab configuration for testing.
- Lifecycle and recovery automation. Node add/remove, drain and rebalance , daemon replacement, MON re-quorum after loss, MDS/OSD failover validation, and re-discovery after re-imaging — integrated with Hive's Fast Recovery + BMC work.
- Reconciliation and drift control. Continuously compare declared desired state against observed cluster state and converge — the same idea Ceph's orchestrator applies to service specs — with clear reporting when reality diverges from intent.
- CI/lab integration. Wire the pipeline into automated test so any commit can spin up a correctly-composed multi-node cluster on real hardware and tear it down cleanly.
Core (must-have)
- Deep experience in lab hardware discovery, inventory, and automated provisioning at fleet scale.
- Bare-metal automation: PXE/DHCP boot, BMC out-of-band management (Redfish/IPMI), image/OS deployment, cloud-init/first-boot .
- Infrastructure-as-code and config automation (e.g., Ansible/Terraform-class tooling) with a declarative, reconcile-to-desired-state mindset.
- Strong scripting/automation (Python and shell) and CI systems (Jenkins/GitLab CI or equivalent).
- Comfort designing systems that are autonomous by default with reliable manual override .
Strongly preferred (or ramp-up expected)
- Ceph operational knowledge: OSD/MON/MGR/MDS roles, the orchestrator/cephadm model, placement specs , ceph-volume, CRUSH maps and failure domains , MON quorum, and cluster health/lifecycle.
- Distributed-systems fluency: quorum/Paxos intuition, rebalance/recovery behavior, failure-domain reasoning.
- Networking bring-up for storage fabrics (RoCEv2/Ethernet, port discovery).
- Familiarity with DPU/SoC-based nodes and constrained-node environments.
Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.
Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at [email hidden] to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.
 
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
