Senior Platform Reliability Engineer
firmus · Melbourne, Victoria, Australia
About The Role
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
AI FactoryOS Operations
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.
Role Summary
The Senior Platform Reliability Engineer is part of the team that operates Firmus AI FactoryOS in production: the GPU compute fleet, and the platform services it depends on, including exabyte-scale storage, the shared core services, the virtualisation hosting the management plane, and the observability infrastructure the estate is measured through. This is state-of-the-art AI infrastructure, among the largest deployments in Asia Pacific, built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation.
This is a hands-on senior role with deep technical expertise, working in a team that shares accountability for the compute fleet and the platform services it depends on. The team runs those to a declared service level and sets the acceptance requirements each service has to meet before it goes live. The team also builds the shared administrative infrastructure the estate is run from, in consultation with the AI Infrastructure team, and operates it as a shared service. Automation is a first-class part of this role: the team builds and maintains the guarded automation and remediation tooling that turns manual response into a self-healing capability.
Key Responsibilities
Operate the multi-tenant control and management plane that Firmus' AI and infrastructure services depend on, including the tenancy, quota and access controls that support separation between tenant workloads and data.
Operate Firmus' exabyte-scale distributed and high-performance filesystems (for example VAST, WEKA, Ceph) and S3-compatible object storage to their declared service levels, own the operational automation around them, and drive continuous improvement in how they are operated, feeding platform improvement requirements to AI Infrastructure with evidence.
Operate the GPU compute fleet in production: node health and readiness, GPU and node fault detection and handling, firmware and driver currency to the supported baselines, remediation and return-to-service, and the hardware fault and replacement workflow with vendors and site operations.
Build and maintain the guarded automation and remediation tooling that turns manual response into a self-healing capability, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour.
Diagnose and tune performance across the full data path, applying a deep understanding of operating systems, computer networks and software-defined storage, and working with technologies including RDMA, GPU Direct Storage, RoCE and InfiniBand.
Deliver change as code, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
Operate the observability infrastructure as a shared service across metrics, logs, traces, alerting and retention, and work with service owners so that telemetry becomes alerting that is actionable and tied to a runbook.
Operate each service in the portfolio to its published service level, carry new services through production readiness review, and execute the monthly patching cycle and urgent vulnerability remediation.
Provide the deepest technical expertise for these platforms, taking on and diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure, and manage vendor escalations at engineering level.
Lead technical recovery during major incidents, drive the changes that remove repeat causes, share the follow-the-sun on-call roster, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results.
Skills & Experience
Required Skills
Strong skills in infrastructure, systems and platform engineering, with 8+ years of experience including substantial ownership of production storage and shared infrastructure services in a 24/7 environment.
Extensive experience with scale-out, parallel or enterprise storage supporting demanding workloads (for example VAST, WEKA, Ceph, Lustre, GPFS or NetApp), including diagnosis of performance and capacity problems across the full data path using evidence rather than assumption.
Substantial experience operating a virtualisation platform (for example Proxmox, VMware or KVM) and building shared services such as databases and object storage to a defined service level.
Expert-level knowledge of Linux systems, including storage and file system internals, kernel and driver behaviour, networking, memory and I/O subsystems, and systematic performance analysis.
Extensive experience operating observability infrastructure as a service across metrics, logs, traces and alerting (for example Prometheus, Grafana, OpenTelemetry, Loki or Elasticsearch), and designing alerting that is actionable.
Strong skills in infrastructure automation, infrastructure-as-code and GitOps practices (for example OpenTofu or Terraform, Ansible, Argo CD, CircleCI), with change delivered through peer review, automated testing and progressive rollout.
<li data-leveltext="" data-font="Symbol" data-listid="16" data-list-defn-props="{"335552541":1,"335559685":720,"335559991":360,"469769226":"Symbol","469769242":[8226],"469777803":"left","469777804":"","469777815":"multilevel"}" data-aria-posins
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
