Skip to content
← Back to job listings

Stellar - Director of SRE

deCircle · New York, United States

Software DevelopmentImported listingfull-timeabout 2 hours ago

About The Role

About Stellar Development Foundation
The Stellar Development Foundation (SDF) is a mission-driven organization supporting the development and growth of the Stellar blockchain network, an open-source platform designed to expand access to the global financial system.
Since 2014, Stellar has grown into a global blockchain ecosystem used by developers and companies building financial applications and infrastructure around the world.
SDF is now looking for a Director of Site Reliability Engineering to lead its SRE function and shape how engineering teams own, operate and improve production services.
The Role
This is a senior engineering leadership position reporting directly to the CTO.
You’ll lead a small, high-leverage SRE team while defining the broader SRE vision, operating model and reliability culture across engineering.
Rather than SRE acting as the operational owner of every production system, engineering teams at SDF own the services they build. Your role will be to create the infrastructure, frameworks, tooling, standards and observability practices that allow those teams to operate their services reliably and independently.
You’ll combine hands-on technical judgment with organizational leadership, helping SDF improve reliability, infrastructure maturity and developer productivity without introducing unnecessary process or complexity.
What You’ll Work On
You’ll lead, coach and develop a distributed SRE team while establishing its charter, priorities, operating model and measures of success.
A major focus will be defining and rolling out a Service Ownership & Maturity Framework, establishing appropriate reliability and operational standards based on the criticality of individual services.
You’ll own and evolve core engineering infrastructure across:
-
Cloud infrastructure and foundations
-
Kubernetes and containerized compute
-
CI/CD and deployment infrastructure
-
Observability, monitoring and alerting
-
Secrets and access management
-
GitHub workflows
-
Infrastructure-as-code and automation
You’ll help engineering teams become stronger owners of their production services through better dashboards, runbooks, alerting, escalation paths, deployment practices and operational readiness.
You’ll also improve deployment automation, resilience, self-healing systems, disaster recovery and service reliability, prioritizing improvements based on real operational risk and impact.
Another important part of the role will be evolving incident response, postmortems, escalation and on-call practices across a geographically distributed engineering organization.
You’ll build paved paths and self-service infrastructure that reduce engineering toil and cognitive load while allowing teams to ship faster without compromising reliability.
The role also works closely with Security, Compliance, Legal, Finance, Procurement and Corporate IT wherever cloud infrastructure, access management, vendors or operational controls intersect with engineering.
SDF is also interested in pragmatically exploring AI-assisted and agentic workflows where they can improve infrastructure operations, observability, developer productivity and service ownership.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing