Staff Site Reliability Engineer
levistraussandco · Remote, Spain
About The Role
Job Location: Spain
Calling all originals: At Levi Strauss & Co., you can be yourself — and be part of something bigger.We’rea company of people who like to forge our own path and leave the world better than we found it. Whobelievethat what makes us different makes usstronger.Soadd your voice. Make an impact. Find your fit — and your future.
We're seeking an exceptionalStaff Site Reliability Engineer to join our Data & AI Platform Engineering team. In this role, you'll own and elevate the reliability, scalability, and operability of our enterprise data and AI platforms — the platforms that power everything from the design of our iconic jeans to the optimization of our global retail and supply chain.
As a hands-on technical leader, you'll embody the principles of Google's SRE discipline: eliminating toil, engineering for reliability, and building a culture of shared ownership between development and operations. This is a unique opportunity to shape how a legendary brand runs production at scale on Google Cloud Platform, with a growing multi-cloud footprint across GCP and Azure.
About the Job
Reliability & Incident Management
- -
- Define, instrument, and enforceSLOs, SLIs, and error budgets across all platform services, ensuring alignment with business and product commitments
- -
- Drive continuous reduction inMTTD and MTTR through improved observability, automated alerting, and runbook-driven incident response
- -
- Leadblameless post-mortems and translate findings into durable reliability improvements, ensuring systemic issues are eliminated rather than patched
Toil Reduction & Automation
- -
- Systematically identify, measure, and eliminate operational toil; track toil percentage per sprint and enforce guardrails to keep it below 50% of engineering capacity
- -
- Build and maintainself-serve infrastructure capabilities — enabling product and data engineering teams to provision, scale, and operate their own resources safely and consistently
- -
- Automate deployment pipelines, configuration management, and operational workflows usingInfrastructure-as-Code principles (Terraform, Helm, GitOps)
Platform Engineering & Architecture
- -
- Serve as the primaryGCP subject matter expert — architecting and optimizing workloads across GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
- -
- Leadmulti-cloud architecture decisions across GCP and Azure, ensuring consistent security posture, cost efficiency, and operational practices across environments
- -
- Design and implementself-healing infrastructure patterns, auto-scaling strategies, and capacity planning models to support high-availability data and AI platforms
- -
- Championdata security and governance best practices — including encryption at rest and in transit, IAM least-privilege, secrets management, and audit logging
AI, Agentic Systems & Modern Observability
-
Apply SRE principles toagentic AI workloads — defining reliability expectations for LLM-based and multi-agent systems, including latency SLOs, fallback patterns, and model observability
-
Partner with AI Platform teams to productionize agentic pipelines with robust monitoring, drift detection, and rollback capabilities
-
Drive adoption ofAI-assisted operations tooling to enhance observability, anomaly detection, and predictive incident management
Leadership & Culture
- -
- Guide and mentor junior and mid-level SREs — conducting code reviews, running reliability reviews, and elevating the team's engineering craft
- -
- Collaborate cross-functionally with Data Engineering, Software Engineering, Security, and Product teams to embed reliability as a shared value from design through deployment
- -
- Champion a culture ofpsychological safety, continuous learning, and reliability excellence
- -
- Communicate platform health, risk posture, and reliability roadmaps clearly to both technical and executive audiences
About You
Required Qualifications
- -
- Master's degree in Computer Science, Engineering, or related field (or equivalent practical experience)
- -
- 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a strong track record in large-scale production environments
- -
- Deep, hands-on expertise in GCP — including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
- -
- Proficiency withInfrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux)
- -
- Strong command ofobservability tooling — distributed tracing, structured logging, metrics pipelines, and alerting platforms (e.g., Cloud Monitoring, Datadog, Prometheus/Grafana)
- -
- Proven experience defining and operating againstSLOs, SLIs, and error budgets in production environments
- -
- Solid understanding ofdata security principles: IAM, encryption, secrets management, network policies, and compliance frameworks
- -
- Experience withmulti-cloud environments (GCP + Azure), including cross-cloud networking, identity federation, and cost governance
- -
- Demonstrated ability tolead without authority — influencing engineers across teams and driving reliability improvements at the organizational level
- -
- Excellent written and verbal communication skills; ability to translate complex reliability concepts for non-technical stakeholders
Technical Depth
- -
- Fluency in at least one systems or scripting language (Python, Go, or Bash) for automation and tooling
- -
- Experience withcontainer orchestration (Kubernetes/GKE), service mesh, and traffic management patterns
- -
- Familiarity withdata engineering patterns: batch and streaming pipelines, data warehouses, and the operational challenges of large-scale data platforms
- -
- Understanding ofagentic AI architectures and the unique reliability challenges of LLM-based, event-driven, and multi-agent systems
- -
- Working knowledge ofdata governance frameworks, data lineage tooling, and platform-level data quality enforcement
Desirable Experience
- -
- Experience operating data platforms inretail or e-commerce environments
- -
- Familiarity withSRE principles in practice — error budget policies, CRE engagements, production readiness reviews
- -
- Exposure toFinOps practices — cloud cost attribution, commitment optimization, and unit economics for data workloads
- -
- Experience withAzure-native services (AKS, Azure Data Factory, Event Hubs) and cross-cloud identity management
- -
- Prior experience in aStaff or Principal-level SRE role with organization-wide scope
LOCATION
Spain - Remote
FULL TIME/PART TIME
- Full time
- Current LS&Co Employees, apply via your Workday account.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
