Skip to content
← Back to job listings

Site Reliability Engineer

Levi Strauss & Co · DF, Mexico

Software DevelopmentImported listingfull-timeabout 17 hours ago

About The Role

Job Location: Mexico City, Mexico

Calling all originals: At Levi Strauss & Co., you can be yourself — and be part of something bigger.We’rea company of people who like to forge our own path and leave the world better than we found it. Whobelievethat what makes us different makes usstronger.Soadd your voice. Make an impact. Find your fit — and your future.

We'reseekinga curious and drivenSite Reliability Engineerto join our Data & AI Platform Engineering team. In this role,you'llhelp keep our data and AI platforms running reliably, efficiently, and securely — platforms that power decisions across our global retail operations.

You'llwork alongside experienced SREs and engineers tomonitorproduction systems, respond to incidents, reduce operational toil, and build the automation that makes our infrastructure more resilient. This is an excellent opportunity to grow your SRE craft in a fast-paced, collaborative environment on Google Cloud Platform, with exposure to multi-cloud technologies and modern data engineering.

About the Job

Reliability & Incident Response

  • -
  • Monitor production systems usingobservability tooling— dashboards, alerts, and logs — to detect and triage issues before theyimpactend users
  • -
  • Participate inon-call rotations, respond to incidents following established runbooks, and escalate appropriately when needed
  • -
  • Contribute toblameless post-mortems, documenting rootcausesand follow-up action items to prevent recurrence
  • -
  • Helpmaintainand improveSLO dashboards and alerting thresholdsto ensure platform health is visible and measurable

Toil Reduction & Automation

  • -
  • Identifyrepetitive manual tasks and buildautomation toeliminatethem, reducing toil for yourself and the broader team
  • -
  • Write andmaintainscripts, tooling, andCI/CD pipeline componentsthat improve deployment reliability and operational efficiency
  • -
  • Supportself-serve infrastructure initiativesthat allow engineering teams to safely provision and manage their own resources

Platform Operations & Cloud Infrastructure

  • -
  • Operate andmaintainworkloads running onGCP— including GKE, Cloud Run,BigQuery, Pub/Sub, GCS, and Composer
  • -
  • ApplyInfrastructure-as-Codepractices (Terraform, Helm) to consistently and safely manage and version infrastructure changes
  • -
  • Supportmulti-cloud awarenessacross GCP and Azure, following team standards for consistency and security across environments
  • -
  • Adhere todata security and governancepolicies — IAM best practices, secrets management, encryption, and audit logging

Collaboration & Growth

  • -
  • Work closely with Data Engineering, AI Platform, and Software Engineering teams to ensure reliability is considered from design through deployment
  • -
  • Participate inreliability reviews, design discussions, and team ceremonies, contributing ideas and raising operational concerns early
  • -
  • Engage withAI and agentic platform workloads, gaining exposure to the operational patterns of LLM-based systems and data pipelines
  • -
  • Continuously develop your technical skills and SRE craft, supported by team knowledge-sharing, documentation, and hands-on experience

About You

Required Qualifications

  • -
  • Bachelor's degreein Computer Science, Engineering, or related field (or equivalent practical experience)
  • -
  • 6+ years of experiencein Site Reliability Engineering, DevOps, or Platform/Infrastructure Engineering in production environments
  • -
  • Hands-on experience with GCP services— particularly GKE, Cloud Run,BigQuery, Pub/Sub, and GCS
  • -
  • WorkingproficiencywithInfrastructure-as-Codetools such as Terraform or Helm
  • -
  • Familiarity withobservability tooling— metrics, logging, tracing, and alerting (e.g., Cloud Monitoring, Datadog, or Prometheus/Grafana)
  • -
  • Understanding ofSLO/SLI conceptsand how they relate to production reliability and on-call operations
  • -
  • Exposure todata security fundamentals: IAM, encryption, secrets management, and network policies
  • -
  • Proficiencyin at least one scripting or systems language (Python, Bash, or Go) for automation and operational tooling
  • -
  • Strong communicationskills with the ability to clearly document incidents, runbooks, and technical processes

Technical Familiarity

  • -
  • Experience withcontainer orchestration— Kubernetes or GKE — and the operational patterns around deploying and managing containerized workloads
  • -
  • Basic understanding ofCI/CD pipelinesandGitOpsworkflows (ArgoCD, GitHub Actions, or similar)
  • -
  • Comfort working withdata platforms— familiarity with batch or streaming data pipelines is a plus
  • -
  • Awareness ofmulti-cloud concepts, particularly across GCP and Azure

Desirable Experience

  • -
  • Experience working inretail, e-commerce, or consumer goodsenvironments
  • -
  • Familiarity withGoogle's SRE principles— error budgets, toil tracking, and production readiness reviews
  • -
  • Exposure toAI or ML platform operations, including monitoring model serving infrastructure
  • -
  • Experience withFinOpsor cloud cost visibility tooling
  • Why Join Us?
  • Ifyou'rean engineer who is passionate about reliability, loves solving operational problems, and wants to grow your SRE craft at a global iconic brand,we'dlove to hear from you.

LOCATION

Mexico, D.F., Mexico

FULL TIME/PART TIME

  • Full time
  • Current LS&Co Employees, apply via your Workday account.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing