Skip to content
← Back to job listings

System Engineer – Site Reliability Engineering (SRE)

Ejwl · Bethesda, MD, United States

Software DevelopmentExternal listingfull-timeabout 2 hours ago

About The Role

OB SUMMARY

The Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission-critical cloud and on-prem services that support millions of Marriot customers globally. This role involves overseeing incident management, driving automation efforts, and working closely with cross-functional teams to ensure alignment between SRE strategy and business objectives. Partners closely with Product Teams, Applications teams, Infrastructure, and the broader Applications and Infrastructure Delivery teams to develop key metrics and KPIs to improve applications stability, availability and performance. The ideal candidate will bring strong communication skills, collaborating with key stakeholders across the company to optimize cloud infrastructure and uphold the highest standards of operational excellence in a dynamic, fast-paced environment

CANDIDATE PROFILE

Required Education and Experience

  • Undergraduate degree in an engineering or computer science discipline and/or equivalent experience/certification
  • 5+ years of experience as a Site Reliability Engineer (SRE), building and managing highly available and mission critical systems
  • Expertise in AWS services including designing highly available, multi-AZ and multi-region architectures including:
  • Compute: EC2, Auto Scaling, Lambda
  • Containers: EKS (Mandatory), ECS (good to have)
  • Networking: VPC, subnets, route tables, NAT gateways, Transit Gateway
  • Security: IAM roles/Policies, KMS, Secret manager
  • Storage and Databases: S3, EBS, EFS, RDS, DocumentDB.
  • Deep understanding of SRE practices such as Service Level Objectives, Error Budgets, Toil Management, Observability & Monitoring, Blameless Postmortems, Incident Response Process, Capacity Planning
  • Strong understanding of Cloud Security best practices and responsibility model.
  • Experience driving cloud cost optimization initiatives (rightsizing, reserved instances, autoscaling strategies, cost observability)
  • Proven automation and programming experience in one or more of the following languages: Python, Bash, PowerShell
  • Strong working knowledge of modern, continuous development techniques and pipelines (Agile, Kanban, Jira, CI/CD, Helm, Harness, Jenkins, Git, Artifactory, Vault)
  • Production level expertise with containerization orchestration engines such as Kubernetes (EKS, AKS, ACK)
  • Hands-on experience with service mesh technologies to enable secure and resilient service communication, including mTLS, traffic shaping, and policy enforcement.
  • Strong experience troubleshooting API-related issues in distributed systems, including latency, authentication/authorization failures, rate limiting, and upstream/downstream dependency failures.
  • Ability to analyze API traffic and debug issues using logs, traces, and metrics.
  • Experience with Infrastructure as Code (Iac) tools like Terraform and CloudFormation.
  • Experience with configuration management and automation tools such as Ansible.
  • Deep expertise and hands-on experience with Linux administration (RHEL, Ubuntu, CentOS, AWS Linux)
  • Solid understanding of Virtualization Technologies (VMware vSphere, KVM etc)
  • Strong understanding of networking fundamentals such as Load Balancing, Firewalls, Security Groups, NACLs, TCP/IP, DNS, HTTP/HTTPS, SSL/TLS etc
  • Deep understanding and/or experience with Cloud Native, Relational and NoSQL databases like RDS, MySQL, PostgreSQL, Cassandra or Couchbase
  • Strong experience designing and implementing end-to-end observability solutions across metrics, logs, and traces using tools like Prometheus, Grafana, ELK Stack, and OpenTelemetry.
  • Proven ability to define SLIs/SLOs, build actionable alerting systems, and leverage telemetry data for incident response, root cause analysis, and performance optimization.
  • Experience with deploying, monitoring, and troubleshooting large-scale, distributed applications in cloud environments such as AWS
  • Experience in vulnerability management, OS hardening, patching, security compliance of infrastructure, applications and databases
  • Experience in implementing OS and cloud hardening guidelines and perform regular vulnerability remediation.
  • Familiarity with security frameworks such as ISO27001, SOCII, PCI-DSS, and/or HIPAA
  • 5+ years progressive technology experience
  • 3+ years’ experience in the operational support of critical solutions in large scale environments and organizations with specific experience in
  • Windows Servers
  • RHEL
  • VMWare (vCenter and ESXi Hypervisor)
  • Ability to work outside of normal business hours and days in a 24x7x365 environment.
  • Some travel required

Preferred

  • CI/CD pipeline technologies such as Git, Docker Trusted Registry, Artifactory, Hashicorp Vault, Maven, etc.
  • Security Protocols like SSL, SAML, LDAP etc.
  • AIX Administration
  • Automation and Scripting languages (Ansible, PowerShell, VB, ShellScript)
  • Strong organizational, written and verbal communication skills
  • ITIL v4 certification
  • Windows and Redhat Certification
  • Experience delivering technology solutions in a fast-paced, deadline driven enterprise environment
  • Experience learning and applying new technologies to solve business needs
  • Excellent understanding of change management, testing requirements, techniques, and tools to ensure high availability of systems
  • Experience in researching emerging technologies and trends, standards, and products

CORE WORK ACTIVITIES

  • Provide support for priority incidents as directed by SRE leader
  • Collaborates through the incident with key team members (network, application, etc.) to engage service providers and other stakeholders to identify problem root cause and drive service restoration
  • Provides closure to incidents or initiates a process-based, hand-off to next shift
  • Works with support vendors and providers to ensure proper global coverage, phone support, parts replacement and smart hands
  • Work to create and mature incident response processes
  • Provides oversight to service providers or less experienced engineers
  • Drive the objectives associated with Problem Management; such as customer communication and Root Cause Analysis reports
  • Trains and/or mentors other team members, and peers as appropriate
  • Identifies opportunities to enhance the service delivery, operations and continual service improvement processes
  • Identifies solutions that may contribute to greater stability and reliability of property infrastructure
  • Develop implementation plans, test plans, and timelines for projects and tasks

Delivering Technology

  • Create and enhance administrative, operational and technical policies and procedures, adopting best practice guidelines, standards and procedures for employees, contractors and vendor engagements
  • Maintains a proper balance between business and operational risk
  • Establishes cadence of communication with other organizations to stay abreast of deployment and production support activities
  • Works in a concerted effort with application development and engineering teams to resolve complex issues
  • Provides oversight, collaboration, provisioning, management and maintenance of technology products and service alternatives that improve the production services environment
  • Responsible for the establishment and continuous development of monitoring and alerting for all production environments
  • Contributes to continuous improvement of internal processes
  • Attends training to ensure skillset and tools support the production environments and deliver on project commitments
  • Performs quantitative and qualitative analyses for operational availability to promote a zero-defect environment
  • Facilitates achievement of expected deliverables and obligations of Services Providers
  • Assists operational teams in system updates & upgrades
  • Provides consultation for routine systems development
  • Ensures early warning to the business stakeholder executives regarding degraded or missed service levels

Service Provider Management

  • Actively coordinates with IT service providers and vendors to bring incidents to resolution
  • Manages Service Providers with a focus on continuous service improvement and service restoration
  • Monitors, manages and leads Service Provider outcomes required to ensure operational availability and a zero-defect production environment
  • Consults with internal Service Management & external Service Providers on performance, business reporting, analytics metrics and business value dashboards

Maintaining Goals

  • Submits reports in a timely manner, ensuring delivery deadlines are met.
  • Promotes the documenting of project progress accurately.
  • Provides input and assistance to other teams regarding projects.

Managing Work, Projects, and Policies

  • Manages and implements work and projects as assigned.
  • Generates and provides accurate and timely results in the form of reports, presentations, etc.
  • Analyzes information and evaluates results to choose the best solution and solve problems.
  • Provides timely, accurate, and detailed status reports as requested.

Demonstrating and Applying Discipline Knowledge

  • Provides technical expertise and support to people inside and outside of the department.
  • Demonstrates knowledge of job-relevant issues, products, systems, and processes.
  • Demonstrates knowledge of function-specific procedures.
  • Keeps up-to-date technically and applies new knowledge to job.
  • Uses computers and computer systems (including hardware and software) to enter data and/ or process information.

Delivering on the Needs of Key Stakeholders

  • Understands and meets the needs of key stakeholders.
  • Develops specific goals and plans to prioritize, organize, and accomplish work.
  • Determines priorities, schedules, plans and necessary resources to ensure completion of any projects on schedule.
  • Collaborates with internal partners and stakeholders to support business/initiative strategies
  • Communicates concepts in a clear and persuasive manner that is easy to understand.
  • Generates and provides accurate and prompt results in the form of reports, presentations, etc.
  • Demonstrates an understanding of business priorities

At Marriott International, we are dedicated to being an equal opportunity employer, welcoming all and providing access to opportunity. We actively foster an environment where the unique backgrounds of our associates are valued and celebrated. Our greatest strength lies in the rich blend of culture, talent, and experiences of our associates. We are committed to non-discrimination on any protected basis, including disability, veteran status, or other basis protected by applicable law.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing