Skip to content
← Back to job listings

Site Reliability Engineer II (NGPOS Operations Support)

hudsonmanpower (recruitee) · Cincinnati, OH, United States

Software DevelopmentQuick applycontract23 days ago

About The Role

Position Overview

We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support .

The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.

This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements.

Location: Blue Ash, OH (Cincinnati) – Onsite (5 Days/Week)

Employment Type: W2 – Contract-to-Hire

Duration: Full-Time

Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship.

Experience Required: 3+ Years

Nice to Have

  • Enterprise Point of Sale (POS) systems
  • Retail technology experience
  • Automation scripting
  • Monitoring optimization
  • Runbook creation
  • Store technology deployments

Key Responsibilities

  • Lead major incident response during production outages
  • Serve as Incident Commander during P1/P2 incidents
  • Coordinate technical bridge calls
  • Communicate outage status to engineering teams and business leadership
  • Lead Root Cause Analysis (RCA) activities
  • Track corrective actions through completion
  • Improve production reliability and system stability
  • Enhance monitoring and observability
  • Reduce alert fatigue
  • Partner with Software Engineering and Platform Engineering teams
  • Support retail store deployments
  • Develop operational documentation, runbooks, and playbooks
  • Participate in after-hours support rotations and maintenance windows
  • Improve service health using SLIs and SLOs

Technical Environment

Monitoring & Observability

  • Dynatrace
  • Azure Monitor
  • Log Analytics
  • Metrics
  • Dashboards

Cloud

  • Microsoft Azure
  • Google Cloud Platform (GCP)

Containers

  • Kubernetes
  • Docker

Operating Systems

  • Linux

Scripting Languages

  • Bash
  • Python

Agile Tools

  • Jira

Enterprise Environment

  • Retail systems
  • Point of Sale (POS)
  • Production Support
  • Hybrid Infrastructure

Ideal Candidate Profile

The ideal candidate will demonstrate

  • Strong leadership during production incidents
  • Excellent troubleshooting and analytical skills
  • Effective communication under pressure
  • Ownership and accountability
  • Experience coordinating multiple engineering teams
  • Strong operational discipline
  • Continuous improvement mindset
  • Passion for reliability engineering
  • Excellent documentation skills

This listing was posted by a verified recruiter at hudsonmanpower (recruitee). Report this listing