Site Reliability Engineer II (NGPOS Operations Support)
hudsonmanpower (recruitee) · Cincinnati, OH, United States
About The Role
Position Overview
We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support .
The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.
This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements.
Location: Blue Ash, OH (Cincinnati) – Onsite (5 Days/Week)
Employment Type: W2 – Contract-to-Hire
Duration: Full-Time
Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship.
Experience Required: 3+ Years
Nice to Have
- Enterprise Point of Sale (POS) systems
- Retail technology experience
- Automation scripting
- Monitoring optimization
- Runbook creation
- Store technology deployments
Key Responsibilities
- Lead major incident response during production outages
- Serve as Incident Commander during P1/P2 incidents
- Coordinate technical bridge calls
- Communicate outage status to engineering teams and business leadership
- Lead Root Cause Analysis (RCA) activities
- Track corrective actions through completion
- Improve production reliability and system stability
- Enhance monitoring and observability
- Reduce alert fatigue
- Partner with Software Engineering and Platform Engineering teams
- Support retail store deployments
- Develop operational documentation, runbooks, and playbooks
- Participate in after-hours support rotations and maintenance windows
- Improve service health using SLIs and SLOs
Technical Environment
Monitoring & Observability
- Dynatrace
- Azure Monitor
- Log Analytics
- Metrics
- Dashboards
Cloud
- Microsoft Azure
- Google Cloud Platform (GCP)
Containers
- Kubernetes
- Docker
Operating Systems
- Linux
Scripting Languages
- Bash
- Python
Agile Tools
- Jira
Enterprise Environment
- Retail systems
- Point of Sale (POS)
- Production Support
- Hybrid Infrastructure
Ideal Candidate Profile
The ideal candidate will demonstrate
- Strong leadership during production incidents
- Excellent troubleshooting and analytical skills
- Effective communication under pressure
- Ownership and accountability
- Experience coordinating multiple engineering teams
- Strong operational discipline
- Continuous improvement mindset
- Passion for reliability engineering
- Excellent documentation skills
This listing was posted by a verified recruiter at hudsonmanpower (recruitee). Report this listing
JobSpring