Skip to content
← Back to job listings

Observability Engineer / Site Reliability Engineer

Ontrac Solutions · Chicago, IL

RemoteExternal listingcontract30 days ago

About The Role

About Ontrac Solutions

Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible

results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.

Role Overview

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.

Key Responsibilities

GCP & Cloud Management

Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).

Platform Management

Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.

Automation & IaC

Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.

CI/CD Integration

Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.

System Performance

Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.

SRE Evangelism

Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.

Required Skills & Qualifications

Cloud Infrastructure

Proven engineering experience within

Google Cloud Platform (GCP)

environments, particularly managing cloud-native monitoring and compute resources.

Observability Tooling

Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with

GEM (Grafana Enterprise Metrics)

is highly desirable.

OS & Scripting

Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.

Programming

Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).

Containers & Orchestration

  • Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.
  • __________________________________
  • Ontrac Solutions has partnered with

PinpointVerify

to help genuine applicants rise above the noise. Today, qualified candidates are too often overshadowed by fake and fraudulent applications. PinpointVerify gives our recruiters confidence that you are exactly who you say you are — and gives you a

portable verification credential

you can share with any employer.

Applicants who complete verification are

prioritized over non-verified candidates

with comparable experience.

And if you're hired, Ontrac reimburses the full cost of your verification.

Get verified →

<https://pinpointverify.com/ontrac>

This is an external listing. JobSpring does not represent or verify the employer. Report this listing