
Senior Site Reliability Engineer
jobgether · Brazil
About The Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Brazil.
This role offers the opportunity to design, operate, and continuously improve large-scale, distributed infrastructure in a highly technical <environment.You> will work across cloud, Kubernetes, networking, storage, observability, and Linux systems to deliver resilient and high-performing platforms.The position combines hands-on engineering with automation, intelligent monitoring, and architectural <problem-solving.You> will collaborate closely with internal engineering teams, clients, and AI/ML specialists to ensure infrastructure is reliable and ready for demanding workloads.The role is ideal for an experienced SRE who enjoys solving complex operational challenges and improving systems through <automation.You> will contribute to a culture focused on scalability, reliability, continuous learning, and operational excellence.This is a strong opportunity to make a direct impact on modern cloud and AI/ML infrastructure while expanding your technical expertise.
Accountabilities
- Operate, maintain, and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems to ensure reliability, scalability, and performance.
- Design and implement automation workflows using Go, Python, and Shell scripting to reduce manual operations and improve efficiency.
- Build and maintain monitoring and observability solutions using Prometheus, Grafana, and Loki.
- Diagnose and resolve complex issues involving networking, storage, infrastructure, and system performance.
- Collaborate with AI/ML teams to ensure infrastructure is prepared to support model training, data pipelines, and other demanding workloads.
- Contribute to infrastructure architecture, deployment, automation, and continuous improvement initiatives.
- Participate in on-call rotations and incident response activities to maintain system availability and reliability.
- Lead and contribute to postmortem reviews, identifying root causes and implementing improvements to prevent recurring incidents.
- Work collaboratively with clients and multidisciplinary engineering teams to deliver resilient, high-performing infrastructure solutions.
Requirements
- Strong professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, or a closely related field.
- Hands-on experience with Google Cloud Platform (GCP) and Infrastructure as Code tools such as Terraform.
- Strong knowledge of microservices, containers, Kubernetes, Docker, and modern networking concepts.
- Practical experience with Linux systems administration and troubleshooting.
- Experience with PKI and service mesh technologies, preferably including Istio.
- Strong understanding of SRE principles, with a focus on automation, scalability, availability, observability, and reliability.
- Experience troubleshooting complex infrastructure, networking, storage, and performance problems.
- Strong scripting and automation capabilities using tools such as Python and Shell.
- Experience with Golang is considered an asset.
- Ability to work effectively in fast-paced, problem-solving environments and collaborate with both technical teams and clients.
- Strong ownership, analytical thinking, communication, and continuous-improvement mindset.
- Willingness and ability to participate in an on-call rotation.
Benefits
- Competitive total rewards package.
- Remote-work equipment, including a laptop with your choice of operating system.
- Annual budget to personalize and improve your home-work environment.
- Substantial training and professional development allowance.
- Opportunities to attend training, pursue certifications, and participate in professional development days.
- Annual wellness budget that can be used for activities such as gym memberships, fitness, massages, and other wellness initiatives.
- Generous paid vacation and sick leave.
- Paid day off to volunteer for a charity of your choice.
- Opportunity to collaborate with highly experienced professionals in a global technical environment.
- Support for continuous learning, career development, and technical growth.
- Background-check requirements apply to the successful candidate.
- Reasonable accommodations are available upon request during the selection process.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring