
Site Reliability / Gitops Engineer
jobgether · Ireland
About The Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability / Gitops Engineer based in Ireland.
Join a global Information Systems team responsible for operating and evolving critical production services at significant <scale.In> this role, you’ll use automation, Infrastructure as Code, and GitOps practices to make cloud operations more reliable, consistent, and <efficient.You>’ll work across private and public cloud environments, strengthening infrastructure resilience, scalability, observability, and performance.Your expertise will also influence the evolution of open-source infrastructure technologies through hands-on feedback, bug reporting, and <collaboration.You>’ll troubleshoot complex systems, improve operational processes, and help eliminate repetitive manual work through thoughtful automation.Working with a distributed team of experienced SREs, you’ll have opportunities to share knowledge, mentor colleagues, and contribute to major engineering initiatives.This is an ideal opportunity for an automation-first technologist who is passionate about Linux, open source, and building robust systems at scale.
Accountabilities
- Apply Infrastructure as Code expertise to continuously improve automation practices, processes, and operational consistency.
- Automate software operations across private and public clouds while accounting for the complexities of distributed systems.
- Develop new capabilities and improve the resilience, scalability, and reliability of cloud and container infrastructure.
- Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliable day-to-day performance.
- Troubleshoot complex systems, perform capacity planning and performance investigations, and develop strong operational expertise.
- Set up, maintain, and use observability and monitoring solutions such as Prometheus, Grafana, and Elasticsearch.
- Design and maintain monitoring and alerting for critical systems and services.
- Collaborate with development teams on service architecture, documentation, playbooks, policies, and operational procedures.
- Work closely with globally distributed engineering, operations, and support teams to resolve issues and improve services.
- Dedicate focused development time to larger engineering projects and the automation of repetitive manual processes.
- Share technical knowledge and best practices through design sessions, mentoring, and collaborative problem-solving.
- Take final responsibility for resolving time-critical operational escalations.
Requirements
- Deep experience defining IT operations through code, using version control, peer review, and CI/CD to deploy application and infrastructure changes.
- Strong modern software engineering practices, including peer review, unit testing, source control management, CI/CD, and Agile methodologies.
- Significant Python development experience, including work on large or complex projects.
- Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
- Familiarity with Linux storage technologies, ranging from Ceph to database systems.
- Hands-on experience administering enterprise Linux servers.
- Extensive understanding of cloud computing concepts, architectures, and technologies.
- Bachelor’s degree or higher, preferably in computer science, software engineering, or a related technical discipline.
- Strong English communication skills across written and spoken channels, including email, chat, video calls, voice calls, and in-person collaboration.
- Strong troubleshooting abilities, with the curiosity and persistence to investigate issues from the Linux kernel through to the web layer.
- Ability to collaborate effectively while knowing when to seek input from teammates and subject-matter experts.
- Adaptability, willingness to learn quickly, and comfort working in fast-changing technical environments.
- Ability to thrive within globally distributed teams and collaborate across different locations and time zones.
- Strong interest in open-source technologies, particularly Ubuntu or Debian.
Benefits
- Opportunity to work on production infrastructure supporting large-scale global services.
- Exposure to private and public cloud environments, Infrastructure as Code, GitOps, CI/CD, observability, and open-source technologies.
- Dedicated development time for impactful automation and larger engineering projects.
- Collaboration with a highly experienced, globally distributed SRE and engineering community.
- Opportunities for mentoring, knowledge sharing, and cross-functional technical collaboration.
- Remote work flexibility, with the role available across time zones.
- Opportunities to meet colleagues in person 2–4 times per year at internal events, typically lasting 1–2 weeks.
- International exposure through collaboration with distributed teams and participation in global company events.
- Compensation and benefits are determined according to the role, location, experience, and applicable company policies.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
