Skip to content
← Back to job listings

AI infrastructure Engineer

Together AI · Amsterdam, Netherlands

Quick applyfull-time22 days ago

About The Role

Join Together as a Site Reliability Engineer (SRE) and be responsible for keeping all user-facing services and production systems running smoothly. You will apply sound engineering principles, operational discipline, and mature automation to our operating environments and codebase. Your expertise in systems, availability, reliability, and scalability will be crucial in ensuring the highest quality service for our customers.

  • As a Site Reliability Engineer (SRE) at Together, you will be responsible for ensuring the smooth operation of user-facing services and production systems, applying sound engineering principles and operational discipline.
  • You will specialize in systems, implementing best practices for availability, reliability, and scalability, while also being involved in debugging production issues and identifying improvements for the product architecture.
  • You will build and run infrastructure using tools like Ansible, Terraform, and Kubernetes, and design and implement operational processes such as deployments and upgrades.
  • Ability to thrive in a collaborative environment involving different stakeholders and subject matter experts
  • Proficiency in programming/scripting languages
  • Bachelor's degree in Computer Science or a related field or equivalent work experience
  • Expert knowledge of Ansible (roles, playbooks), Terraform, and Kubernetes
  • Direct experience in monitoring and observability practices
  • Advanced knowledge of cloud services
  • 7+ years of professional SRE or related experience

This listing was posted by a verified recruiter at Together AI. Report this listing