Skip to content
← Back to job listings

ML Platform Engineer

jobgether · Spain

RemoteExternal listingfull-time18 days ago

About The Role

  • **This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a ML Platform Engineer based in Spain.**
  • This role offers the opportunity to build and evolve the infrastructure powering advanced AI products used at enterprise scale.
  • You will design reliable, scalable systems that enable machine learning teams to train, deploy, and operate complex models efficiently.
  • Working at the intersection of software engineering, cloud infrastructure, and machine learning, you will help shape the future of AI platform capabilities.
  • The position focuses on automation, reliability, performance optimization, and creating tools that improve how teams build and operate ML systems.
  • You will collaborate closely with researchers and product engineers to transform technical challenges into robust platform solutions.
  • This is an ideal opportunity for a systems-focused engineer who enjoys solving complex infrastructure problems and driving meaningful improvements in AI development workflows.

### Accountabilities

  • Design, develop, and improve platform systems supporting machine learning model training, evaluation, deployment, and production serving.
  • Build scalable infrastructure and internal tooling that improve the reliability, efficiency, and cost-effectiveness of machine learning workloads.
  • Develop automation workflows, internal tools, and agent-oriented systems that reduce operational complexity for researchers and engineers.
  • Architect and maintain systems that enable efficient model deployment, monitoring, and operation across research and product environments.
  • Improve workload scheduling, monitoring, debugging, and resource management for GPU-based and cloud infrastructure environments.
  • Drive improvements across observability, automation, reliability, developer experience, and platform usability.
  • Create abstractions and developer tools that enable engineering teams to work more effectively with complex ML systems.
  • Collaborate with research and product teams to identify technical challenges and turn them into scalable platform capabilities.
  • Contribute to architectural decisions, technical strategy, and long-term platform evolution.
  • Take ownership of open-ended engineering challenges while making pragmatic decisions that balance scalability, simplicity, and reliability.

## **Requirements:**

  • Strong professional experience building or operating production systems with a focus on reliability, scalability, performance, and maintainability.
  • Strong systems mindset with the ability to reason about bottlenecks, failure scenarios, interfaces, resource utilization, and long-term operational needs.
  • Hands-on experience with cloud infrastructure, Linux environments, and infrastructure automation.
  • Experience operating distributed systems and workloads in production, including Kubernetes-based environments.
  • Strong programming skills in Python or similar backend-oriented programming languages.
  • Experience building internal platforms, developer tooling, infrastructure abstractions, or systems used by engineering teams.
  • Understanding of machine learning infrastructure, model serving systems, or data-intensive workloads.
  • Experience working with GPU-based systems, performance-sensitive environments, or large-scale computing resources.
  • Familiarity with observability, monitoring, and debugging practices for distributed systems.
  • Knowledge of infrastructure and development tools such as Terraform, Datadog, GitHub Actions, or similar technologies.
  • Ability to work effectively in ambiguous environments, take ownership, and solve complex technical problems independently.
  • Pragmatic approach to engineering, focusing on delivering valuable solutions without unnecessary complexity.

**Preferred Skills & Experience:**

  • Experience building agentic systems or internal tools powered by large language models.
  • Familiarity with workflow orchestration platforms such as Temporal.
  • Experience working between research and production engineering teams.
  • Background in performance optimization, scheduling, or resource allocation challenges.
  • Experience developing lightweight tools or products for engineers and technical users.

## **Benefits:**

  • Fully remote work environment with flexibility across Europe.
  • Opportunity to work on advanced AI infrastructure supporting large-scale enterprise applications.
  • High ownership role with significant influence over platform architecture and technical direction.
  • Collaborative environment with close interaction between engineering, research, and product teams.
  • Opportunity to solve complex challenges involving machine learning systems, automation, and distributed infrastructure.
  • Professional growth opportunities within a fast-moving AI-focused organization.
  • Ability to contribute to tools and systems that improve productivity for technical teams worldwide.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing