Skip to content
← Back to job listings

Research Engineer (Reinforcement Learning)

jobgether · Switzerland

RemoteExternal listingfull-time5 days ago

About The Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Switzerland.

Join a small, senior engineering team building the next generation of voice- and text-driven AI <agents.You>’ll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions.Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production <deployment.You>’ll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior.The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model <performance.You>’ll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued.Your contributions will directly shape AI systems operating at significant production scale.

Accountabilities

  • Build training environments, verifiers, and supporting infrastructure for post-training models.
  • Own the synthetic data pipeline from data generation through quality assurance and validation.
  • Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements.
  • Design and maintain evaluations that models must pass before production releases.
  • Select and adapt suitable open-weight foundation models for specific agent and product requirements.
  • Develop trained behaviors that perform consistently across both voice and text-based agents.
  • Deploy trained models to production and continuously improve them based on real-world usage and feedback.
  • Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations.

Requirements

  • Strong Python engineering skills and the ability to build reliable, production-quality systems.
  • Demonstrated experience taking a machine learning model from raw data through experimentation and into production.
  • A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage.
  • The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms.
  • Practical experience working with GPUs and a realistic understanding of their capabilities and limitations.
  • Strong judgment around when model training is the right solution—and when a simpler approach is preferable.
  • Ability to collaborate effectively within a remote, distributed, and highly autonomous team.
  • Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable.
  • Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus.
  • Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous.
  • Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable.
  • Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus.

Benefits

  • Opportunity to make a significant impact on a fast-growing developer platform and help shape its future.
  • Collaboration with a small, highly experienced team that values technical excellence, creativity, and ownership.
  • Competitive salary and equity package.
  • Health, dental, and vision benefits.
  • Flexible vacation policy.
  • Remote-friendly working environment with flexibility and autonomy.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing