Skip to content
← Back to job listings

Staff / Senior Machine Learning Engineer (Reinforcement Learning)

Wayve · Sunnyvale, United States

Imported listingfull-time24 days ago

About The Role

Join Wayve, a leading company in the autonomous vehicle industry, as a Senior / Staff Machine Learning Engineer. In this role, you will advance reinforcement learning methods for end-to-end driving models, lead the technical direction and delivery of a learned emergency trajectory model, and shape and execute the reinforcement learning roadmap. You will work closely with researchers and engineers across various teams and have a significant impact on driving behavior and safety.

  • Advancing reinforcement learning methods for end-to-end driving models and taking promising ideas from design through large-scale experiments.
  • Leading the technical direction and delivery of a learned emergency trajectory model for low-frequency, high-consequence maneuvers.
  • Shaping and executing the reinforcement learning roadmap for Driving Core / Core Model Safety, selecting problems and methods against clear behavioral gaps.
  • Senior-level ownership and collaboration: able to lead a substantial technical area, work across research and engineering boundaries, and bring others along through clear written and verbal communication
  • Excellent experimental judgement: able to turn an ambiguous behavioral problem into falsifiable hypotheses, useful metrics, disciplined ablations, and clear technical decisions
  • Hands-on experience with behaviour cloning, reinforcement learning, or related methods
  • Deep understanding of modern reinforcement learning fundamentals, including policy and value learning, off-policy learning, function approximation, distribution shift, and the failure modes of learned objectives
  • Proficiency in Python and PyTorch, with strong software engineering practices and hands-on experience building reliable machine learning training and evaluation systems
  • A strong track record developing and experimentally validating reinforcement learning or closely related sequential decision-making methods on complex, high-dimensional problems
  • Experience with offline reinforcement learning, imitation learning, reward modeling, preference learning, or post-training of large neural policies
  • Experience in autonomous vehicles, robotics, control, or another domain where policies interact with safety-critical physical systems, including an understanding of motion planning, vehicle dynamics, control, or collision avoidance
  • Experience training multimodal, transformer-based, or generative policy models at scale
  • Experience with closed-loop simulation, off-policy evaluation, uncertainty or calibration, and evaluation under rare or shifted conditions
  • Proficiency in C++, CUDA, distributed training, or performance optimization for production machine learning systems

This is an external listing. JobSpring does not represent or verify the employer. Report this listing