Skip to content
← Back to job listings

Forward Deployed Engineer (RL Environments)

Labelbox · San Francisco, United States

RemoteImported listingfull-time18 days ago

About The Role

Join our team as a Forward Deployed Engineer, where you'll design, develop, and operationalize reinforcement learning environments. This hands-on engineering role involves writing production-quality infrastructure code, integrating with open-source RL tooling, and collaborating with our data operations team. You'll build sandboxed RL environments, develop reproducible execution environments, and own environment deployment and reliability. Enjoy a flexible vacation policy, 401k program, college savings account, daily lunches, virtual wellness programs, and more.

  • Conception, développement et opérationnalisation d'environnements d'apprentissage par renforcement.
  • Création d'environnements d'exécution reproductibles et conteneurisés qui soutiennent les déploiements de tâches déterministes.
  • Collaboration avec l'équipe des opérations de données pour concevoir des curricula de tâches et des protocoles d'évaluation.
  • Comfort working with browser automation frameworks or terminal interaction tooling
  • Familiarity with RL concepts: MDPs, reward shaping, episode structure, observation/action spaces. You don’t need to have trained models, but you need to understand what an environment must provide to an RL training loop
  • Ability to read and implement from academic papers and open-source benchmark repositories without extensive hand-holding
  • 2+ years of professional software engineering experience, with strong fundamentals in Python and at least one systems-level language (Go, Rust, C++)
  • Demonstrated experience with containerization and sandboxing (Docker, Podman, Firecracker, or similar) in production or near-production contexts
  • Strong debugging instincts—you can trace failures across process boundaries, container layers, and network calls
  • Experience building or maintaining developer tooling, CLI tools, or infrastructure automation
  • Direct experience building or contributing to RL environments (Gymnasium/Gym, PettingZoo, or custom environment implementations)
  • Prior work at an AI data company, ML platform company, or AI research lab
  • Experience with agentic AI evaluation frameworks (SWE-bench, WebArena, OSWorld, TerminalBench, or similar)
  • Familiarity with GCP or AWS infrastructure (Compute Engine, ECS/EKS, Cloud Build)
  • Contributions to open-source projects in the RL, agents, or dev-tools space
  • The ideal candidate is a strong software engineer first, with genuine curiosity and working knowledge of reinforcement learning
  • You’re the kind of engineer who reads an RL benchmark paper and immediately thinks about how to make the environment more robust, not how to improve the policy gradient
  • You’ve probably built infrastructure or developer tooling at a startup or mid-stage company, and you’ve been pulled toward the ML/AI space—maybe through side projects, open-source contributions, or a prior role adjacent to an ML team
  • You move fast, but you care about reliability because you know environments that break silently poison training data
  • You thrive in ambiguity
  • You can take a loosely defined project requirement—“build an environment that tests an agent’s ability to navigate a file system and execute multi-step bash workflows”—and deliver a working, tested, documented system without needing a detailed spec

This is an external listing. JobSpring does not represent or verify the employer. Report this listing