Machine Learning Engineer - MLOps & AI Platform
beigene · 上海市, 上海市, 中国
About The Role
We are seeking a hands-on Machine Learning Engineer to join the Biologics AI CoE. This is primarily an ML engineering role with meaningful ownership of MLOps and GPU-compute enablement. The successful candidate will turn research code into reliable, reproducible, and scalable training and inference workflows for biologics AI applications. Working closely with ML scientists, computational biologists, and infrastructure partners, this individual will improve engineering quality, developer experience, and compute efficiency across the model lifecycle. The role includes practical support for SLURM-based GPU workloads, experiment tracking, profiling, source control, testing, and CI/CD. Essential Responsibilities: • Engineer and maintain reusable training, evaluation, and inference pipelines for deep-learning models used in biologics discovery. • Translate research prototypes into well-tested, modular, documented, and maintainable production-quality code. • Enable efficient single-node and distributed GPU training using PyTorch DDP/FSDP, DeepSpeed, or comparable frameworks. • Profile model training and inference to identify bottlenecks in GPU utilization, memory, data loading, kernels, and interconnect communication; implement and validate improvements. • Support day-to-day use of SLURM-based GPU clusters, including job templates, queues/partitions, resource requests, failure diagnosis, and user guidance; partner with IT/HPC teams on underlying infrastructure issues. • Build dashboards, reports, or alerts for GPU utilization, job efficiency, capacity trends, and failed or stalled workloads. • Establish reproducible experiment-management practices using Weights & Biases (W&B) or equivalent tools, including run metadata, artifacts, model versions, and comparisons. • Own and promote GitHub engineering practices: branching, pull requests, peer review, code ownership, issue tracking, release practices, and repository hygiene. • Develop and maintain automated tests and CI/CD workflows for ML code, containers, and deployment artifacts. • Create containerized, reproducible runtime environments and manage dependencies for on-premises and/or cloud execution. • Implement model packaging and serving workflows for batch and online inference, with appropriate monitoring for latency, throughput, reliability, and data/model drift where applicable. • Document platform standards, troubleshooting playbooks, and self-service workflows; coach scientists on efficient and responsible use of shared compute resources. • Evaluate new ML systems and MLOps technologies pragmatically, balancing scientific velocity, reliability, security, maintainability, and cost. Qualifications Required: • Bachelor's or Master's degree in a relevant quantitative or engineering discipline, with typically 4+ years of relevant industry experience; equivalent practical experience will be considered. • Strong Python engineering skills and hands-on experience with PyTorch or another modern deep-learning framework. • Demonstrated ability to build, debug, and optimize ML training and inference pipelines on NVIDIA GPUs. • Working knowledge of Linux, Bash, networking/storage fundamentals, environment and dependency management, and systematic troubleshooting. • Hands-on experience with SLURM or a comparable workload scheduler in an HPC or multi-GPU environment. • Experience with GPU and application profiling tools such as PyTorch Profiler, NVIDIA Nsight Systems/Compute, nvidia-smi/DCGM, or equivalent tooling. • Practical experience with Git and GitHub, pull-request review, automated testing, and CI/CD (for example, GitHub Actions). • Experience with experiment tracking and artifact/model management using W&B, MLflow, or a comparable platform. • Experience with Docker or another container technology; familiarity with Kubernetes is beneficial but not mandatory. • Understanding of distributed training, mixed precision, checkpointing, reproducibility, and performance/cost trade-offs. • Strong ownership, communication, and collaboration skills, including the ability to work effectively with both research and infrastructure teams. • Professional working proficiency in English; Chinese communication capability is strongly preferred for the Shanghai-based team. Preferred Qualifications: • Experience with protein language models, large language models, diffusion/flow models, graph neural networks, or other foundation-model workloads. • Experience in computational biology, bioinformatics, biologics, pharmaceutical R&D, or another scientific-computing environment. • Familiarity with Azure, AWS, or GCP; infrastructure-as-code or configuration-management experience is a plus. • Experience with model-serving technologies such as NVIDIA Triton, vLLM, TorchServe, or KServe. • Experience building internal ML platforms, developer tooling, observability, or chargeback/capacity-planning solutions. 百济神州全球胜任力 当我们通过以下十二项全球胜任力,展现出 "患者为先"、"无界协作"、"锐意创新 "和 "追求卓越 "的价值观时,我们就能帮助全世界更多患者获得更多负担得起的药品。 ●团队协作 ●提供并征求坦诚及可行的反馈 ●自我认知 ●兼容并蓄 ●积极主动 ●开拓精神 ●持续学习 ●拥抱变化 ●结果导向 ●分析性思维/数据分析 ●卓越财务 ●清晰沟通 BeOne Global Competencies When we exhibit our values of Patients First, Collaborative Spirit, Bold Ingenuity and Driving Excellence, through our twelve global competencies below, we help get more affordable medicines to more patients around the world. ●Fosters Teamwork ●Provides and Solicits Honest and Actionable Feedback ●Self-Awareness ●Acts Inclusively ●Demonstrates Initiative ●Entrepreneurial Mindset ●Continuous Learning ●Embraces Change ●Results-Oriented ●Analytical Thinking/Data Analysis ●Financial Excellence ●Communicates with Clarity 求职者隐私申明: 百济神州致力于尊重和保护您的个人信息权利,并承诺依据合法、正当、必要和诚信的原则处理您的个人信息(包括个人敏感信息 )。 由于百济神州在全球范围内开展业务,我们可能需要基于人力资源管理等合理业务目的而将您的个人信息发送和/或存储在位于您所在国家以外其他国家(例如:美国)的服务器和数据库中,详情参见百济神州《求职者隐私政策》(百济神州官网 - 隐私政策 - 求职者隐私政策)。 如您主动向我们提供您的简历信息或其他个人信息,则视为您已经充分理解并确认接受百济神州《求职者隐私政策》内容。如您对此有任何疑问的,请勿提交简历信息或其他个人信息。 BeOne is committed to respect and protect your personal information rights, and will process your personal information, including your sensitive personal information, based on the principles of legality, legitimacy, necessity, and integrity. Due to the reasonable business need for human resource management as a result of BeOne’s global operation, your personal information may be transferred and/ or stored in a server/database located in a third country (e.g., the United States) other than your own country. For further details, please refer to BeOne Job Applicant Privacy Policy (BeOne official website - Privacy Policy - Job Applicant Privacy Policy). If you voluntarily provide your resume or other personal information to us, it is deemed as you have thoroughly acknowledged and accepted BeOne Job Applicant Privacy Policy. If you have any concern, please DO NOT submit your resume or any other personal information.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring