
Senior ML Engineer (Token Factory)
jobgether · Germany
About The Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML Engineer (Token Factory) based in Germany.
This role offers the opportunity to work at the forefront of large-scale AI infrastructure and machine learning <systems.You> will help build inference and fine-tuning technologies for foundation models spanning language, vision, audio, and multimodal architectures.Your work will focus on improving model quality, training efficiency, inference performance, and hardware utilization at massive <scale.You> will tackle technically challenging problems involving distributed training, low-precision computation, optimization, and reinforcement learning.Working primarily with Python and JAX, you will turn advanced research ideas into reliable, production-ready systems.The role combines deep technical ownership with opportunities to influence engineering practices and contribute to the evolution of AI <platforms.You> will collaborate with highly experienced engineers and researchers in a fast-moving, international environment where your work can have significant impact.
Accountabilities
- Develop and improve advanced fine-tuning methodologies, including LoRA-based and full-parameter approaches, for cutting-edge foundation models.
- Optimize model quality and training efficiency across large-scale machine learning workloads.
- Identify and address bottlenecks in large language model inference to improve production performance and resource efficiency.
- Build training and evaluation pipelines using JAX for techniques such as speculative decoding and advanced inference optimization.
- Experiment with different model architectures, including dense and mixture-of-experts models as well as autoregressive and parallel approaches.
- Develop and evaluate scaling laws to inform model development, performance optimization, and resource allocation.
- Investigate low-precision training and inference approaches, including FP8, NVFP4, and MXFP4, for supervised fine-tuning and reinforcement learning.
- Work with distributed training environments spanning multiple computational nodes and large GPU clusters.
- Analyze performance considerations such as sharding strategies, custom kernels, and modern hardware capabilities.
- Translate research concepts and experimental results into robust, scalable, production-quality machine learning systems.
- Apply strong software engineering practices, including CI/CD, version control, unit testing, and maintainable code design.
- Collaborate across engineering and research teams while communicating technical concepts clearly and contributing to technical direction.
Requirements
- Deep understanding of the theoretical foundations of machine learning and reinforcement learning.
- Strong expertise in modern deep learning techniques for language processing and generation.
- Demonstrated experience training large machine learning models across multiple computational nodes.
- Solid understanding of performance optimization for large neural network training, including sharding strategies, custom kernels, and hardware-specific capabilities.
- Strong software engineering skills, particularly with Python.
- Extensive experience with modern deep learning frameworks, particularly JAX.
- Proficiency in contemporary software development practices, including CI/CD, version control, unit testing, and production-quality engineering.
- Strong communication, collaboration, and technical leadership abilities.
- Experience working with language models or related NLP technologies is highly valued.
- Familiarity with concepts such as multi-head attention, RoPE, ZeRO/FSDP, Flash Attention, and quantization is advantageous.
- Experience building and delivering products in dynamic, startup-like environments is a plus.
- Strong engineering background in distributed systems or high-load web services is beneficial.
- Open-source projects demonstrating advanced engineering capabilities are valued.
- Excellent English communication skills, including strong technical writing and articulation.
Benefits
- Competitive compensation.
- Career growth and continuous learning opportunities.
- Flexible working environment with a high degree of ownership.
- Opportunity to work on impactful, large-scale AI and machine learning projects.
- Collaborative culture with experienced engineers and researchers.
- International environment with diverse and highly skilled teams.
- Opportunity to contribute to advanced foundation model training, fine-tuning, inference optimization, and AI infrastructure.
- Exposure to cutting-edge GPU computing, distributed systems, and modern machine learning technologies.
- Inclusive workplace committed to equal employment opportunities.
- Support and reasonable accommodations throughout the hiring process when required.
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring