Skip to content
← Back to job listings

Senior Inference Engineer

jobgether · Canada

RemoteExternal listingfull-timeabout 20 hours ago

About The Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Inference Engineer based in Canada.

This is an opportunity to become the first dedicated engineer responsible for building and owning an inference platform from the ground <up.You>’ll work closely with the CTO to turn large language models into reliable, production-grade query-to-response systems running at scale.The role combines hands-on infrastructure engineering with performance optimization across GPU-based <environments.You>’ll work with modern serving technologies such as vLLM, SGLang, and TensorRT-LLM while solving real-world challenges around latency, cost, and <throughput.As> the platform matures, you’ll take increasing ownership of its technical direction and collaborate with Product on the future inference <roadmap.You>’ll join a remote-first, open-source-oriented startup where engineering ownership, speed, and measurable customer impact are highly valued.This role is ideal for a senior engineer who wants significant autonomy and the opportunity to define how production inference infrastructure is built and scaled.

Accountabilities

  • Build and deploy production-grade LLM inference systems across one or multiple GPU machines, owning the complete pipeline from customer query to served response.
  • Design, implement, and operate model-serving infrastructure using technologies such as vLLM, SGLang, and TensorRT-LLM.
  • Optimize inference workloads for scale, balancing latency, throughput, reliability, and infrastructure costs.
  • Apply techniques such as quantization, batching, caching, and intelligent request routing to improve inference performance.
  • Develop robust production infrastructure using Python or Golang, with an emphasis on maintainable, scalable engineering rather than configuration-only work.
  • Establish the initial inference platform in close collaboration with the CTO and take ownership of its evolution as the organization scales.
  • Define and execute the technical roadmap for inference infrastructure, identifying opportunities to improve performance, reliability, and developer or customer experience.
  • Partner with Product to translate evolving customer and market requirements into practical inference-platform capabilities.
  • Evaluate emerging inference technologies and approaches and determine where they can create meaningful improvements.
  • Collaborate with engineering stakeholders to establish reliable operational practices for production GPU infrastructure.

Requirements

  • Significant professional experience building and operating production software or infrastructure systems, with strong hands-on engineering capabilities.
  • Demonstrated experience deploying and serving large language models in production, ideally using vLLM, SGLang, TensorRT-LLM, or comparable inference frameworks.
  • Practical expertise optimizing inference workloads through quantization, batching, caching, routing, or similar techniques.
  • Strong programming skills in Python or Golang, with a track record of writing and maintaining production-quality code.
  • Strong understanding of production inference architectures, including the journey from user request through model execution to a reliable served response.
  • Excellent problem-solving skills and the ability to independently investigate complex performance, scalability, and reliability challenges.
  • Strong communication skills, with the ability to explain sophisticated technical concepts clearly to engineers, product stakeholders, and other audiences.
  • Comfortable taking significant ownership, working with ambiguity, and making pragmatic technical decisions in a fast-moving environment.
  • Familiarity with Docker and Kubernetes is a plus.
  • Hands-on experience with generative AI technologies such as PyTorch or Transformers is advantageous.
  • Knowledge of the GPU software stack, including CUDA, NCCL, drivers, and related libraries, is beneficial.
  • Understanding of model architectures and fine-tuning techniques is a plus.
  • Experience with NVIDIA Dynamo is an additional advantage.

Benefits

  • Competitive compensation package including equity.
  • Health, dental, vision, and life insurance, with coverage for eligible dependents where available.
  • Benefits adapted to the country of employment.
  • Flexible working schedule focused on outcomes rather than fixed working hours.
  • High degree of workplace flexibility, supporting remote work and changing personal circumstances.
  • Remote-first environment with a globally distributed team.
  • Significant ownership over the architecture, implementation, and long-term roadmap of the inference platform.
  • Direct collaboration with senior technical leadership and Product teams.
  • Opportunity to work on production-grade GPU and AI infrastructure at scale.
  • Exposure to modern LLM serving, inference optimization, Kubernetes, cloud infrastructure, and open-source technologies.
  • A culture built around ownership, action, continuous improvement, open-source collaboration, and technically ambitious engineering.
  • Opportunity to help define emerging standards for AI infrastructure and production inference.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing