Skip to content
← Back to job listings

AI Inference Platform Engineer

DRW · Chicago, United States

Software DevelopmentImported listingfull-time1 day ago

About The Role

<p><strong>DRW</strong> is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot to capture opportunities, so we operate using our own capital and trading at our own risk.</p>
<p>Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets. We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets.</p>
<p>We operate with respect, curiosity and open minds. The people who thrive here share our belief that it’s not just what we do that matters–it's how we do it. <strong>DRW</strong> is a place of high expectations, integrity, innovation and a willingness to challenge consensus.</p>
<p><strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">About the Role</span></span></strong><span data-ccp-props="{"134233117":false,"134233118":false,"134245418":true,"134245529":true,"335559738":299,"335559739":299}"> </span></p>
<p><span data-contrast="auto">We're looking for an AI Inference </span><span data-contrast="auto">Platform Engineer to build, </span><span data-contrast="auto">operate</span><span data-contrast="auto">,</span><span data-contrast="auto"> and optimize the systems that serve large language</span><span data-contrast="auto">, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use</span><span data-contrast="auto">.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":240,"335559739":240}"> </span></p>
<p><span data-contrast="auto">You'll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy.  You'll own </span><span data-contrast="auto">the serving platform end-to-end: onboarding</span><span data-contrast="auto"> newly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":240,"335559739":240}"> </span></p>
<p><strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">What You'll Do</span></span></strong><span data-ccp-props="{"134233117":false,"134233118":false,"134245418":true,"134245529":true,"335559738":299,"335559739":299}"> </span></p>
<ul>
<li><span data-contrast="auto">Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
</ul>
<p><strong><span data-contrast="none"><span data-ccp-parastyle="heading 2">What We're Looking For</span></span></strong><span data-ccp-props="{"134233117":false,"134233118":false,"134245418":true,"134245529":true,"335559738":299,"335559739":299}"> </span></p>
<p><strong><span data-contrast="none"><span data-ccp-parastyle="heading 3">The Tech</span></span></strong><span data-ccp-props="{"134233117":false,"134233118":false,"134245418":true,"134245529":true,"335559738":281,"335559739":281}"> </span></p>
<ul>
<li><span data-contrast="auto">Hands-on experience </span><span data-contrast="auto">serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, </span><span data-contrast="auto">NVLink</span><span data-contrast="auto">/</span><span data-contrast="auto">NVSwitch</span><span data-contrast="auto">, and the compute and memory bottlenecks that shape serving decisions</span><span data-contrast="auto">.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
</ul>
<p><strong><span data-contrast="none"><span data-ccp-parastyle="heading 3">The Intangibles</span></span></strong><span data-ccp-props="{"134233117":false,"134233118":false,"134245418":true,"134245529":true,"335559738":281,"335559739":281}"> </span></p>
<ul>
<li><span data-contrast="auto">You take a measurement-driven approach to performance optimization.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">You take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">You can move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">You understand the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision-making processes.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">You can evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
<li><span data-contrast="auto">You communicate clearly and can explain complex performance tradeoffs across engineering teams.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":0,"335559739":0}"> </span></li>
</ul>
<p> </p>
<p>The annual base salary range for this position is $200,000 to $250,000 depending on the candidate’s experience, qualifications, and relevant skill set. The position is also eligible for an annual discretionary bonus. In addition, DRW offers a comprehensive suite of employee benefits including group medical, pharmacy, dental and vision insurance, 401k (with discretionary employer match), short and long-term disability, life and AD&D insurance, health savings accounts, and flexible spending accounts.</p>
<p><strong>For more information about DRW's processing activities and our use of job applicants'

This is an external listing. JobSpring does not represent or verify the employer. Report this listing