Skip to content
← Back to job listings

Analyst-Data Science

Egug · Gurugram, HR, India

IT - Network / Systems / DB AdminExternal listingfull-timeabout 4 hours ago

About The Role

The role

We are looking for an early-career engineer who enjoys understanding how modern AI systems actually work.

You will work across model experimentation, inference, evaluation, deployment, and ML systems performance . This is a hands-on engineering role: you will build prototypes, run experiments, investigate failures, profile systems, and turn promising ideas into working implementations.

What you will do

  • Experiment with LLMs and multimodal models.
  • Build prototypes and production-quality components in Python and PyTorch .
  • Run and optimise model inference using tools such as vLLM .
  • Measure and improve latency, throughput, batching, caching, GPU utilisation, and memory usage.
  • Build evaluation frameworks to understand model quality and behaviour.
  • Work with embeddings, retrieval, reranking, structured generation, and tool-using/agentic systems.
  • Deploy model-backed services and troubleshoot them in realistic workloads.
  • Read relevant research papers and reproduce or test promising ideas.
  • Design controlled experiments and analyse failures across data, model, software, and infrastructure.
  • Document findings clearly, including what worked, what failed, and why.

Core requirements

  • Strong Python skills.
  • Working knowledge of PyTorch and modern neural-network architectures.
  • Understanding of transformers, tokenisation, embeddings, attention, sampling, and decoding.
  • Ability to write maintainable software beyond notebooks.
  • Comfortable working with Linux, Git, Docker, APIs, and basic cloud infrastructure .
  • Strong analytical and debugging skills.

Useful ML systems knowledge

You should understand, or be motivated to learn

  • model serving and inference;
  • vLLM or similar runtimes;
  • continuous batching and KV caching;
  • quantization;
  • throughput vs latency trade-offs;
  • GPU memory constraints;
  • mixed precision and device placement;
  • profiling and out-of-memory debugging;
  • structured/constrained generation.
  • Experience with CUDA, Triton, distributed systems, Kubernetes, NCCL, or low-level optimisation is useful but not required .
  • What we look for
  • We care more about demonstrated technical depth than years of experience.

Good evidence includes

  • a substantial ML or systems project;
  • research or thesis work;
  • reproducing or implementing a research paper;
  • open-source contributions;
  • building or profiling an inference/training system;
  • technically serious side projects;
  • benchmarks or experiments where you measured and improved performance.
  • You should be able to explain what you built, why you built it that way, what you measured, what failed, and what you learned .
  • Academic background
  • A strong foundation in a quantitative discipline such as Computer Science, Mathematics, Statistics, Engineering, Physics, Operations Research, or a related field is preferred.
  • Research experience is useful but not mandatory.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing