Skip to content
← Back to job listings

Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt Ltd · Ahmedabad, GJ, India

Imported listingfull-time23 days ago

About The Role

Why Aubergine

  • Aubergine is a
  • global transformation and innovation partner
  • , shaping next-gen digital products through consulting-informed execution that integrates strategy, design, and development.
  • Since 2013, we’ve built
  • 400+ B2B and B2C products worldwide
  • , turning powerful ideas into impact-driven experiences. We are one of the
  • top global B2B companies on Clutch
  • , rated highest among more than 80,000 technology service providers.
  • With more than
  • 150 digital thinkers
  • , we are home to some of the brightest, most passionate people around the world who are committed to delivering excellence.
  • We’re not just another workplace. We’re officially
  • Great Place To Work® certified
  • , with an exceptional trust index rating, making Aubergine a community where you can thrive and grow.

Role Overview: Build AI Systems We Can Trust

We are looking for a Senior AI Evaluation & Reliability Engineer who is passionate about solving one of the most important challenges in AI:

How do we know an AI system is actually working, improving, and delivering business value?

You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.

This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.

What You Will Own

Make AI Performance & ROI Measurable

  • Build evaluation strategies that connect AI performance to business outcomes and ROI.
  • Define quality benchmarks, SLOs, risk thresholds, and success criteria.
  • Track hallucinations, reliability issues, and production risks.
  • Build executive-friendly AI quality and ROI scorecards.
  • Help clients determine where AI should be autonomous, supervised, or avoided.
  • Don't just measure model accuracy. Measure business impact.

Build Production-Grade Evaluation Pipelines

  • Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.
  • Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.
  • Build regression suites for prompts, models, tools, and workflows.
  • Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.
  • Establish statistically meaningful benchmarks and continuously monitor AI quality.
  • Evaluation should become part of engineering, not a final QA step.

Engineer LLM-as-a-Judge Systems

  • Design reference-based and reference-free LLM evaluation frameworks.
  • Create structured rubrics and scoring systems.
  • Build calibration loops using human-labelled datasets.
  • Identify and mitigate judge biases such as position, verbosity, and self-preference bias.
  • Measure judge reliability and optimize evaluation quality, latency, and cost.

Own Evaluation Economics

LLM evaluations can become expensive quickly.

You will

  • Design tiered evaluation strategies.
  • Use deterministic and heuristic graders for simple checks.
  • Reserve powerful LLM judges for complex evaluations.
  • Optimise batching, concurrency, and parallel execution.
  • Track evaluation cost and balance quality, latency, and inference spend.

Benchmark RAG & Agentic Systems

Build systematic benchmarks across

  • Retrieval quality, faithfulness, and answer relevance
  • Embedding, chunking and re-ranking strategies
  • Vector databases and retrieval architectures
  • Agent task completion and tool-calling accuracy
  • State retention, multi-step workflows and failure recovery
  • Latency, reliability and cost
  • Build synthetic datasets, golden test suites, and adversarial scenarios to continuously expand evaluation coverage.

Be a Trusted AI Consultant

You will work directly with North American enterprise clients as a technical AI advisor.

You will

  • Lead AI architecture and evaluation discussions.
  • Translate complex AI metrics into clear business recommendations.
  • Define AI quality, reliability, and risk frameworks.
  • Present evaluation telemetry and ROI scorecards to technical and business stakeholders.
  • Challenge assumptions and recommend the right AI solution, even when that means saying "don't use AI here."
  • You should be equally comfortable discussing LLM evaluation with engineers and ROI with a CTO/COO/CEO/CFO.

Drive AI Consulting & Presales

Your consulting mindset will extend beyond delivery. You will actively participate in presales and new business opportunities, helping us win strategic AI engagements with enterprise clients.

You will

  • Participate in discovery calls, solution workshops, and technical discussions with prospective clients.
  • Understand client challenges and identify where AI, agents, RAG, or automation can create meaningful business value.
  • Shape AI solution approaches, evaluation strategies, and technical proposals.
  • Contribute to statements of work, estimates, solution architectures, and project proposals.
  • Build compelling technical narratives and demonstrations that showcase our AI capabilities.
  • Present AI solutions and recommendations to technical and executive stakeholders.
  • Help identify opportunities to expand existing engagements through new AI capabilities.
  • Partner closely with Sales, Delivery, Product, and Engineering teams to convert opportunities into successful engagements.
  • Bring a consultative, outcome-focused approach rather than simply responding to client requirements.
  • You won't just help us deliver AI solutions. You'll help us win the right AI problems to solve.
  • Coach Engineers. Raise the Bar.
  • As our organisation evolves towards an AI superagency, you will help shape how we build AI.

You will

  • Mentor engineers working on AI and LLM systems.
  • Establish AI engineering and evaluation best practices.
  • Conduct technical workshops and knowledge-sharing sessions.
  • Review AI architectures and evaluation strategies.
  • Help engineers move from prompt experimentation to disciplined AI engineering.
  • Build reusable frameworks, playbooks, and internal accelerators.
  • Stay ahead of emerging AI technologies and bring valuable ideas into the organisation.
  • Your impact shouldn't be limited to the systems you build. It should be reflected in the engineers you make better.

What We Are Looking For

AI Evaluation & Reliability

  • Proven experience building production-grade AI evaluation systems.
  • Strong practical experience with LLM-as-a-Judge.
  • Experience with calibration, human ground truth, and evaluation bias.
  • Strong understanding of RAG and agent evaluation.
  • Knowledge of Faithfulness, Context Precision, Answer Relevance, Task Completion, and semantic similarity.
  • Experience with DeepEval, Ragas, TruLens, LangSmith, DSPy, Promptfoo, or equivalent.

Engineering

  • Advanced Python and/or TypeScript.
  • Strong backend, API, and asynchronous systems experience.
  • Experience with CI/CD and automated testing.
  • Experience with vector databases such as Pinecone, Qdrant, Milvus, or Chroma.
  • Understanding of LLM infrastructure, observability, and inference economics.

Consulting & Leadership

  • Experience working directly with enterprise clients.
  • Excellent communication and presentation skills.
  • Strong consulting and problem-solving mindset.
  • Ability to translate technical complexity into business outcomes.
  • Strong mentoring and coaching abilities.
  • Curiosity, ownership, and enthusiasm for emerging AI technologies.

Why This Role Matters

  • The next generation of AI companies won't simply be the ones with access to the best models.
  • They will be the ones that can measure AI, trust AI, improve AI, and turn AI into business value.
  • If you want to build that future with us, we'd love to hear from you.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing