Senior AI Evaluation & Reliability Engineer
Aubergine Solutions Pvt Ltd · Ahmedabad, GJ, India
About The Role
Why Aubergine
- Aubergine is a
- global transformation and innovation partner
- , shaping next-gen digital products through consulting-informed execution that integrates strategy, design, and development.
- Since 2013, we’ve built
- 400+ B2B and B2C products worldwide
- , turning powerful ideas into impact-driven experiences. We are one of the
- top global B2B companies on Clutch
- , rated highest among more than 80,000 technology service providers.
- With more than
- 150 digital thinkers
- , we are home to some of the brightest, most passionate people around the world who are committed to delivering excellence.
- We’re not just another workplace. We’re officially
- Great Place To Work® certified
- , with an exceptional trust index rating, making Aubergine a community where you can thrive and grow.
Role Overview: Build AI Systems We Can Trust
We are looking for a Senior AI Evaluation & Reliability Engineer who is passionate about solving one of the most important challenges in AI:
How do we know an AI system is actually working, improving, and delivering business value?
You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.
This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.
What You Will Own
Make AI Performance & ROI Measurable
- Build evaluation strategies that connect AI performance to business outcomes and ROI.
- Define quality benchmarks, SLOs, risk thresholds, and success criteria.
- Track hallucinations, reliability issues, and production risks.
- Build executive-friendly AI quality and ROI scorecards.
- Help clients determine where AI should be autonomous, supervised, or avoided.
- Don't just measure model accuracy. Measure business impact.
Build Production-Grade Evaluation Pipelines
- Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.
- Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.
- Build regression suites for prompts, models, tools, and workflows.
- Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.
- Establish statistically meaningful benchmarks and continuously monitor AI quality.
- Evaluation should become part of engineering, not a final QA step.
Engineer LLM-as-a-Judge Systems
- Design reference-based and reference-free LLM evaluation frameworks.
- Create structured rubrics and scoring systems.
- Build calibration loops using human-labelled datasets.
- Identify and mitigate judge biases such as position, verbosity, and self-preference bias.
- Measure judge reliability and optimize evaluation quality, latency, and cost.
Own Evaluation Economics
LLM evaluations can become expensive quickly.
You will
- Design tiered evaluation strategies.
- Use deterministic and heuristic graders for simple checks.
- Reserve powerful LLM judges for complex evaluations.
- Optimise batching, concurrency, and parallel execution.
- Track evaluation cost and balance quality, latency, and inference spend.
Benchmark RAG & Agentic Systems
Build systematic benchmarks across
- Retrieval quality, faithfulness, and answer relevance
- Embedding, chunking and re-ranking strategies
- Vector databases and retrieval architectures
- Agent task completion and tool-calling accuracy
- State retention, multi-step workflows and failure recovery
- Latency, reliability and cost
- Build synthetic datasets, golden test suites, and adversarial scenarios to continuously expand evaluation coverage.
Be a Trusted AI Consultant
You will work directly with North American enterprise clients as a technical AI advisor.
You will
- Lead AI architecture and evaluation discussions.
- Translate complex AI metrics into clear business recommendations.
- Define AI quality, reliability, and risk frameworks.
- Present evaluation telemetry and ROI scorecards to technical and business stakeholders.
- Challenge assumptions and recommend the right AI solution, even when that means saying "don't use AI here."
- You should be equally comfortable discussing LLM evaluation with engineers and ROI with a CTO/COO/CEO/CFO.
Drive AI Consulting & Presales
Your consulting mindset will extend beyond delivery. You will actively participate in presales and new business opportunities, helping us win strategic AI engagements with enterprise clients.
You will
- Participate in discovery calls, solution workshops, and technical discussions with prospective clients.
- Understand client challenges and identify where AI, agents, RAG, or automation can create meaningful business value.
- Shape AI solution approaches, evaluation strategies, and technical proposals.
- Contribute to statements of work, estimates, solution architectures, and project proposals.
- Build compelling technical narratives and demonstrations that showcase our AI capabilities.
- Present AI solutions and recommendations to technical and executive stakeholders.
- Help identify opportunities to expand existing engagements through new AI capabilities.
- Partner closely with Sales, Delivery, Product, and Engineering teams to convert opportunities into successful engagements.
- Bring a consultative, outcome-focused approach rather than simply responding to client requirements.
- You won't just help us deliver AI solutions. You'll help us win the right AI problems to solve.
- Coach Engineers. Raise the Bar.
- As our organisation evolves towards an AI superagency, you will help shape how we build AI.
You will
- Mentor engineers working on AI and LLM systems.
- Establish AI engineering and evaluation best practices.
- Conduct technical workshops and knowledge-sharing sessions.
- Review AI architectures and evaluation strategies.
- Help engineers move from prompt experimentation to disciplined AI engineering.
- Build reusable frameworks, playbooks, and internal accelerators.
- Stay ahead of emerging AI technologies and bring valuable ideas into the organisation.
- Your impact shouldn't be limited to the systems you build. It should be reflected in the engineers you make better.
What We Are Looking For
AI Evaluation & Reliability
- Proven experience building production-grade AI evaluation systems.
- Strong practical experience with LLM-as-a-Judge.
- Experience with calibration, human ground truth, and evaluation bias.
- Strong understanding of RAG and agent evaluation.
- Knowledge of Faithfulness, Context Precision, Answer Relevance, Task Completion, and semantic similarity.
- Experience with DeepEval, Ragas, TruLens, LangSmith, DSPy, Promptfoo, or equivalent.
Engineering
- Advanced Python and/or TypeScript.
- Strong backend, API, and asynchronous systems experience.
- Experience with CI/CD and automated testing.
- Experience with vector databases such as Pinecone, Qdrant, Milvus, or Chroma.
- Understanding of LLM infrastructure, observability, and inference economics.
Consulting & Leadership
- Experience working directly with enterprise clients.
- Excellent communication and presentation skills.
- Strong consulting and problem-solving mindset.
- Ability to translate technical complexity into business outcomes.
- Strong mentoring and coaching abilities.
- Curiosity, ownership, and enthusiasm for emerging AI technologies.
Why This Role Matters
- The next generation of AI companies won't simply be the ones with access to the best models.
- They will be the ones that can measure AI, trust AI, improve AI, and turn AI into business value.
- If you want to build that future with us, we'd love to hear from you.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
