AI Evaluation Engineer - Proofline
fermi ai · Bangalore, India
About The Role
AI Evaluation Engineer — Evals & Systems VerificationLocation: HSR, Bengaluru (On-site) Experience: 3–6 yearsAbout the RoleBuilding with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.ResponsibilitiesBuild and maintain evaluation harnesses for AI-facing features to measure and tune system quality (e.g., capture quality, retrieval quality, guidance quality)Own end-to-end and API test infrastructure (Playwright-class), supporting a continuous, daily-release cycleDesign and execute host-behavior probes using scripted user sessions across diverse AI assistants, ensuring product behavior aligns with expectationsGate production releases through thorough user acceptance testing (UAT), regression analysis, and quality reportingRequirementsBackground in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure and not just executing testsStrong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or CypressFamiliarity with LLM applications or a demonstrated interest in evaluating non-deterministic systemsHighly autonomous and able to define your own workflows and processes for evaluation and verification
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring