Skip to content
← Back to job listings

【JAPAN AI】AI Evaluation Scientist / English

株式会社ジーニー · 東京都新宿区西新宿住友不動産新宿オークタワー 5/6階

External listingfull-timeRecently

About The Role

About JAPAN AI JAPAN AI, Inc. was established in April 2023 as a group company of Geniee, Inc. (TSE Growth Market) with the mission of dramatically expanding human potential through AI technology. We drive cutting-edge AI R&D both domestically and internationally. Related URLs Our Website Company Introduction Materials Tech Blog Careers Why We're Hiring JAPAN AI is rapidly expanding its enterprise AI agent suite, including JAPAN AI AGENT / CHAT / SPEECH. As the core of our products shifts to LLMs and multi-agent systems, we are establishing a new specialized organization to scientifically evaluate the quality, safety, and reliability of AI outputs. Mission "Make AI Output Quality a Science — Prove Agent Reliability through Research and Development of Evaluation Methods." You will quantitatively evaluate and improve the output quality of LLMs and AI agents using methods from machine learning, statistics, and psychometrics. This position is not for "people who test" — it is for "scientists who define and measure what makes a good AI." Role & Expectations As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent quality-evaluation infrastructure. Research and develop evaluation metrics — scientifically define "what constitutes quality" through LLM-as-Judge calibration, reward modeling, and benchmark design Design and build automated evaluation pipelines — integrate research outcomes into production CI/CD to deliver scalable quality gates Red teaming and safety verification — automate adversarial testing and build policy compliance verification frameworks Drive quality improvement through statistical experimental design — quantitatively verify the effectiveness of prompt strategies and model changes through A/B tests and significance testing Feed evaluation signals back to research and development teams — build a compound-interest loop for model improvement Ensure the quality of products used in production by ~200 companies through a "science of quality" approach Why You'll Love This Role Evaluation Science in practice : Practice "AI Evaluation Science" — the discipline that Apple, Anthropic, Scale AI, and others are investing in — within the context of Japanese enterprise AI. This is a globally rare position where evaluation methodology itself is the research subject. A new application of ML/DS skills : Apply your machine learning and statistics expertise not to "building models" but to "evaluating models." Intellectual challenges span both research and implementation — reward modeling, LLM-as-Judge calibration theory, and benchmark design. Quality determines product trust : In a production environment used by ~200 companies, the evaluation infrastructure you build becomes the last line of defense for release quality. You will feel the direct business impact of quality assurance. Greenfield position : Design and build the entirely new specialized domain of AI agent evaluation science from scratch. You will have significant autonomy — from evaluation metric R&D to production deployment of automated evaluation pipelines. Frontline of AI safety : Engage in Responsible AI practices including automated red teaming, adversarial testing, and policy compliance verification. You will play a key role in scientifically guaranteeing safety in a world where AI agents autonomously execute business operations as "the brain of the enterprise." Rapid-growth environment : In a startup that has grown to 200+ people and 9 products in just 3 years, you will have significant autonomy in technical decision-making. You will work closely with Research Engineers and Agent Harness Engineers, influencing quality across the entire product suite. Job Description As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent Evaluation Infrastructure. Evaluation Metric Research & Development Research and implement LLM-as-Judge calibration methods (rubric design, bias detection, proper scoring rules) Design, build, and validate evaluation benchmarks (construct validity, contamination detection) Research the application of reward modeling / preference learning to evaluation Select and design evaluation metrics (win rate, task success, factuality, harm detection) Design, build, and maintain evaluation sets (synthetic data + real logs) Automated Evaluation Pipeline Design & Development Design and implement scalable automated evaluation pipelines Integrate evaluation pipelines into CI/CD and build quality gates Design agent evaluation harnesses (multi-turn, tool use, long-context support) Ensure reproducibility and reliability of evaluation pipelines Safety & Quality Verification Research and implement automated red teaming (automated adversarial testing) Build safety and policy compliance verification frameworks Research and implement hallucination detection and calibration methods Design and execute prompt / tool regression tests Statistical Analysis & Experimental Design Design and analyze statistical experiments (A/B tests, significance testing) Visualize quality trends and automate regression detection Create quality reports and improvement proposals Feed evaluation signals back to research and development teams Key Results (KR/Metrics) Evaluation coverage rate (test case coverage) Regression detection rate (pre-release quality degradation detection ≥ 95%) Evaluation pipeline execution time (completed within CI/CD) LLM-as-Judge and human evaluation agreement rate False positive / false negative rate Safety incident rate (post-release) Team Structure Approximately 120 members are part of the development organization. The AI Evaluation Scientist operates as a dedicated quality assurance function, collaborating closely with: Agentic Product Engineer — Agent feature development Research Engineer — Research and development, model improvement Agent Harness Engineer / Software Engineer (AI Platform) — AI execution infrastructure development Product Manager — Product design and quality requirements definition You May Be a Good Fit If You Education & Experience Master's degree or higher (or equivalent practical experience) in Computer Science, Machine Learning, Statistics, Mathematics, Physics, Psychometrics, or related fields 3+ years of practical experience as an ML Engineer, Data Scientist, Research Engineer, or in ML/AI evaluation-related roles Technical Skills Deep knowledge of LLM / generative AI evaluation methods (benchmark design, LLM-as-Judge, quantitative output quality measurement, hallucination detection, etc.) Practical knowledge of statistics and experimental design (hypothesis testing, A/B testing, confidence intervals, effect sizes, etc.) Experience building ML / evaluation pipelines in Python Practical experience with machine learning frameworks (PyTorch, JAX, TensorFlow, etc.) Experience designing and implementing evaluation metrics (task-specific metric design beyond precision/recall) Language requirement (at least one of the following): Japanese: Fluent — able to discuss product development without friction English: Business level This position is a research and development role responsible for AI output Evaluation Science. Research or implementation experience in ML model evaluation / LLM evaluation is required. Strong Candidates May Also Have Publication experience at top ML/NLP conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, etc.) Research or implementation experience with reward modeling / preference learning (RLHF, DPO, etc.) Experience with LLM-as-Judge calibration and rubric design Knowledge or experience in AI safety, Responsible AI, and red teaming Experience with benchmark design and validity verification (IRT, construct validity) Experience evaluating multi-agent workflows, tool use, and long-context scenarios Large-scale data processing experience (Spark / BigQuery, etc.) Experience integrating ML / evaluation pipelines into CI/CD Ability to read, comprehend, and reproduce research papers Technical communication ability in English Tech Stack Languages : Python (evaluation pipelines & analysis) , TypeScript / React / Next.js (frontend) / NX Evaluation/QA : pytest, LangSmith, Weights & Biases, custom eval frameworks Data : BigQuery, Spark, Pandas Infrastructure : GCP (containers / K8s) , Docker, Terraform CI/CD : GitHub Actions Tools : Slack, Confluence, Linear, Google Workspace, GitHub, Notion AI Dev Support: Claude Code MAX Plan, Cursor, ChatGPT, Devin Work environment : Mac (Apple Silicon) , dual monitors available

This is an external listing. JobSpring does not represent or verify the employer. Report this listing