Senior AI Engineer — GenAI & Autonomous Agents
Jus Mundi · Paris, Ile-de-France, France
About The Role
Join Jus Mundi, a company dedicated to revolutionizing international law and arbitration through AI. As a Senior AI Engineer, you will be responsible for improving search relevance and retrieval, designing and building agentic systems, and conducting rigorous evaluations and benchmarks. You will also fine-tune models, deploy AI features to production, and provide technical leadership across squads. The ideal candidate will have a proven track record in applied AI, strong command of retrieval and RAG architecture, and deep expertise in designing and building agent orchestration loops.
- Own and improve retrieval and ranking over a large, multilingual legal corpus, designing, measuring, and iterating on RAG pipelines.
- Design and build agentic systems from the ground up, including orchestration loops, tool/function calling, short- and long-term memory, state management, planning, and multi-step reasoning.
- Design and run rigorous evaluations and benchmarks with a critical understanding of the methods themselves, ensuring high standards for hallucinations and citation accuracy.
### ⚙️ Preferred Experience and Skills
- **Applied AI Depth:** Proven track record shipping LLM-powered features to production — API integration (OpenAI, Anthropic, Mistral) and open-source models (Hugging Face) — with clear, measurable impact.
- **Retrieval & RAG:** Strong command of vector databases, retrieval and ranking, and RAG architecture, including how to actually improve relevance and evaluate it.
- **Agentic Systems (critical requirement):** Deep, hands-on expertise designing and building agent orchestration loops — tool/function calling, memory, planning, error recovery, and state management. You must be able to point to something you have actually built and shipped in this space (custom architectures or frameworks like LangGraph, LangChain, LlamaIndex, AutoGen). This is more important to us than RAG experience alone; strong RAG skills without proven agent-building experience will not be sufficient.
- **Evaluation & Benchmarking (critical requirement):** Deep, critical understanding of evaluation methodology — designing benchmarks, choosing and interpreting metrics, and building automated eval pipelines that reliably drive accuracy improvements. You should be able to reason about the limits of a given eval method, not just run one.
- **Fine-Tuning & Synthetic Data (critical requirement):** Demonstrated experience fine-tuning models (LoRA/QLoRA), preparing training data, and generating synthetic data — plus the judgment to know when fine-tuning is and isn't worth it.
- **Engineering Foundations:** Strong Python for AI and backend work, with enough full-stack fluency (TypeScript/Node, some React) to integrate AI into product surfaces end to end.
- **Systems Judgment:** Scalable API design and a habit of building for graceful degradation, testability, and reproducibility.
- **Senior Behaviours:** Turns vague business problems into technical roadmaps in collaboration with product, design, and stakeholders; communicates results clearly to non-technical audiences; mentors peers; and takes accountability for outcomes.
**Nice to Have**
- Experience in legal-tech or another high-stakes, accuracy-critical domain.
- Multilingual NLP / retrieval experience.
- Model inference optimization (vLLM, TensorRT-LLM).
- Cloud infrastructure (AWS, GCP, or Azure) and containerization.
- Contributions to open-source AI projects or a strong portfolio of shipped AI applications.
This listing was posted by a verified recruiter at Jus Mundi. Report this listing
JobSpring