Skip to content
← Back to job listings

AI Operation Lead

Fa Errt Saasfaprod1 · Bucuresti - Ilfov, Romania

External listingfull-timeabout 1 hour ago

About The Role

The AI Operations Lead contributes hands-on to integration, deployment and monitoring of AI systems such as ML, GenAI, RAG and agentic workflows. The role focuses on a subset of products/services and ensures operational excellence, observability, and performance.

Key duties and responsibilities

  • Implement and operate integration of AI capabilities into enterprise products following standard patterns
  • Contribute to deployment of:
  • RAG pipelines
  • Copilots and AI assistants
  • Agentic workflows
  • Predictive ML services
  • Support delivery squads in integrating AI services into business applications
  • roubleshoot and resolve integration or runtime issues in production

AI Observability & Monitoring (Core focus)

  • Design and implement AI observability frameworks, including:
  • Model performance monitoring (drift, quality, hallucination signals)
  • Usage and adoption metrics
  • Latency, reliability, and system health
  • Ensure proper logging, tracing, and monitoring of AI pipelines
  • Contribute to definition of AI SLAs/SLOs aligned with business expectations
  • Support incident management and post-mortem analysis for AI systems

Cost & Performance Optimization

  • Monitor AI-related cloud consumption and inference costs
  • Optimize pipelines for efficiency (model selection, caching, orchestration)
  • Contribute to FinOps practices specific to AI workloads

Business Acumen

  • Understands operational impact of AI systems on business processes
  • Able to balance performance, cost, and quality trade-offs
  • Communicates effectively with technical and business stakeholders

Required experience & competencies

  • 5–8 years in software/ML engineering
  • Cloud (Azure), Kubernetes, Python
  • Experience with GenAI and ML systems

Technical Skills

  • Strong hands-on experience in:
  • Python, APIs, microservices architecture
  • Cloud environments (Azure preferred, AWS/GCP acceptable)
  • Kubernetes and containerized deployments
  • Experience with:
  • MLOps / LLMOps tooling
  • Monitoring/observability tools (e.g., logs, metrics, tracing)
  • Data pipelines and distributed systems
  • Understanding of:
  • GenAI / LLM systems (RAG, embeddings, prompting)
  • ML lifecycle and deployment patterns

Soft skills

  • Hands-on and problem-solving mindset
  • Ability to debug complex AI systems in production
  • Strong collaboration with engineering and product teams
  • Ability to explain technical issues clearly to non-experts
  • Proactive and continuous improvement mindset

Business acumen

  • Can adapt his/her speech to make relevant for business users
  • Can interact effectively with top management
  • Can support in produce presentations or architecture material

Required Education

  • Master’s degree (Ph. D. is a plus) in Science, Technology, Engineering, Computer Science,
  • Bachelor’s degree plus ASA or similar work experience is accepted in place of a relevant Master’s degree
  • Certifications on Cloud or Microservices or Kubernetes (CKAD) are plus.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing