AI Product Engineer
lynon · Armenia
About The Role
Lynon builds B2B software for the iGaming industry. We’re scaling an internal agentic automation platform that turns real business workflows into production systems: high-volume, high-quality output with measurable reliability and cost control.
This is a builder’s role, not a research one. You won’t write agent frameworks from scratch. You’ll compose products and automations on top of frontier models — Claude, GPT, Gemini, and others — using provider-native tool calling and structured outputs, MCP integrations, and solid engineering underneath. We care far more about what you ship and how reliably it runs than about which framework you used to get there.
The most important part of this role is safety and evaluation. Non-deterministic systems only earn a place in production when we can measure that they’re correct, safe, and cost-effective — and prove they stay that way as models and prompts change. If evals, guardrails, and observability sound like the boring part, this isn’t the right role. If they sound like the actual engineering, keep reading.
Responsibilities
Build AI products & automations
- Turn business workflows into production pipelines: generation → QA → publish → feedback loops.
- Compose multi-step agent behavior with model-native tool calling, structured outputs, and MCP-based integrations — reaching for a framework only when it clearly earns its place.
- Orchestrate across providers (Claude, GPT, Gemini, others) with retries, fallbacks, structured outputs, and per-run cost tracking.
- Develop RAG and memory layers (pgvector / Qdrant / OpenSearch + Postgres).
- Integrate agents with the tools teams already use: Jira, CRM, Notion/Confluence, repositories, analytics.
- Ship simple internal interfaces — dashboards or chat tools — so non-engineers can actually use what you build.
Keep it production-grade
- Apply reliability fundamentals: queues/workers, rate-limit handling, idempotency, and sensible cost/latency trade-offs.
- Containerize, deploy, and keep CI/CD, environments, and infra healthy.
Safety & evaluation — the core of the job
- Design evaluation for non-deterministic output: test/golden datasets, LLM-as-judge calibrated against humans, and regression suites that run on every change.
- Define and track quality, safety, cost, and latency metrics, and surface them so teams can trust — or challenge — what the systems produce.
- Build guardrails: input/output validation, content and brand safety, PII handling, compliance-aware output, and safe failure behavior.
- Own observability: tracing across multi-step runs, structured logs, and alerts when quality or cost drifts.
Mandatory Requirements
- 4–6+ years across software engineering, backend, DevOps, or automation — comfortable owning something end to end, from API to deployment.
- Hands-on experience building LLM-powered applications: RAG, tool calling, and multi-step agent workflows on top of frontier models.
- A real, demonstrable focus on evaluation and safety — you’ve measured output quality, not just eyeballed it.
- Strong reliability fundamentals: retries, queues/workers, rate limits, cost/latency trade-offs.
- Working knowledge of LLM APIs, vector search, structured outputs, and MCP / tool-calling integrations.
- Python (FastAPI a plus), Docker, CI/CD, and observability tooling.
- Pragmatic and product-minded: you reach for the simplest thing that works, and you ship.
Nice to have
- High-volume AI generation pipelines
- LangGraph or similar orchestration tools — useful context, not a requirement; we lean on model-native tooling first.
- Azure infrastructure.
- Prior work in a regulated or high-trust domain (iGaming, fintech, health) where output safety and auditability matter.
This listing was posted by a verified recruiter at lynon. Report this listing
JobSpring