Skip to content
← Back to job listings

AI Evaluation Engineer

siteground · Sofia, Bulgaria

Data Science / AI / Machine LearningImported listingfull-timeabout 4 hours ago

About The Role

YOUR ROLE

At SiteGround, we’ve built and launched our own suite of AI products

Coderick AI — our vibe coding platform.

AI Studio — access to flagship AI models, providers, and intelligent AI Agents.

WordPress AI Agent — one AI assistant for virtually every WordPress task.

We’re moving fast: adding new capabilities, launching new AI agents, and serving a growing number of customers who use our AI products every day.

This role sits at the heart of how we measure and improve those products. Your job is to make our LLM-powered features provably reliable — not just impressive in a demo.

Prompt engineering is a central part of the role. You’ll design, test, and iterate on prompts and agent behaviour directly. But your defining contribution will be rigour: turning “this seems to work” into measurable, repeatable, production-ready behaviour through evaluation and observability.

We work across the major model ecosystems — OpenAI, Google Gemini, and Anthropic — and you’ll regularly make pragmatic decisions about model choice, quality, latency, and cost.

You’ll work closely with backend engineers, fellow evaluation engineers, and product managers. Together, you’ll dig into real customer interactions, identify where our AI succeeds or fails, and turn those insights into better prompts, agents, tools, and product experiences.

YOUR RESPONSIBILITIES

Evaluation & Observability

  • Build high-quality evaluation datasets with ground truth and define task-specific metrics such as correctness, task completion, efficiency, cost, token usage, and safety;
  • Run model bake-offs and before/after prompt experiments — every meaningful prompt change should be a measured experiment, not a guess;
  • Design evaluations for harder cases where there is no clean ground truth, including agentic workflows, tool-calling correctness, and structured-output validity;
  • Own production observability through trace analysis, per-tool dashboards, error classification, and session-level review;
  • Root-cause failures, identify recurring patterns, understand their impact on quality and cost, and turn findings into prompt improvements and engineering tickets;
  • Help close the loop between users and product teams by translating real customer behaviour and product questions into clear, measurable insights about how our AI features are used and where they can improve.

Prompting & Agent Engineering

  • Design, version, test, and continuously improve system prompts for domain-specific agents;
  • Architect agent behaviour, including tool orchestration, multi-turn flows, and safeguards such as plan-before-execute, read-before-write, and always-confirm for destructive actions;
  • Build reusable, modular prompt components for areas such as safety, tone, context, and tool documentation;
  • Select the right model for each task, balancing quality, latency, and cost.

Tooling & Integrations

  • Work across the API layer our AI agents act through, integrating the REST endpoints that expose application data and operations to the LLM;
  • Help design and implement guardrails that prevent agents from corrupting user data, including permission enforcement and safe defaults;
  • Debug the full request path end to end — from an agent’s decision and tool call to the live result — correlating application logs with agent traces.

OUR EXPECTATIONS

  • Strong prompt engineering and LLM application experience — you’ve shipped agents or LLM-powered features to production, not just prototyped them;
  • Hands-on evaluation experience and a data-driven mindset — you’ve built evaluation datasets and metrics, and you measure results before claiming something works;
  • Python experience and comfort consuming and integrating REST APIs;
  • Familiarity with modern LLM tooling and patterns, including function/tool calling, RAG, agentic workflows, observability platforms, and current model families;
  • Healthy scepticism toward AI output and a strong instinct for safety, failure modes, and guardrail design.

GREAT ADVANTAGE WILL BE

  • Experiment design and statistics literacy, including A/B testing, significance, and sampling, so your evaluation conclusions hold up under scrutiny;
  • Experience designing LLM-as-a-judge evaluations and calibrating automated judges against human labels;
  • Red-teaming and adversarial testing experience, including prompt injection, jailbreaks, and structured-output abuse;
  • Comfort working with data tools for exploring and slicing evaluation results, such as SQL, pandas, DuckDB, notebooks, or similar;
  • Experience with platforms/solutions such as Langfuse, Langchain, n8n, Google ADK.

WHAT WE OFFER

At SiteGround, we work hard and challenge ourselves to exceed expectations. To support your drive, we provide a thoughtfully designed benefits package created to help you thrive both at work and beyond.

  • Competitive remuneration and a performance-based bonus system;
  • Shorter workday on Fridays – we finish at 3:00 PM;
  • 5 extra days off at 5 years of service, plus 1 more each year after – up to 10;
  • 2 extra paid days off for volunteering to support the causes you care about;
  • Premium additional health insurance coverage and annual medical check-ups;
  • Modern and cozy offices across Bulgaria, plus flexible work options;
  • Chef-prepared breakfast and lunch provided daily at our HQ restaurant – all on us;
  • Free parking and a metro shuttle;
  • Fully equipped in-office gym with professional coaches;
  • On-site sports at our HQ: yoga, spinning, group conditioning, Brazilian Jiu-Jitsu, dancing, table tennis, and more;
  • Multisport or Coolfit cards fully covered by the company;
  • Free massages;
  • Memorable team events, internal gatherings, and company festivals;
  • Opportunities for professional development through trainings and conferences;
  • Knowledge-sharing culture through meetups and events;
  • Anniversary and company gifts along the way.
  • Come and make a difference at SiteGround with people who inspire you to grow!
  • Please note: Only shortlisted candidates will be contacted for further steps.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing