Staff Software Engineer - AI
Outreach · Hyderabad, India
About The Role
About Outreach
Outreach, founded in 2014, is the only complete agentic AI platform for revenue teams. Outreach infuses agentic AI, conversation intelligence, and assistive AI to power hundreds of use cases across revenue motions. From new logo prospecting to expansions, deal acceleration, driving retention, and forecasting, Outreach AI automates workflows and frees sellers to focus on more strategic conversations and actions. Revenue leaders benefit from connected account visibility, performance insights, and higher forecasting accuracy across every GTM team. World leading enterprise organizations use Outreach to power their revenue teams, including Databricks, SAP, Siemens, and Verizon to name a few.
The Team
At Outreach, we build the technology that powers the world’s leading sales execution platform — and over the last year, we have been moving fast to bring AI to the center of how revenue teams work. We have shipped agents that research accounts, personalize outreach, run meetings, and drive revenue workflows end-to-end. With Ask Outreach, we have built a fully agentic, conversational platform on LangGraph that lets users interact with their Outreach data and workflows in entirely new ways. As we scale the depth and breadth of our AI platform, quality is not an afterthought — it is foundational.
The Role
We are seeking a Staff-level engineer to own quality for our GenAI platform and agent ecosystem. This is a high-impact, strategic role where you will define and lead testing practices across a rapidly evolving agentic platform — including the agents themselves, the tools they call, the LangGraph orchestration layer, and the underlying ML pipelines and data flows.
This Staff AI Test Engineer who is first and foremost an exceptional quality engineer, and who brings a genuine curiosity and working understanding of how AI and LLM-based systems behave, fail, and improve.
This role requires someone who understands the unique challenges of testing AI systems: outputs are not always deterministic, correctness is often contextual, and traditional pass/fail assertions are insufficient on their own. You will design and implement evaluation frameworks that combine deterministic validation with LLM-based grading, establish quality standards for agent behavior, and partner closely with Data Science, Engineering, and Product teams to make quality a shared discipline.
You will be a senior voice in how we build, ship, and continuously improve AI products at Outreach.If you are passionate about building rigorous test strategies for complex, probabilistic systems at scale, we want to talk to you.
Your Daily Adventures Will Include
- Leading the architecture, design, and delivery of distributed cloud-native applications capable of high concurrency and demanding real-time data needs.
- Designing and building production-grade data pipelines and ETL/ELT workflows — modeling data for both transactional (OLTP) and analytical (OLAP) use, and orchestrating them with modern workflow tooling.
- Integrating ML and GenAI capabilities into product features — from model serving and evaluation to LLM-powered enrichment, retrieval (RAG), and intelligent automation within our services.
- Championing data quality and correctness — building validation, observability, and testing into every step of the pipeline so downstream analytics and AI features can be trusted.
- Collaborating with data science, product, and engineering partners to ship intelligent, complex product features.
- Setting and promoting engineering standards for code quality, security, and operational excellence; nurturing automation and continuous improvement.
- Diagnosing and eliminating performance bottlenecks and proactively addressing reliability risks.
- Mentoring, reviewing code/architectures, and fostering a culture of rapid learning.
- Decomposing legacy systems into SOA/microservices, resolving tech debt, and evolving the architecture for scale.
- Taking end-to-end ownership of major initiatives from planning through impact.
What Sets This Role Apart: Data & AI Engineering
This role blends strong backend engineering with hands-on data and AI capability. You will be expected to grow the team’s fluency in these areas, so we look for depth (or a clear trajectory) across:
- Data modeling & storage: schema design, query optimization, and knowing when to reach for RDBMS, NoSQL, OLAP, or OLTP stores.
- Modern data stack: Spark / Delta Lake, Databricks or Snowflake, dbt for transformation, and Airflow (or similar) for orchestration.
- AI/ML in production: deploying and monitoring models, feature pipelines, and — increasingly — GenAI/LLM application patterns (embeddings, vector search, RAG, prompt/context engineering, evaluation).
- Analytics enablement: building the data models, frameworks, and artifacts that make trustworthy analytics and dashboards easy for others to build on.
Our Vision of You
Required Qualifications
- Demonstrated excellence designing and operating large-scale distributed systems with cloud service-oriented architecture.
- Proven leadership in fast-paced environments, setting standards, and inspiring technical teams to exceed delivery goals.
- Mastery in backend programming (Go required; Python strongly valued for data/AI work; Java/Ruby a plus) and hands-on with distributed data platforms (Kafka, RabbitMQ, NoSQL).
- Hands-on data engineering: building and operating data pipelines/ETL, data modeling and schema design, and a rigorous focus on data quality and correctness.
- Experience building APIs and analytics/data infrastructure, and deploying ML algorithms in production.
- Excellent communication and cross-team collaboration skills.
- Commitment to security, compliance, and robust, scalable design.
- Growth mindset — always learning, always elevating the technical bar for the team.
Strong Pluses
Not required on day one, but a meaningful differentiator — and where we most want this hire to help the team grow.
- Experience with the modern data stack: Spark/Delta Lake, Databricks or Snowflake, dbt, and Airflow.
- Hands-on GenAI/LLM application experience: RAG, vector databases, embeddings, prompt/context engineering, and model/output evaluation.
- MLOps exposure: feature stores, model serving, and monitoring in production.
- Experience shaping data as a product — dashboards, semantic/metrics layers, and analytics enablement for other teams.
This listing was posted by a verified recruiter at Outreach. Report this listing
JobSpring