Skip to content
← Back to job listings

Sr. Product Manager - Observability

Workday, Inc. · Pleasanton, CA, United States

Business StrategyExternal listingfull-timeabout 3 hours ago

About The Role

Your work days are brighter here.

We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too.

About the Team

The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, ClickHouse, Tempo, Mimir, Grafana, S3, Kafka, and Elasticsearch — serving traces, metrics, and logs for every workload at Workday.

How we help our customers: Every engineering, SRE, and product team at Workday depends on us to see inside their systems — from a single service's latency spike to a cross-service cascading failure impacting thousands of customers. By owning the full observability data lifecycle at petabyte scale, we give internal teams the speed and confidence to find and fix issues before they affect Workday's customers. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale — turning observability from a reactive debugging tool into a proactive, AI-assisted safety net for every workload running on Workday's platform.

About the Role

We are looking for a hands-on, technical Senior Product Manager (P4) to own and drive our Observability strategy, with a strong emphasis on Distributed Tracing . You will define the product vision for how engineers understand, debug, and optimize complex distributed systems, with a particular focus on building AI-powered detection and triage capabilities that reduce time-to-detect and time-to-resolve production issues. This role requires someone comfortable diving deep into technical architecture discussions, reading code/traces, and partnering closely with engineering and applied ML teams to ship technically sound, high-impact products.

What You'll Do

  • Own the product vision, strategy, and roadmap for Observability, with a primary focus on Distributed Tracing capabilities (trace context propagation, sampling strategies, span analysis, service maps, latency/error analysis, etc.)
  • Define and drive the roadmap for AI-enabled anomaly detection , including specifying requirements for statistical and ML-based detection methods (e.g., time-series forecasting, seasonality-aware baselining, change-point detection, multivariate anomaly detection across correlated metrics/traces/logs)
  • Partner with ML engineering to define model requirements, evaluation metrics (precision/recall, false-positive rate, alert-to-incident ratio), and feedback loops for continuous model improvement
  • Define requirements for LLM-based root cause summarization and triage assistance — e.g., generating human-readable incident summaries from raw trace/log/metric data, ranking probable root causes, suggesting remediation runbooks based on historical incident patterns
  • Specify how confidence scores, explainability, and human-in-the-loop review are surfaced in the triage workflow so on-call engineers can trust and act on AI-generated recommendations
  • Partner closely with engineering teams to define technical requirements, evaluate architectural trade-offs, and make hands-on contributions to product decisions (e.g., reviewing design docs, understanding OpenTelemetry/OTel standards, tracing protocols, and instrumentation approaches)
  • Define and track success metrics (detection precision/recall, MTTD, MTTR, alert noise reduction, triage automation rate, trace coverage) to measure product impact
  • Conduct customer and internal stakeholder research to identify pain points in debugging, alert fatigue, and incident triage
  • Write clear, detailed product requirements, user stories, and specs; work closely with design, engineering, and data science to bring them to life
  • Stay current on industry trends in observability (OpenTelemetry, eBPF-based tracing, service mesh telemetry) and applied AI/ML for anomaly detection, incident correlation, and LLM-based operational tooling

About You

Basic Qualifications

  • 8+ years of product management experience, with meaningful time spent in Observability, Monitoring, APM, or related infrastructure/developer tooling domains
  • Deep, demonstrable expertise in Distributed Tracing concepts and technologies (e.g., OpenTelemetry, Jaeger, Zipkin, trace sampling, span/context propagation)
  • Strong technical background — comfortable reading code, understanding system architecture, APIs, and engaging directly in technical design discussions
  • Concrete, hands-on understanding of AI/ML techniques applied to detection and triage , including:
  • Anomaly detection methods (statistical thresholds, time-series forecasting models, change-point/seasonality detection)
  • Alert correlation and clustering techniques (embedding/vector similarity, graph-based dependency analysis, unsupervised clustering)
  • LLM-based summarization and reasoning for incident root-cause analysis and runbook suggestion
  • Model evaluation frameworks (precision/recall, false-positive/false-negative trade-offs, drift monitoring)
  • Experience partnering directly with ML engineering teams to translate detection and triage requirements into shippable models and features
  • Proven track record of shipping complex, technical products from concept to launch
  • Excellent written and verbal communication skills; ability to translate complex technical and ML concepts for varied audiences
  • Strong analytical and data-driven decision-making skills

Other Qualifications

  • Prior experience as a software engineer, SRE, ML engineer, or data scientist before transitioning to product management
  • Familiarity with observability platforms (Datadog, New Relic, Honeycomb, Grafana, Splunk, Dynatrace, etc.) and their AI/ML-driven detection features
  • Direct experience productionizing ML models (e.g., anomaly detectors, classifiers, LLM-based tools) for operational/incident-management use cases
  • Familiarity with vector databases/embedding search, graph-based correlation techniques, or LLM prompt/agent design for triage automation
  • Background in large-scale, cloud-native, microservices-based systems

Workday Pay Transparency Statement

The annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday’s comprehensive benefits, please click here .

Primary Location: USA.CA.Pleasanton

Primary Location Base Pay Range: $168,000 USD - $252,000 USD

Additional US Location(s) Base Pay Range: $140,600 USD - $252,000 USD

Our Approach to Flexible Work

With Flex Work, we’re combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.

Pursuant to applicable Fair Chance law, Workday will consider for employment qualified applicants with arrest and conviction records.

Workday is an Equal Opportunity Employer including individuals with disabilities and protected veterans.

Workday is committed to providing reasonable accommodations for qualified individuals during our application process, in order to perform one or more essential functions of their job, as well as regarding the use of AI tools for employment decision-making to any degree. Please see below for more details including how to request an accommodation as a qualified veteran, due to a disability or for religious reasons, or as otherwise provided under applicable law.

Workday prohibits taking adverse action against any candidate or employee for reporting a possible violation of this policy, requesting one or more work accommodations, exercising a privacy right, or cooperating in an investigation in accordance with applicable law. Any employee who retaliates against a candidate or employee for doing so may be subject to disciplinary action, up to and including termination of employment, to the fullest extent allowable under applicable law.

If you require a reasonable accommodation, you may email accommodations@workday.com , as far in advance as possible.

Are you being referred to one of our roles? If so, ask your connection at Workday about our Employee Referral process!

At Workday, we value our candidates’ privacy and data security. Workday will never ask candidates to apply to jobs through websites that are not Workday Careers.

Please be aware of sites that may ask for you to input your data in connection with a job posting that appears to be from Workday but is not.

In addition, Workday will never ask candidates to pay a recruiting fee, or pay for consulting or coaching services, in order to apply for a job at Workday.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing