Skip to content
← Back to job listings

AI Policy Generalist - Seattle Onsite

handshake · Seattle, WA

External listingfull-time2 days ago

About The Role

About HandshakeHandshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational <institutions.In> 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.Why join Handshake now:Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feelPartner hand-in-hand with world-class AI labs, Fortune 500 partners and the world’s top educational institutionsWork together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC foundersBuild a massive, fast-growing business with billions in revenueAbout Handshake AIHuman data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.About the RoleAs an AI Policy Generalist, you will turn complex customer policies into consistent, well-reasoned evaluations of AI model <behavior.You> will read user requests, model responses, and relevant conversation history, then determine which policy category best applies. The most interesting cases will not have obvious answers. Two examples may look almost identical until a single word, contextual detail, or difference in intent changes the correct classification.We are looking for people who enjoy splitting hairs in a healthy way. You form clear opinions, explain precisely why two cases should be treated differently, challenge interpretations respectfully, and change your mind when better evidence emerges. You understand that productive disagreement is not about winning an argument. It is how a team finds the most accurate and consistent interpretation.This is not rote annotation. Policies cannot anticipate every possible edge case, and good evaluators do not apply them mechanically. You will balance the policy’s text and intent with customer expectations, conversation context, precedent, and team calibration.The subject matter will vary. One project may involve distinguishing benign assistance from meaningful facilitation of harm. Another may require evaluating whether an interaction reflects ordinary emotional support or unhealthy reliance. A third may focus on nuanced boundaries within sexual-safety policy. Success requires learning each customer’s framework on its own terms rather than carrying assumptions from one domain into another.What You Will DoLearn new customer policies, definitions, taxonomies, and evaluation rubrics quicklyEvaluate user requests and AI model responses within the full relevant conversation contextDistinguish between closely related labels, severity levels, and policy boundariesSelect the most defensible classification when a case is genuinely ambiguousWrite concise, evidence-based rationales that cite relevant policy language and conversation detailsIdentify policy gaps, contradictions, unclear definitions, and emerging edge casesRaise thoughtful questions when existing guidance does not resolve a caseParticipate actively in calibration discussions with evaluators, project leads, policy teams, and researchersChallenge interpretations respectfully and update your judgment when new guidance or stronger reasoning emergesApply customer policy consistently without substituting personal beliefs for the policy standardMaintain accuracy and attention to detail across repeated evaluationsIncorporate feedback quickly and apply clarified guidance to future workHelp improve evaluation frameworks, examples, decision rules, and quality standardsMove effectively between projects covering different policy domains and customer needsYou May Be a Fit IfYou enjoy making precise distinctions between cases that other people might consider equivalentYou notice when one word, contextual detail, or change in intent materially affects the answerYou can hold a strong opinion without becoming attached to being rightYou explain judgment calls clearly enough that another person can audit your reasoningYou ask productive questions when a policy is ambiguous instead of guessing or forcing certaintyYou can separate your personal views from the standard a customer has asked you to applyYou are comfortable discussing disagreement directly, respectfully, and without making it personalYou can follow the letter of a policy while also understanding its purpose and underlying logicYou remain careful and consistent during repetitive, feedback-heavy workYou learn unfamiliar subject matter quickly and know when additional context is neededYou are intellectually curious, self-directed, and comfortable working in a fast-changing environmentYou communicate clearly and precisely in writingYou treat sensitive information and difficult subject matter with maturity and sound judgmentStrong candidates may come from quality assurance, research, editing, law, teaching, operations, trust and safety, content moderation, social science, policy, investigations, compliance, customer support, or other fields that require careful interpretation and defensible decision-making. We care more about how you reason than where you learned to reason.Nice to HaveExperience evaluating or comparing outputs from ChatGPT, Claude, Gemini, or other language models in a professional capacityPrior work in AI evaluation, data annotation, RLHF, model quality, trust and safety, policy operations, or content moderationExperience applying detailed rubrics, taxonomies, regulatory language, editorial standards, or quality frameworksFamiliarity with calibration sessions, inter-rater agreement, quality audits, or adjudication workflowsExperience writing policy guidance, decision trees, evaluation examples, or structured rationalesComfort working with long conversations, incomplete context, and conflicting evidenceFamiliarity with AI safety, responsible AI, or the ways language models can assist, mislead, or cause harmPrior AI evaluation experience is helpful, but it is not required.Sensitive-Content NoticeThis role involves regular and deliberate engagement with sensitive material. Depending on the project, evaluations may include sexual content, emotional distress, self-harm, suicide, violence, weapons, abuse, exploitation, discrimination, and other potentially disturbing subjects.The work is conducted within structured evaluation frameworks and professional guidelines. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality.Role DetailsLocation: Seattle, WA Compensation: $45-55Employment classification: W-2Schedule: 8AM - 5PM PTWeekly commitment: M-F

This is an external listing. JobSpring does not represent or verify the employer. Report this listing