AI safety and policy review jobs
Overview · 1 week ago
AI safety and policy review jobs pay you to judge whether a model handled a sensitive request correctly under a written policy. What one task involves, who hires, and what the listings on this site pay.
What AI safety and policy review jobs are
AI safety and policy review jobs pay you to decide whether a model handled a sensitive request correctly under a written policy. The model has already answered; you judge whether it helped where it should have, declined where it should have, and got the facts right along the way. It is the defensive counterpart to red teaming, and the two overlap: many listings ask for both. Where it sits among other kinds of AI training work is covered in what AI training work actually is.
What one task looks like
Mercor's AI Safety Practitioner listing is the clearest description on the board. It concerns "grey-area" topics, where a question is neither clearly fine nor clearly forbidden. For each model response you:
- Assess it for safety, factual accuracy, policy compliance and overall quality.
- Flag unsafe outputs, hallucinations, reasoning failures and policy violations.
- Write structured feedback explaining the verdict.
- Help refine the rubrics that other reviewers apply.
The core skill is consistency. Two reviewers looking at the same borderline response should reach the same verdict for the same reasons, which means real time reading the policy and calibrating against examples before your judgements count.
Language versions add a classification step. On Mercor's Bilingual Ukrainian Generalist Expert listing you write prompts in Ukrainian, sort prompts and whole conversations into categories using written guidelines, flag phrasings that try to get around a model's safeguards, and record your reasoning for each call. The STEM (PhD) versions do the same on scientific subjects, where the judgement is about telling legitimate research and teaching questions from requests that should be declined.
Who hires for it
On this site it is overwhelmingly one platform. Of the 699 listings live on 29 September 2026, 63 fall under our AI Safety & Evaluation field, and Mercor posted 61 of them; micro1's two are salaried research engineering roles rather than review work. Twenty listings name policy review, content policy, content classification, trust and safety or moderation as a skill.
What it pays
- Bilingual generalists: 11 listings at $18–62 an hour, all pitched at entry level. The Ukrainian listing states that no AI background is needed and that Mercor trains you on the workflow.
- Bilingual STEM (PhD): 12 listings at $23–81 an hour, priced by language. Hindi sits at $23–27, Korean at $63–67, Norwegian at $77–81.
- Experienced practitioners: $60–70 an hour for the AI Safety Practitioner role.
- Domain safety panels: Mercor's specialist panels pay $65–75 per task, with no stated task length.
Mercor's safety listings state weekly payment via Stripe or Wise.
Before you apply
The content is heavy by design. The practitioner listing names self-harm, violence, misinformation and political persuasion among its topics, and it does not mention wellbeing support. If reviewing that material at volume would wear on you, weigh that before the rate.
The adjacent work, grading whether AI descriptions of video clips follow guidelines and flagging clips that need a ruling, pays far less: micro1's clip review roles list $8–10 an hour.
Browse current roles under AI Safety & Evaluation and Law, Policy & Security.
Questions
- What does an AI safety policy reviewer do?
- You assess model responses on sensitive topics for safety, factual accuracy and policy compliance, flag unsafe outputs, hallucinations and policy violations, write structured feedback, and in some roles classify prompts and conversations against written guidelines.
- How much do AI safety review jobs pay?
- On this site on 29 September 2026, Mercor's bilingual generalist AI safety roles paid $18 to $62 an hour, its bilingual STEM PhD roles $23 to $81 depending on language, its AI Safety Practitioner role $60 to $70, and its specialist domain panels $65 to $75 per task.
- Is AI safety review work emotionally difficult?
- It can be. Mercor's AI Safety Practitioner listing names self-harm, violence, misinformation and political persuasion among its topics and does not mention wellbeing support. Weigh that before applying.
Platforms covered here
Put this into practice
Every listing shows its pay and who it is open to.