Operational health physicists from DOE sites, national labs, reactors and decommissioning projects red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide59 jobs
- United States2 jobs
- United Kingdom2 jobs
- Canadano roles alongside your other filters
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English63 jobs
- German1 job
- Spanishno roles alongside your other filters
- French1 job
- Japanese2 jobs
- Portuguese2 jobs
Field: AI Safety & Evaluation
- Languages & Linguistics119 jobs
- Audio & Voice114 jobs
- Engineering109 jobs
- Business & Finance107 jobs
- Software & IT88 jobs
- Health & Medicine78 jobs
- Law, Policy & Security70 jobs
- General & Data Collection68 jobs
- Science & Math65 jobs
- AI Safety & Evaluation63 jobs, applied. Activate to remove
- Video, Image & Design44 jobs
- Data, AI & ML35 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
63 open roles matching these filters · page 3 of 3
- $65 – $75 / TaskWorldwide
Korean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.
$63 – $67 / HourWorldwideNative Korean speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models stay safe in Korean. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; South Korea or East Asia preferred, not required; 7 hires this month.
$48 – $52 / HourWorldwideHindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.
$23 – $27 / HourWorldwidePortuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.
$50 – $54 / HourWorldwideNamed RSOs and radiation protection managers on broad-scope, hospital, university or industrial licences red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideAI safety testing in Urdu and English: provoke jailbreaks, bias and harmful output from chat models and judge whether their Urdu answers are accurate and appropriate, in Nastaliq script or Roman Urdu. Evaluation judgment is the core ask. Remote hourly contract at $16–22/hour, weekly pay.
$16 – $22 / HourWorldwideDesign adversarial prompts, find jailbreaks and policy failures, and document vulnerabilities in frontier AI models across cyber, biosecurity, fraud, misinformation and political content. Hourly remote contract at $70–84/hour for residents of Europe, the UK and the US; 5+ years' relevant experience required.
$70 – $84 / HourOpen to Albania, Austria and 38 more countriesNative German speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models behave safely in German. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; Germany or Western Europe preferred, not required; 8 hires this month.
$48 – $52 / HourWorldwideNative French speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models behave safely in French. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; France or Western Europe preferred, not required; 7 hires this month.
$48 – $52 / HourWorldwideArabic-speaking PhD chemists and biologists write specialised science prompts in Arabic and grade AI answers for accuracy and dual-use safety. Part-time remote at $38–42/hour; Saudi Arabia or MENA preferred, not required. 13 hires this month, the most active listing in the series.
$38 – $42 / HourWorldwideUmbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's radiological safety red-team panel: RSOs, health physicists, source security, emergency response and nuclear medicine specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 8 hired this month.
$65 – $75 / TaskWorldwideA full-time, salaried research role at micro1 designing benchmarks, rubrics, datasets and evaluation pipelines for frontier coding agents. Base salary $200,000–260,000 plus equity and benefits, remote, one opening. Three years in software engineering, ML or evaluation.
$200000 – $260000 / YearWorldwideA full-time research engineering role at micro1 building RL environments, reward functions, verifiers, synthetic data pipelines and automated evaluation systems. Base salary $200,000–300,000 plus equity and benefits, remote, one opening. Deep reinforcement learning experience required.
$200000 – $300000 / YearWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.