Skip to content
Labeling Jobs

AI Safety & Evaluation AI training jobs

Open AI training, data labeling, and annotation roles in AI Safety & Evaluation, with the pay each platform reports and the countries it accepts.

Open roles

  • AI safety testing in Telugu and English: provoke and document jailbreaks, bias and harmful output from chat models, and judge the quality of their Telugu answers. Evaluation skill matters more than security experience. Remote hourly contract at $16–22/hour, weekly pay; H-1B and STEM OPT excluded.

    $16 – $22 / HourWorldwide
  • Odia and English AI safety work: probe chat models for jailbreaks, bias and harmful output and judge whether their Odia answers hold up. One of the lowest-resource languages in the family and the busiest Indian variant, with 88 hired this month. Remote hourly contract, $16–22/hour.

    $16 – $22 / HourWorldwide
  • Red-team AI models in European and other non-Brazilian Portuguese plus English: jailbreaks, prompt injection, bias and manipulation, logged as reproducible safety data. Brazilian Portuguese is explicitly excluded. Remote hourly contract at $29–45/hour, paid weekly.

    $29 – $45 / HourWorldwide
  • Evaluate frontier AI responses on grey-area and policy-sensitive topics (misinformation, political persuasion, self-harm, violence, cyber, biosecurity), apply safety rubrics and write structured feedback. 5+ years in trust and safety, journalism, policy, research or security. US, UK and most of Europe. $60–70/hour.

    $60 – $70 / HourOpen to Albania, Austria and 38 more countries
  • Native Japanese speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models behave safely in Japanese. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; Japan or East Asia preferred, not required; 8 hires this month.

    $48 – $52 / HourWorldwide
  • Certified explosives specialists, forensic analysts and licensee inspectors red-team frontier AI models: write benign, dual-use and adversarial prompts from casework, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Red-team conversational AI models in Norwegian and English: attempt jailbreaks, prompt injections and multi-turn manipulation, then document what broke. Prior red-teaming, security or adversarial testing experience expected. Remote hourly contract at $48–62/hour, paid weekly.

    $48 – $62 / HourWorldwide
  • Native Croatian speakers write prompts on sensitive subjects, classify prompts and conversations, and flag adversarial phrasing to harden AI models in Croatian. Bachelor's (in progress is fine) and business English. Part-time remote at $38–42/hour; Croatia or Eastern Europe preferred, not required.

    $38 – $42 / HourWorldwide
  • AI safety work in Kannada and English: test chat models for jailbreaks, bias and harmful output, and judge whether their Kannada answers are accurate and appropriate. Unlike the European variants, prior red-teaming is not listed as a requirement. Remote hourly contract, $16–22/hour, weekly pay.

    $16 – $22 / HourWorldwide
  • Red-team AI models in Dutch and English: jailbreaks, prompt injection, bias exploitation and multi-turn manipulation, written up as reproducible attack cases and labelled data. Native Dutch plus prior adversarial or security experience. Remote hourly contract at $48–62/hour, paid weekly.

    $48 – $62 / HourWorldwide
  • Engineers with propulsion, initiation or effects test experience red-team frontier AI models: write benign, dual-use and adversarial prompts, judge the replies against a policy standard, and write reference answers with the reasoning. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Umbrella listing for Mercor's nuclear red-team panel: fuel-cycle engineers, safeguards inspectors, nuclear security, forensics and nonproliferation specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 21 hired this month.

    $65 – $75 / TaskWorldwide
  • Red-team AI chat models and agents in Thai and English, then document every jailbreak, injection or biased answer as reproducible safety data. Native Thai required, plus prior adversarial or security experience. Remote hourly contract at $24–35/hour, paid weekly; 205 hired this month.

    $24 – $35 / HourWorldwide
  • Native Thai speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing to make AI models safer in Thai. Business English and a bachelor's (in progress is fine) required. Part-time remote at $18–22/hour; Southeast Asia preferred, not required.

    $18 – $22 / HourWorldwide
  • The highest-paid generalist role in Mercor's bilingual AI safety series: native Norwegian speakers write sensitive-topic prompts, classify conversations and flag adversarial phrasing. Bachelor's (in progress is fine) and business English. Part-time remote at $58–62/hour; Norway preferred, not required.

    $58 – $62 / HourWorldwide
  • Ukrainian-speaking PhD chemists and biologists write specialised science prompts in Ukrainian and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $48–52/hour, $10 above the Ukrainian generalist role. Eastern Europe preferred, not required; 11 hires this month.

    $48 – $52 / HourWorldwide
  • Radiological emergency planners, field monitoring teams and consequence modellers red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Japanese-speaking PhD chemists and biologists write specialised science prompts in Japanese and grade AI answers for accuracy and dual-use safety. Part-time remote at $68–72/hour, second-highest in the series and $20 above the Japanese generalist role. East Asia preferred, not required.

    $68 – $72 / HourWorldwide
  • Test AI models for safety failures in Finnish and English: jailbreaks, prompt injection, bias and multi-turn manipulation, logged as structured red-team data. For native Finnish speakers with prior adversarial, security or abuse-analysis experience. Remote hourly contract, $48–62/hour.

    $48 – $62 / HourWorldwide
  • Swedish and English red-teaming of AI models: jailbreaks, injected instructions, bias and multi-turn manipulation, written up as reproducible attack cases. The busiest listing in this family, with 243 hires this month. Remote hourly contract at $48–62/hour, paid weekly.

    $48 – $62 / HourWorldwide
  • Test AI models in Gujarati and English for jailbreaks, bias and harmful answers, and judge whether their Gujarati is accurate and appropriate. Evaluation judgment is the core requirement. Open to the Gujarati diaspora, with no residence rule published. Remote hourly contract at $16–22/hour; 56 hired this month.

    $16 – $22 / HourWorldwide
  • Probe AI chat models and agents for safety failures in Danish and English, then turn each failure into labelled, reproducible red-team data. For native Danish speakers with a security, adversarial ML or abuse-analysis background. Remote hourly contract, $48–62/hour, weekly pay.

    $48 – $62 / HourWorldwide
  • Licensed blasters and blasting engineers red-team frontier AI models: write benign, dual-use and adversarial prompts from quarry, mine and demolition practice, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Write sensitive-topic prompts in Ukrainian and classify prompts and conversations for an AI safety project, flagging adversarial phrasing as you go. Native Ukrainian, business English and a bachelor's (finished or in progress). Part-time remote at $38–42/hour; Eastern Europe preferred, not required.

    $38 – $42 / HourWorldwide
  • PhD chemists and biologists who write fluent Finnish: author specialised science prompts and judge how AI models answer them, including where a question strays into dual-use territory. Part-time remote at $61–65/hour, $13 above the Finnish generalist role. Finland or Western Europe preferred, not required.

    $61 – $65 / HourWorldwide
  • Thai-speaking PhD chemists and biologists write specialised science prompts in Thai and grade how AI models answer them, with a focus on dual-use safety. Part-time remote at $24–28/hour, $6 above the Thai generalist role. Southeast Asia preferred, not required; 10 hires this month.

    $24 – $28 / HourWorldwide
  • For Danish-speaking PhD scientists in chemistry or biology: write specialised prompts in Danish, grade AI answers for accuracy and safe handling, and classify conversations against guidelines. Part-time remote at $61–65/hour. Denmark or Western Europe preferred, not required; PhD candidates eligible.

    $61 – $65 / HourWorldwide
  • Formulation and synthesis chemists from pyrotechnics or propellant work red-team frontier AI models: write benign, dual-use and adversarial prompts, grade the model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Certified bomb technicians, EOD veterans and bomb squad leaders red-team frontier AI models: write benign, dual-use and adversarial prompts from public-safety practice, judge model responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • The best-paid role in Mercor's bilingual AI safety series: Norwegian-speaking PhD chemists and biologists write specialised science prompts and grade how AI models handle dual-use questions. Part-time remote at $77–81/hour. Norway preferred, not required; PhD candidates eligible.

    $77 – $81 / HourWorldwide
  • Adversarial testing of AI chat models in Bahasa Indonesia and English: jailbreaks, prompt injection, bias and multi-turn manipulation, recorded as structured red-team data. Native Indonesian plus prior red-teaming or security experience. Remote hourly contract, $17–25/hour, weekly pay.

    $17 – $25 / HourWorldwide
  • Hands-on ML researchers take on scoped, open-ended empirical problems: training image classifiers and generators from scratch, fine-tuning open-weight LLMs, adversarial robustness, compression under hard budgets, and multilingual pre-training. 3+ years of ML research (PhD counts). Remote hourly contract at $100–120/hour.

    $100 – $120 / HourWorldwide
  • Probe AI chat models and agents for safety failures in Vietnamese and English (jailbreaks, prompt injection, bias, multi-turn manipulation) and document each as reproducible data. Native Vietnamese plus prior red-teaming or security experience. Remote hourly contract, $17–25/hour.

    $17 – $25 / HourWorldwide
  • Export control, treaty and proliferation analysts red-team frontier AI models: write benign, dual-use and adversarial prompts from nonproliferation work, judge model responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Nuclear forensics, radiochemistry and detection specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from characterisation and attribution work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • MC&A, physical protection and vulnerability assessment specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from nuclear security work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • For Flemish speakers living in Belgium: write sensitive-topic prompts in Belgian Dutch, classify conversations and flag adversarial phrasing so AI models handle Belgian usage safely. Belgium residence required; bachelor's (in progress is fine) and business English. Part-time remote at $48–52/hour.

    $48 – $52 / HourOpen to Belgium
  • Croatian-speaking PhD chemists and biologists write specialised science prompts in Croatian and evaluate how AI models answer, including how they handle dual-use questions. Part-time remote at $48–52/hour, $10 above the Croatian generalist role. Eastern Europe preferred, not required; 9 hires this month.

    $48 – $52 / HourWorldwide
  • Probe AI models in Tamil and English for jailbreaks, bias and harmful output, and assess whether their Tamil answers are accurate and appropriate. Evaluation judgment is the core requirement, not prior red-teaming. Remote hourly contract, $16–22/hour, weekly pay; 57 hired this month.

    $16 – $22 / HourWorldwide
  • Native Finnish speakers write prompts on sensitive topics, classify conversations and flag adversarial phrasing so AI models stay safe in Finnish. Needs business English and a bachelor's, finished or in progress. Part-time remote at $48–52/hour; Finland or Western Europe preferred, not required.

    $48 – $52 / HourWorldwide
  • Red-team AI models in Malay and English: jailbreaks, prompt injection, bias and multi-turn manipulation, captured as labelled, reproducible safety data. Native Malay required, along with prior adversarial, security or abuse-analysis experience. Remote hourly contract at $17–25/hour, paid weekly.

    $17 – $25 / HourWorldwide
  • Fuel-cycle and criticality safety engineers red-team frontier AI models: write benign, dual-use and adversarial prompts from enrichment, fabrication, reprocessing or criticality work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Current and former IAEA, Euratom and state-system safeguards inspectors and accountancy analysts red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Test AI chat models in Malayalam and English for safety failures and judge whether their answers are accurate, complete and appropriate. Native Malayalam plus careful evaluation skills; prior red-teaming is a plus rather than the core requirement. Remote hourly contract at $16–22/hour, paid weekly.

    $16 – $22 / HourWorldwide
  • Belgium-based PhD chemists and biologists who write Belgian Dutch: author specialised science prompts and grade how AI models handle accuracy and dual-use safety. Belgium residence required. Part-time remote at $61–65/hour, $13 above the Belgian Dutch generalist role; 9 hires this month.

    $61 – $65 / HourOpen to Belgium
  • Native Danish speakers write sensitive-topic prompts, classify prompts and conversations, and flag adversarial phrasing to make AI models safer in Danish. Business English and a bachelor's (in progress counts). Part-time remote at $48–52/hour; Denmark or Western Europe preferred, not required.

    $48 – $52 / HourWorldwide
  • Radioactive source security and vulnerability assessment specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from Category 1 and 2 source security work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Test AI chat models in Marathi and English for safety failures (jailbreaks, bias, harmful answers) and judge whether their Marathi is accurate and appropriate rather than Hindi in disguise. Evaluation judgment is the core ask. Remote hourly contract, $16–22/hour, weekly pay.

    $16 – $22 / HourWorldwide

See all 63

Jobs in other fields

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.