PhD chemists and biologists who write fluent Finnish: author specialised science prompts and judge how AI models answer them, including where a question strays into dual-use territory. Part-time remote at $61–65/hour, $13 above the Finnish generalist role. Finland or Western Europe preferred, not required.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide59 jobs
- United States2 jobs
- United Kingdom2 jobs
- Canadano roles alongside your other filters
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English63 jobs
- German1 job
- Spanishno roles alongside your other filters
- French1 job
- Japanese2 jobs
- Portuguese2 jobs
Field: AI Safety & Evaluation
- Languages & Linguistics119 jobs
- Audio & Voice114 jobs
- Engineering109 jobs
- Business & Finance107 jobs
- Software & IT88 jobs
- Health & Medicine78 jobs
- Law, Policy & Security70 jobs
- General & Data Collection68 jobs
- Science & Math65 jobs
- AI Safety & Evaluation63 jobs, applied. Activate to remove
- Video, Image & Design44 jobs
- Data, AI & ML35 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
63 open roles matching these filters · page 2 of 3
- $61 – $65 / HourWorldwide
Thai-speaking PhD chemists and biologists write specialised science prompts in Thai and grade how AI models answer them, with a focus on dual-use safety. Part-time remote at $24–28/hour, $6 above the Thai generalist role. Southeast Asia preferred, not required; 10 hires this month.
$24 – $28 / HourWorldwideFor Danish-speaking PhD scientists in chemistry or biology: write specialised prompts in Danish, grade AI answers for accuracy and safe handling, and classify conversations against guidelines. Part-time remote at $61–65/hour. Denmark or Western Europe preferred, not required; PhD candidates eligible.
$61 – $65 / HourWorldwideFormulation and synthesis chemists from pyrotechnics or propellant work red-team frontier AI models: write benign, dual-use and adversarial prompts, grade the model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideCertified bomb technicians, EOD veterans and bomb squad leaders red-team frontier AI models: write benign, dual-use and adversarial prompts from public-safety practice, judge model responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideThe best-paid role in Mercor's bilingual AI safety series: Norwegian-speaking PhD chemists and biologists write specialised science prompts and grade how AI models handle dual-use questions. Part-time remote at $77–81/hour. Norway preferred, not required; PhD candidates eligible.
$77 – $81 / HourWorldwideAdversarial testing of AI chat models in Bahasa Indonesia and English: jailbreaks, prompt injection, bias and multi-turn manipulation, recorded as structured red-team data. Native Indonesian plus prior red-teaming or security experience. Remote hourly contract, $17–25/hour, weekly pay.
$17 – $25 / HourWorldwideHands-on ML researchers take on scoped, open-ended empirical problems: training image classifiers and generators from scratch, fine-tuning open-weight LLMs, adversarial robustness, compression under hard budgets, and multilingual pre-training. 3+ years of ML research (PhD counts). Remote hourly contract at $100–120/hour.
$100 – $120 / HourWorldwideProbe AI chat models and agents for safety failures in Vietnamese and English (jailbreaks, prompt injection, bias, multi-turn manipulation) and document each as reproducible data. Native Vietnamese plus prior red-teaming or security experience. Remote hourly contract, $17–25/hour.
$17 – $25 / HourWorldwideExport control, treaty and proliferation analysts red-team frontier AI models: write benign, dual-use and adversarial prompts from nonproliferation work, judge model responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideNuclear forensics, radiochemistry and detection specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from characterisation and attribution work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideMC&A, physical protection and vulnerability assessment specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from nuclear security work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideFor Flemish speakers living in Belgium: write sensitive-topic prompts in Belgian Dutch, classify conversations and flag adversarial phrasing so AI models handle Belgian usage safely. Belgium residence required; bachelor's (in progress is fine) and business English. Part-time remote at $48–52/hour.
$48 – $52 / HourOpen to BelgiumCroatian-speaking PhD chemists and biologists write specialised science prompts in Croatian and evaluate how AI models answer, including how they handle dual-use questions. Part-time remote at $48–52/hour, $10 above the Croatian generalist role. Eastern Europe preferred, not required; 9 hires this month.
$48 – $52 / HourWorldwideProbe AI models in Tamil and English for jailbreaks, bias and harmful output, and assess whether their Tamil answers are accurate and appropriate. Evaluation judgment is the core requirement, not prior red-teaming. Remote hourly contract, $16–22/hour, weekly pay; 57 hired this month.
$16 – $22 / HourWorldwideNative Finnish speakers write prompts on sensitive topics, classify conversations and flag adversarial phrasing so AI models stay safe in Finnish. Needs business English and a bachelor's, finished or in progress. Part-time remote at $48–52/hour; Finland or Western Europe preferred, not required.
$48 – $52 / HourWorldwideRed-team AI models in Malay and English: jailbreaks, prompt injection, bias and multi-turn manipulation, captured as labelled, reproducible safety data. Native Malay required, along with prior adversarial, security or abuse-analysis experience. Remote hourly contract at $17–25/hour, paid weekly.
$17 – $25 / HourWorldwideFuel-cycle and criticality safety engineers red-team frontier AI models: write benign, dual-use and adversarial prompts from enrichment, fabrication, reprocessing or criticality work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideCurrent and former IAEA, Euratom and state-system safeguards inspectors and accountancy analysts red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideTest AI chat models in Malayalam and English for safety failures and judge whether their answers are accurate, complete and appropriate. Native Malayalam plus careful evaluation skills; prior red-teaming is a plus rather than the core requirement. Remote hourly contract at $16–22/hour, paid weekly.
$16 – $22 / HourWorldwideBelgium-based PhD chemists and biologists who write Belgian Dutch: author specialised science prompts and grade how AI models handle accuracy and dual-use safety. Belgium residence required. Part-time remote at $61–65/hour, $13 above the Belgian Dutch generalist role; 9 hires this month.
$61 – $65 / HourOpen to BelgiumNative Danish speakers write sensitive-topic prompts, classify prompts and conversations, and flag adversarial phrasing to make AI models safer in Danish. Business English and a bachelor's (in progress counts). Part-time remote at $48–52/hour; Denmark or Western Europe preferred, not required.
$48 – $52 / HourWorldwideRadioactive source security and vulnerability assessment specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from Category 1 and 2 source security work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideTest AI chat models in Marathi and English for safety failures (jailbreaks, bias, harmful answers) and judge whether their Marathi is accurate and appropriate rather than Hindi in disguise. Evaluation judgment is the core ask. Remote hourly contract, $16–22/hour, weekly pay.
$16 – $22 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.