Croatian-speaking PhD chemists and biologists write specialised science prompts in Croatian and evaluate how AI models answer, including how they handle dual-use questions. Part-time remote at $48–52/hour, $10 above the Croatian generalist role. Eastern Europe preferred, not required; 9 hires this month.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide34 jobs
- United States5 jobs
- United Kingdom1 job
- Canadano roles alongside your other filters
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English41 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japanese1 job
- Portuguese1 job
Field: Science & Math
- Languages & Linguistics58 jobs
- Audio & Voice43 jobs
- Engineering40 jobs
- Business & Finance52 jobs
- Software & IT32 jobs
- Health & Medicine28 jobs
- Law, Policy & Security17 jobs
- General & Data Collection42 jobs
- Science & Math41 jobs, applied. Activate to remove
- AI Safety & Evaluation61 jobs
- Video, Image & Design9 jobs
- Data, AI & ML15 jobs
- Writing & Education1 job
- Other fieldsno roles alongside your other filters
Level
- Entryno roles alongside your other filters
- Juniorno roles alongside your other filters
- Medium13 jobs
- Senior28 jobs
Newest
41 open roles matching these filters · page 2 of 2
- $48 – $52 / HourWorldwide
Author or verify expert multiple-choice biology questions (one correct answer, nine subtle distractors, chain-of-thought solution, references) for an AI benchmark in pharma manufacturing, synthetic biology, drug discovery and agricultural, environmental and food biology. PhD or candidate preferred. $60–75/hour, 10+ hours a week.
$60 – $75 / HourWorldwideQuantum and computational chemistry PhDs write original, runnable research problems for a scientific-computing AI benchmark, with grading criteria, calibrated until frontier models fail them more often than they pass. 6 weeks at 20+ hours a week, Git and Docker workflow. Flat $70/hour; 182 hired this month.
$70 / HourWorldwideHands-on preclinical scientists from companies that develop their own drugs annotate R&D data and model outputs, and advise an AI lab building foundation models for drug discovery. Any modality, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwidePreclinical scientists with direct antibody-drug conjugate or bispecific antibody experience annotate R&D data and advise an AI lab building drug discovery foundation models. In-house asset developers only, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwideBelgium-based PhD chemists and biologists who write Belgian Dutch: author specialised science prompts and grade how AI models handle accuracy and dual-use safety. Belgium residence required. Part-time remote at $61–65/hour, $13 above the Belgian Dutch generalist role; 9 hires this month.
$61 – $65 / HourOpen to BelgiumRed-team frontier AI models from the chemical defence side: write benign, dual-use and adversarial prompts drawn from countermeasures, protection and detection work, judge how models respond against a policy standard, and write the reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideKorean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.
$63 – $67 / HourWorldwideHindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.
$23 – $27 / HourWorldwidePortuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.
$50 – $54 / HourWorldwideArabic-speaking PhD chemists and biologists write specialised science prompts in Arabic and grade AI answers for accuracy and dual-use safety. Part-time remote at $38–42/hour; Saudi Arabia or MENA preferred, not required. 13 hires this month, the most active listing in the series.
$38 – $42 / HourWorldwideUmbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.
$65 – $75 / TaskWorldwideWrite or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.
$61 – $77 / HourWorldwideAuthor executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.
$70 / HourWorldwideRed-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from forensic casework, judge the model's answers against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideThe UK posting of Mercor's molecular biology project: design primers, plasmids, gRNAs, mRNA constructs and repair templates as ground truth for a frontier model, and write the rubrics that judge them. PhD strongly preferred, first-author record expected, 20 hours a week. $70–105/hour.
$70 – $105 / HourOpen to United KingdomDesign the primers, plasmids, gRNAs, mRNA constructs and repair templates that become ground truth for a frontier model, and write the rubrics that judge sequence design quality. PhD strongly preferred, first-author record expected, 20 hours a week minimum. US only, $70–105/hour.
$70 – $105 / HourOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.