Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.
Remote AI training and data labeling jobs
Filter jobs
Location: Worldwide
- Worldwide34 jobs, applied. Activate to remove
- United States5 jobs
- United Kingdom1 job
- Canadano roles alongside your other filters
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English34 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japanese1 job
- Portuguese1 job
Field: Science & Math
- Languages & Linguistics50 jobs
- Audio & Voice28 jobs
- Engineering31 jobs
- Business & Finance26 jobs
- Software & IT16 jobs
- Health & Medicine14 jobs
- Law, Policy & Security9 jobs
- General & Data Collection31 jobs
- Science & Math34 jobs, applied. Activate to remove
- AI Safety & Evaluation57 jobs
- Video, Image & Design6 jobs
- Data, AI & ML9 jobs
- Writing & Education1 job
- Other fieldsno roles alongside your other filters
Level
- Entryno roles alongside your other filters
- Juniorno roles alongside your other filters
- Medium12 jobs
- Senior22 jobs
Newest
34 open roles matching these filters · page 1 of 2
- $90 – $110 / HourWorldwide
Condensed matter PhDs create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Nineteen narrow research areas, from bosonization to SYK. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideAMO physicists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, including levitated optomechanics, cavity QED and ultracold atoms in optical lattices. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideResearchers who have published on stochastic autocatalytic growth create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark. A single narrow area: chemical master equations, branching processes and reaction-network moments. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideHigh energy and nuclear theorists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, from AdS/BCFT to quasi-PDFs and dark photon searches. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideMathematical physicists create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark, where the standard of proof sits closer to mathematics than physics. Four narrow areas, from hypergeometric identities to Fefferman-Graham geometry. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideRemote hourly contract for experimental physicists, chemists and materials scientists who already use Origin or OriginPro on their own Windows PC for graphing and curve fitting. $45–55/hour, paid weekly via Stripe or Wise. Bring your own licence and a screen above 2.5 megapixels.
$45 – $55 / HourWorldwideMaterials science PhDs author original, executable research problems for a scientific-computing AI benchmark, with depth in both semiconductor materials and molecular modeling. Tasks ship only when frontier models fail them more often than not. 6 weeks, 20+ hours a week, Git and Docker workflow. $70/hour, 212 hired this month.
$70 / HourWorldwideUkrainian-speaking PhD chemists and biologists write specialised science prompts in Ukrainian and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $48–52/hour, $10 above the Ukrainian generalist role. Eastern Europe preferred, not required; 11 hires this month.
$48 – $52 / HourWorldwideJapanese-speaking PhD chemists and biologists write specialised science prompts in Japanese and grade AI answers for accuracy and dual-use safety. Part-time remote at $68–72/hour, second-highest in the series and $20 above the Japanese generalist role. East Asia preferred, not required.
$68 – $72 / HourWorldwideRed-team frontier AI models on chemical safety: write benign, dual-use and adversarial prompts from exposure and process-hazard work, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's chemistry red-team panel: synthetic, analytical, forensic, defence and process safety chemists write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 5 hired this month.
$65 – $75 / TaskWorldwideRed-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from trace analysis and method validation, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideRed-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from route design and scale-up experience, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideWrite and verify expert-level physics multiple-choice questions (one right answer, nine plausible wrong ones) for an AI benchmark, across semiconductors, photonics, quantum sensing, plasma, turbulence and geophysics. PhD or doctoral candidate preferred. Remote hourly contract at $61–77/hour, 10+ hours a week.
$61 – $77 / HourWorldwidePhD chemists and biologists who write fluent Finnish: author specialised science prompts and judge how AI models answer them, including where a question strays into dual-use territory. Part-time remote at $61–65/hour, $13 above the Finnish generalist role. Finland or Western Europe preferred, not required.
$61 – $65 / HourWorldwideThai-speaking PhD chemists and biologists write specialised science prompts in Thai and grade how AI models answer them, with a focus on dual-use safety. Part-time remote at $24–28/hour, $6 above the Thai generalist role. Southeast Asia preferred, not required; 10 hires this month.
$24 – $28 / HourWorldwideFor Danish-speaking PhD scientists in chemistry or biology: write specialised prompts in Danish, grade AI answers for accuracy and safe handling, and classify conversations against guidelines. Part-time remote at $61–65/hour. Denmark or Western Europe preferred, not required; PhD candidates eligible.
$61 – $65 / HourWorldwideFormulation and synthesis chemists from pyrotechnics or propellant work red-team frontier AI models: write benign, dual-use and adversarial prompts, grade the model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideThe best-paid role in Mercor's bilingual AI safety series: Norwegian-speaking PhD chemists and biologists write specialised science prompts and grade how AI models handle dual-use questions. Part-time remote at $77–81/hour. Norway preferred, not required; PhD candidates eligible.
$77 – $81 / HourWorldwideCroatian-speaking PhD chemists and biologists write specialised science prompts in Croatian and evaluate how AI models answer, including how they handle dual-use questions. Part-time remote at $48–52/hour, $10 above the Croatian generalist role. Eastern Europe preferred, not required; 9 hires this month.
$48 – $52 / HourWorldwideAuthor or verify expert multiple-choice biology questions (one correct answer, nine subtle distractors, chain-of-thought solution, references) for an AI benchmark in pharma manufacturing, synthetic biology, drug discovery and agricultural, environmental and food biology. PhD or candidate preferred. $60–75/hour, 10+ hours a week.
$60 – $75 / HourWorldwideQuantum and computational chemistry PhDs write original, runnable research problems for a scientific-computing AI benchmark, with grading criteria, calibrated until frontier models fail them more often than they pass. 6 weeks at 20+ hours a week, Git and Docker workflow. Flat $70/hour; 182 hired this month.
$70 / HourWorldwideHands-on preclinical scientists from companies that develop their own drugs annotate R&D data and model outputs, and advise an AI lab building foundation models for drug discovery. Any modality, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.