Skip to content
Labeling Jobs

Science & Math AI training jobs

Open AI training, data labeling, and annotation roles in Science & Math, with the pay each platform reports and the countries it accepts.

Open roles

  • Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.

    $90 – $110 / HourWorldwide
  • Condensed matter PhDs create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Nineteen narrow research areas, from bosonization to SYK. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • AMO physicists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, including levitated optomechanics, cavity QED and ultracold atoms in optical lattices. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • Researchers who have published on stochastic autocatalytic growth create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark. A single narrow area: chemical master equations, branching processes and reaction-network moments. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • High energy and nuclear theorists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, from AdS/BCFT to quasi-PDFs and dark photon searches. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • Mathematical physicists create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark, where the standard of proof sits closer to mathematics than physics. Four narrow areas, from hypergeometric identities to Fefferman-Graham geometry. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • Remote hourly contract for experimental physicists, chemists and materials scientists who already use Origin or OriginPro on their own Windows PC for graphing and curve fitting. $45–55/hour, paid weekly via Stripe or Wise. Bring your own licence and a screen above 2.5 megapixels.

    $45 – $55 / HourWorldwide
  • Build code-based drug-discovery benchmark tasks in Docker, curate ChEMBL and PubChem data, and review AI output on SAR, ADMET and docking. Real Python engineering is required, not just analysis scripts. Contractor, remote, 50 openings, $80–110/hour.

    $80 – $110 / HourWorldwide
  • Frontier benchmark work on the PXP model, Rydberg blockade and quantum many-body scars: symmetry-resolved exact diagonalization (QuSpin preferred) at L ≥ 26, block-diagonalization and Z2-state overlaps. Ten openings, contractor, $100–170/hour, remote.

    $100 – $170 / HourWorldwide
  • Write and review advanced problems, solutions and proofs (calculus and beyond) for AI training data, with research-level rigour. Contractor, remote, open worldwide in tone, 100 openings, $80–90/hour. A BSc, MSc or PhD in mathematics and a record of proofs or papers.

    $80 – $90 / HourWorldwide
  • A single opening on a frontier gravity benchmark: first-order Palatini gravity with tetrads and spin connections, the Pontryagin density, dynamical Chern-Simons gravity, Friedmann equations with torsion and quadratic inflation, integrated numerically in Planck units. Contractor, $80–160/hour, remote.

    $80 – $160 / HourWorldwide
  • Act as the senior referee: adjudicate contested physics arguments and competing solutions, say which approach holds under which assumptions, and write evaluations fit for review by other senior physicists. For associate or full professors and PIs with an active research programme. Six openings, contractor, $80–160/hour.

    $80 – $160 / HourWorldwide
  • Build and evaluate training data for a frontier lab's materials science models: DFT, AIMD, classical MD, surface and adsorption modeling, reaction energetics. For US-based computational PhDs fluent in VASP, Quantum ESPRESSO, CP2K, LAMMPS or ASE. Long-term, 10–40 hours a week, $84/hour.

    $84 / HourOpen to United States
  • Materials science PhDs author original, executable research problems for a scientific-computing AI benchmark, with depth in both semiconductor materials and molecular modeling. Tasks ship only when frontier models fail them more often than not. 6 weeks, 20+ hours a week, Git and Docker workflow. $70/hour, 212 hired this month.

    $70 / HourWorldwide
  • Ukrainian-speaking PhD chemists and biologists write specialised science prompts in Ukrainian and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $48–52/hour, $10 above the Ukrainian generalist role. Eastern Europe preferred, not required; 11 hires this month.

    $48 – $52 / HourWorldwide
  • The reviewer seat on micro1's physics work: critique derivations and arguments written by researchers or AI, find the errors and unjustified steps, and write precise feedback, cross-checking in SymPy and Python. Postdocs and junior faculty, US, Canada and UK focused. 30 openings, contractor, $80–150/hour.

    $80 – $150 / HourWorldwide
  • Build life-science benchmark tasks from real lab protocols and raw data (design critiques, protocol analysis, data interpretation) and write the rubrics that grade AI agents. Paid per accepted task, quoted at $30–50/hour. Bachelor's or Master's plus two years of lab research. 50 openings.

    $30 – $50 / HourWorldwide
  • Japanese-speaking PhD chemists and biologists write specialised science prompts in Japanese and grade AI answers for accuracy and dual-use safety. Part-time remote at $68–72/hour, second-highest in the series and $20 above the Japanese generalist role. East Asia preferred, not required.

    $68 – $72 / HourWorldwide
  • Explain physics to an AI and grade what it writes back: write clear explanations of hard concepts, build physics training material, and review model output for conceptual accuracy. A physics PhD is the preferred bar. 100 openings, contractor, $70–90/hour, fully remote.

    $70 – $90 / HourWorldwide
  • Research-level AI training work on magnetic order: magnetic structure factors, propagation vectors, magnetic space groups in BNS notation checked against MAGNDATA, AFM/FM classification, neutron scattering and MOKE signatures. Ten openings, contractor, $80–160/hour, remote.

    $80 – $160 / HourWorldwide
  • Lend lab chemistry expertise (reactions, synthesis, separations, analytical methods) to AI training data, judging experimental reasoning and explaining it plainly. Contractor, remote, 100 openings, $70–90/hour. A PhD is preferred and hands-on wet-lab experience is expected.

    $70 – $90 / HourWorldwide
  • Red-team frontier AI models on chemical safety: write benign, dual-use and adversarial prompts from exposure and process-hazard work, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Full-time Bay Area hybrid role with an AI lab's research team: review model reasoning on life sciences tasks, write golden solutions and instruction specs, and design benchmarks. For life sciences PhDs with 4+ years of substantive research experience. W-2 via Cincinnatus, $65–105/hour.

    $65 – $105 / HourHybridOpen to United States
  • Umbrella listing for Mercor's chemistry red-team panel: synthetic, analytical, forensic, defence and process safety chemists write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 5 hired this month.

    $65 – $75 / TaskWorldwide
  • Red-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from trace analysis and method validation, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Microbiologists apply knowledge of bacteria, fungi and algae, pathogenesis and antimicrobial effects to rubric-based review and documentation for AI training. Contractor, remote, 100 openings, $70–90/hour. A bachelor's degree in biology, microbiology or chemistry is the stated minimum.

    $70 – $90 / HourWorldwide
  • Write and review evaluation tasks that make an AI derive, reproduce or validate clinical trial statistics from TFLs and CDISC datasets, and catch mismatches between outputs, the SAP and the CSR narrative. Remote contractor, $60–65/hour, 30 openings; 5+ years in clinical trials and SAS or R preferred.

    $60 – $65 / HourWorldwide
  • Red-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from route design and scale-up experience, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Write and verify expert-level physics multiple-choice questions (one right answer, nine plausible wrong ones) for an AI benchmark, across semiconductors, photonics, quantum sensing, plasma, turbulence and geophysics. PhD or doctoral candidate preferred. Remote hourly contract at $61–77/hour, 10+ hours a week.

    $61 – $77 / HourWorldwide
  • PhD chemists and biologists who write fluent Finnish: author specialised science prompts and judge how AI models answer them, including where a question strays into dual-use territory. Part-time remote at $61–65/hour, $13 above the Finnish generalist role. Finland or Western Europe preferred, not required.

    $61 – $65 / HourWorldwide
  • Thai-speaking PhD chemists and biologists write specialised science prompts in Thai and grade how AI models answer them, with a focus on dual-use safety. Part-time remote at $24–28/hour, $6 above the Thai generalist role. Southeast Asia preferred, not required; 10 hires this month.

    $24 – $28 / HourWorldwide
  • Research-level benchmark work on bacterial population growth: two-state growth-rate switching with gamma-distributed waiting times, the Euler-Lotka equation, renewal theory and first-passage times. Solver, Auditor or Adjudicator roles. Ten openings, contractor, $80–160/hour, remote.

    $80 – $160 / HourWorldwide
  • Curate and annotate drug-discovery datasets, judge AI answers on medicinal chemistry and molecular analysis, and write feedback that improves the model. Contractor, remote, 30 openings, $90–120/hour. An advanced degree and cheminformatics or omics experience are the preferred background.

    $90 – $120 / HourWorldwide
  • For Danish-speaking PhD scientists in chemistry or biology: write specialised prompts in Danish, grade AI answers for accuracy and safe handling, and classify conversations against guidelines. Part-time remote at $61–65/hour. Denmark or Western Europe preferred, not required; PhD candidates eligible.

    $61 – $65 / HourWorldwide
  • Formulation and synthesis chemists from pyrotechnics or propellant work red-team frontier AI models: write benign, dual-use and adversarial prompts, grade the model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • The best-paid role in Mercor's bilingual AI safety series: Norwegian-speaking PhD chemists and biologists write specialised science prompts and grade how AI models handle dual-use questions. Part-time remote at $77–81/hour. Norway preferred, not required; PhD candidates eligible.

    $77 – $81 / HourWorldwide
  • Experimental scientists create and review training data for a frontier lab's materials science models: inorganic synthesis, superconductors, semiconductors and advanced packaging, characterization (XRD, SEM, TEM) and fabrication. PhD, MS or equivalent hands-on experience. US-based, 10–40 hours a week, $84/hour.

    $84 / HourOpen to United States
  • Full-time W-2 role through Cincinnatus LLC, embedded with a leading AI lab: senior drug development scientists write golden solutions, instruction specs and benchmarks for pharma R&D reasoning. Hybrid in the Bay Area, 40 hours a week for an initial 6 months. PhD, PharmD or MD with 4+ years in industry R&D. $75–115/hour.

    $75 – $115 / HourHybridOpen to United States
  • Croatian-speaking PhD chemists and biologists write specialised science prompts in Croatian and evaluate how AI models answer, including how they handle dual-use questions. Part-time remote at $48–52/hour, $10 above the Croatian generalist role. Eastern Europe preferred, not required; 9 hires this month.

    $48 – $52 / HourWorldwide
  • Author or verify expert multiple-choice biology questions (one correct answer, nine subtle distractors, chain-of-thought solution, references) for an AI benchmark in pharma manufacturing, synthetic biology, drug discovery and agricultural, environmental and food biology. PhD or candidate preferred. $60–75/hour, 10+ hours a week.

    $60 – $75 / HourWorldwide
  • Quantum and computational chemistry PhDs write original, runnable research problems for a scientific-computing AI benchmark, with grading criteria, calibrated until frontier models fail them more often than they pass. 6 weeks at 20+ hours a week, Git and Docker workflow. Flat $70/hour; 182 hired this month.

    $70 / HourWorldwide
  • Hands-on preclinical scientists from companies that develop their own drugs annotate R&D data and model outputs, and advise an AI lab building foundation models for drug discovery. Any modality, 5+ years industry R&D, about 10 hours a week. $60–100/hour.

    $60 – $100 / HourWorldwide
  • Preclinical scientists with direct antibody-drug conjugate or bispecific antibody experience annotate R&D data and advise an AI lab building drug discovery foundation models. In-house asset developers only, 5+ years industry R&D, about 10 hours a week. $60–100/hour.

    $60 – $100 / HourWorldwide
  • Belgium-based PhD chemists and biologists who write Belgian Dutch: author specialised science prompts and grade how AI models handle accuracy and dual-use safety. Belgium residence required. Part-time remote at $61–65/hour, $13 above the Belgian Dutch generalist role; 9 hires this month.

    $61 – $65 / HourOpen to Belgium
  • Red-team frontier AI models from the chemical defence side: write benign, dual-use and adversarial prompts drawn from countermeasures, protection and detection work, judge how models respond against a policy standard, and write the reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Korean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.

    $63 – $67 / HourWorldwide
  • Hindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.

    $23 – $27 / HourWorldwide
  • Portuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.

    $50 – $54 / HourWorldwide

See all 65

Jobs in other fields

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.