Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.

Filter jobs

Location
Show all 72 location options
Language
Show fewer
Field
Level: Senior

131 open roles matching these filters · page 1 of 6

  • Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.

    $90 – $110 / HourWorldwide
  • Economists with a Master's or PhD and a year or more at a top research institution (World Bank, IMF, the Fed, a graduate school) work on a research project for a leading foundation-model AI lab. Remote contract at $120–150/hour, at least 10 hours a week for a minimum of four weeks.

    $120 – $150 / HourWorldwide
  • Condensed matter PhDs create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Nineteen narrow research areas, from bosonization to SYK. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • AMO physicists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, including levitated optomechanics, cavity QED and ultracold atoms in optical lattices. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • Researchers who have published on stochastic autocatalytic growth create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark. A single narrow area: chemical master equations, branching processes and reaction-network moments. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • High energy and nuclear theorists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, from AdS/BCFT to quasi-PDFs and dark photon searches. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • Mathematical physicists create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark, where the standard of proof sits closer to mathematics than physics. Four narrow areas, from hypergeometric identities to Fefferman-Graham geometry. Remote hourly contract at $80–110 per hour.

    $80 – $110 / HourWorldwide
  • Hands-on structural, thermal, mechanical design and dynamics engineers review and write hard engineering problems about real hardware (loads, margins, heat transfer, vibration, tolerances) for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour.

    $100 – $120 / HourWorldwide
  • Hands-on systems, integration, reliability and manufacturing test engineers review and write hard engineering problems about real hardware (V&V, qualification, FMEA, root-cause analysis) for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour, weekly via Stripe or Wise.

    $100 – $120 / HourWorldwide
  • Embedded firmware, FPGA/RTL, flight software and hardware test automation engineers review and write hard problems about software that runs on real hardware (timing, race conditions, interfaces, bring-up) for an AI research initiative. 5+ years, 3 recent hands-on. Remote contract, $100–120/hour.

    $100 – $120 / HourWorldwide
  • Hands-on RF, power electronics, analog/mixed-signal, PCB and signal integrity engineers review and write hard problems about real hardware designs, measurements and bring-up for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour, weekly via Stripe or Wise.

    $100 – $120 / HourWorldwide
  • Evaluate frontier AI responses on grey-area and policy-sensitive topics (misinformation, political persuasion, self-harm, violence, cyber, biosecurity), apply safety rubrics and write structured feedback. 5+ years in trust and safety, journalism, policy, research or security. US, UK and most of Europe. $60–70/hour.

    $60 – $70 / HourOpen to Albania, Austria and 38 more countries
  • Senior civil engineers build infrastructure evaluation tasks for AI: realistic design, permitting and construction scenarios, reference calculations and rubrics, on either a US (ASCE, ACI, AISC, AASHTO) or International (Eurocodes, ISO) standards track. 5+ years, PE or equivalent strongly preferred. Remote hourly contract at $70–80/hour.

    $70 – $80 / HourWorldwide
  • Certified explosives specialists, forensic analysts and licensee inspectors red-team frontier AI models: write benign, dual-use and adversarial prompts from casework, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Generate, structure and evaluate expert ALD and thin-film data for a frontier AI lab building semiconductor and physical-science models: solve hard problems, rate model reasoning, and turn recipes into model-ready data. US-based, 10–40 hours a week, $84/hour.

    $84 / HourOpen to United States
  • Full-time W-2 role through Cincinnatus LLC, embedded with a leading AI lab: senior board-certified physicians write golden solutions, instruction specs and clinical benchmarks for frontier models. Hybrid in the Bay Area, 40 hours a week for an initial 6 months, 4+ years post-residency and a US licence. $70–110/hour.

    $70 – $110 / HourHybridOpen to United States
  • Interpret dermatology cases from images and history, annotate lesions to a schema, grade AI assessments and write the criteria they are judged by. Non-clinical, shared expert pool, open worldwide. Board-certified or board-eligible, 3+ years post-residency, 15 hours a week minimum. A flat $270/hour.

    $270 / HourWorldwide
  • A short, well-paid sprint for very senior software engineers: help a leading foundation-model lab improve its models on hard SWE tasks. 10+ years at top US tech firms, about 20 hours a week for 2–3 weeks. $150–210/hour by geography and level; the 2-hour vetting exercise is paid $100.

    $150 – $210 / HourWorldwide
  • Regulatory counsel and compliance lawyers design AI evaluation scenarios, compliance memos, filings and rubrics across financial services, FDA, antitrust, privacy and energy regulation. US (APA, SEC, FDA, FTC) or EU/UK track. Remote hourly contract at $90–100/hour; 5+ years in regulatory practice.

    $90 – $100 / HourWorldwide
  • Patent attorneys, patent agents and IP counsel build AI evaluation scenarios on prosecution, licensing, FTO and IP litigation, with reference applications, office action responses and rubrics. US (35 U.S.C., MPEP) or EPC/PCT track. Remote hourly contract at $90–100/hour; 5+ years in IP practice.

    $90 – $100 / HourWorldwide
  • Join a leading AI lab's research team as its accounting and audit specialist: review model outputs for misapplied standards, write instruction specs and golden solutions, and design benchmarks. Needs an active CPA, CA, CIA, CFE or CMA and 4+ years; bookkeeping-only roles do not count. Full-time W-2, hybrid Bay Area, $60–100/hour.

    $60 – $100 / HourHybridOpen to United States
  • Build enterprise legal evaluation tasks for AI: realistic Fortune 500 scenarios, model-grade reference work and criterion-referenced rubrics. For in-house counsel or AmLaw 100 lawyers with an active bar admission, US-based, 20+ hours a week. $110–150/hour.

    $110 – $150 / HourOpen to United States
  • Residency-trained physicians in any specialty write grading criteria, evaluate multi-turn clinical dialogues and annotate clinical reasoning for AI systems. Shared pool with several workstreams. US only, 3+ years post-residency, 20 hours a week minimum, $150/hour.

    $150 / HourOpen to United States

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.