Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.
Remote AI training and data labeling jobs
Filter jobs
Location: Worldwide
- Worldwide83 jobs, applied. Activate to remove
- United States45 jobs
- United Kingdom6 jobs
- Canada2 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English81 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japaneseno roles alongside your other filters
- Portugueseno roles alongside your other filters
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering23 jobs
- Business & Finance15 jobs
- Software & IT6 jobs
- Health & Medicine12 jobs
- Law, Policy & Security7 jobs
- General & Data Collectionno roles alongside your other filters
- Science & Math22 jobs
- AI Safety & Evaluation18 jobs
- Video, Image & Design1 job
- Data, AI & ML7 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Newest
83 open roles matching these filters · page 1 of 4
- $90 – $110 / HourWorldwide
Economists with a Master's or PhD and a year or more at a top research institution (World Bank, IMF, the Fed, a graduate school) work on a research project for a leading foundation-model AI lab. Remote contract at $120–150/hour, at least 10 hours a week for a minimum of four weeks.
$120 – $150 / HourWorldwideCondensed matter PhDs create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Nineteen narrow research areas, from bosonization to SYK. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideAMO physicists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, including levitated optomechanics, cavity QED and ultracold atoms in optical lattices. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideResearchers who have published on stochastic autocatalytic growth create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark. A single narrow area: chemical master equations, branching processes and reaction-network moments. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideHigh energy and nuclear theorists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, from AdS/BCFT to quasi-PDFs and dark photon searches. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideMathematical physicists create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark, where the standard of proof sits closer to mathematics than physics. Four narrow areas, from hypergeometric identities to Fefferman-Graham geometry. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideHands-on structural, thermal, mechanical design and dynamics engineers review and write hard engineering problems about real hardware (loads, margins, heat transfer, vibration, tolerances) for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour.
$100 – $120 / HourWorldwideHands-on systems, integration, reliability and manufacturing test engineers review and write hard engineering problems about real hardware (V&V, qualification, FMEA, root-cause analysis) for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour, weekly via Stripe or Wise.
$100 – $120 / HourWorldwideEmbedded firmware, FPGA/RTL, flight software and hardware test automation engineers review and write hard problems about software that runs on real hardware (timing, race conditions, interfaces, bring-up) for an AI research initiative. 5+ years, 3 recent hands-on. Remote contract, $100–120/hour.
$100 – $120 / HourWorldwideHands-on RF, power electronics, analog/mixed-signal, PCB and signal integrity engineers review and write hard problems about real hardware designs, measurements and bring-up for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour, weekly via Stripe or Wise.
$100 – $120 / HourWorldwideSenior civil engineers build infrastructure evaluation tasks for AI: realistic design, permitting and construction scenarios, reference calculations and rubrics, on either a US (ASCE, ACI, AISC, AASHTO) or International (Eurocodes, ISO) standards track. 5+ years, PE or equivalent strongly preferred. Remote hourly contract at $70–80/hour.
$70 – $80 / HourWorldwideCertified explosives specialists, forensic analysts and licensee inspectors red-team frontier AI models: write benign, dual-use and adversarial prompts from casework, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideInterpret dermatology cases from images and history, annotate lesions to a schema, grade AI assessments and write the criteria they are judged by. Non-clinical, shared expert pool, open worldwide. Board-certified or board-eligible, 3+ years post-residency, 15 hours a week minimum. A flat $270/hour.
$270 / HourWorldwideA short, well-paid sprint for very senior software engineers: help a leading foundation-model lab improve its models on hard SWE tasks. 10+ years at top US tech firms, about 20 hours a week for 2–3 weeks. $150–210/hour by geography and level; the 2-hour vetting exercise is paid $100.
$150 – $210 / HourWorldwideRegulatory counsel and compliance lawyers design AI evaluation scenarios, compliance memos, filings and rubrics across financial services, FDA, antitrust, privacy and energy regulation. US (APA, SEC, FDA, FTC) or EU/UK track. Remote hourly contract at $90–100/hour; 5+ years in regulatory practice.
$90 – $100 / HourWorldwidePatent attorneys, patent agents and IP counsel build AI evaluation scenarios on prosecution, licensing, FTO and IP litigation, with reference applications, office action responses and rubrics. US (35 U.S.C., MPEP) or EPC/PCT track. Remote hourly contract at $90–100/hour; 5+ years in IP practice.
$90 – $100 / HourWorldwideBuild HR evaluation tasks for AI on a US employment-law track, an international (UK/EU) track, or both: scenarios, reference policies and investigation reports, and rubrics. For HR leaders and CHROs with 5+ years at large companies. Remote hourly contract at $70–80/hour.
$70 – $80 / HourWorldwideEngineers with propulsion, initiation or effects test experience red-team frontier AI models: write benign, dual-use and adversarial prompts, judge the replies against a policy standard, and write reference answers with the reasoning. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's nuclear red-team panel: fuel-cycle engineers, safeguards inspectors, nuclear security, forensics and nonproliferation specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 21 hired this month.
$65 – $75 / TaskWorldwideBuild AI evaluation tasks set inside Fortune 500 insurance: commercial underwriting, claims adjudication, reserving, reinsurance and NAIC compliance scenarios, with reference guidelines and rubrics. Needs 5+ years at a major carrier or reinsurer. $50–60/hour, remote contract.
$50 – $60 / HourWorldwideMaterials science PhDs author original, executable research problems for a scientific-computing AI benchmark, with depth in both semiconductor materials and molecular modeling. Tasks ship only when frontier models fail them more often than not. 6 weeks, 20+ hours a week, Git and Docker workflow. $70/hour, 212 hired this month.
$70 / HourWorldwideRadiological emergency planners, field monitoring teams and consequence modellers red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideA paid expert conversation for engineers who build and run LLM agents in production: a short AI screening interview (no coding), then, if selected, a 30-minute live call on agent reliability, evaluation and internal adoption, paid $100–500 depending on depth of experience.
$100 – $500 / TaskWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.