Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide83 jobs
- United States45 jobs
- United Kingdom6 jobs
- Canada2 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
- Australia1 job
- Belgium2 jobs
- Ireland2 jobs
- Argentinano roles alongside your other filters
- Switzerland2 jobs
- Chileno roles alongside your other filters
- Colombiano roles alongside your other filters
- Costa Ricano roles alongside your other filters
- Germany2 jobs
- Spain2 jobs
- Panamano roles alongside your other filters
- Brazilno roles alongside your other filters
- Dominican Republicno roles alongside your other filters
- France2 jobs
- Italy2 jobs
- Luxembourg2 jobs
- Netherlands2 jobs
- New Zealandno roles alongside your other filters
- Peruno roles alongside your other filters
- Uruguayno roles alongside your other filters
- Albania2 jobs
- Austria2 jobs
- Bosnia & Herzegovina2 jobs
- Barbadosno roles alongside your other filters
- Bulgaria2 jobs
- Bahamasno roles alongside your other filters
- Czechia2 jobs
- Denmark2 jobs
- Ecuadorno roles alongside your other filters
- Estonia2 jobs
- Finland2 jobs
- Greece2 jobs
- Guatemalano roles alongside your other filters
- Hondurasno roles alongside your other filters
- Croatia2 jobs
- Hungary2 jobs
- Iceland2 jobs
- Jamaicano roles alongside your other filters
- South Koreano roles alongside your other filters
- Liechtenstein2 jobs
- Lithuania2 jobs
- Latvia2 jobs
- Monaco2 jobs
- Moldova2 jobs
- North Macedonia2 jobs
- Malta2 jobs
- Nicaraguano roles alongside your other filters
- Norway2 jobs
- Poland2 jobs
- Portugal2 jobs
- Romania2 jobs
- Serbia2 jobs
- Sweden2 jobs
- Slovenia2 jobs
- Slovakia2 jobs
- San Marino2 jobs
- El Salvadorno roles alongside your other filters
- Kosovo2 jobs
- Boliviano roles alongside your other filters
- Belizeno roles alongside your other filters
- Cubano roles alongside your other filters
- Indonesiano roles alongside your other filters
- Japanno roles alongside your other filters
- Paraguayno roles alongside your other filters
- Venezuelano roles alongside your other filters
- South Africano roles alongside your other filters
Language
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering29 jobs
- Business & Finance29 jobs
- Software & IT14 jobs
- Health & Medicine20 jobs
- Law, Policy & Security13 jobs
- General & Data Collection1 job
- Science & Math28 jobs
- AI Safety & Evaluation20 jobs
- Video, Image & Design3 jobs
- Data, AI & ML11 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Newest
131 open roles matching these filters · page 1 of 6
- $90 – $110 / HourWorldwide
Economists with a Master's or PhD and a year or more at a top research institution (World Bank, IMF, the Fed, a graduate school) work on a research project for a leading foundation-model AI lab. Remote contract at $120–150/hour, at least 10 hours a week for a minimum of four weeks.
$120 – $150 / HourWorldwideCondensed matter PhDs create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Nineteen narrow research areas, from bosonization to SYK. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideAMO physicists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, including levitated optomechanics, cavity QED and ultracold atoms in optical lattices. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideResearchers who have published on stochastic autocatalytic growth create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark. A single narrow area: chemical master equations, branching processes and reaction-network moments. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideHigh energy and nuclear theorists create, solve, review or audit research-level problems for CritPt, a public benchmark testing whether frontier AI models can do real physics research. Seven narrow areas, from AdS/BCFT to quasi-PDFs and dark photon searches. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideMathematical physicists create, solve, review or audit research-level problems for CritPt, a public AI physics benchmark, where the standard of proof sits closer to mathematics than physics. Four narrow areas, from hypergeometric identities to Fefferman-Graham geometry. Remote hourly contract at $80–110 per hour.
$80 – $110 / HourWorldwideHands-on structural, thermal, mechanical design and dynamics engineers review and write hard engineering problems about real hardware (loads, margins, heat transfer, vibration, tolerances) for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour.
$100 – $120 / HourWorldwideHands-on systems, integration, reliability and manufacturing test engineers review and write hard engineering problems about real hardware (V&V, qualification, FMEA, root-cause analysis) for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour, weekly via Stripe or Wise.
$100 – $120 / HourWorldwideEmbedded firmware, FPGA/RTL, flight software and hardware test automation engineers review and write hard problems about software that runs on real hardware (timing, race conditions, interfaces, bring-up) for an AI research initiative. 5+ years, 3 recent hands-on. Remote contract, $100–120/hour.
$100 – $120 / HourWorldwideHands-on RF, power electronics, analog/mixed-signal, PCB and signal integrity engineers review and write hard problems about real hardware designs, measurements and bring-up for an AI research initiative. 5+ years with 3 recent hands-on. Remote contract, $100–120/hour, weekly via Stripe or Wise.
$100 – $120 / HourWorldwideEvaluate frontier AI responses on grey-area and policy-sensitive topics (misinformation, political persuasion, self-harm, violence, cyber, biosecurity), apply safety rubrics and write structured feedback. 5+ years in trust and safety, journalism, policy, research or security. US, UK and most of Europe. $60–70/hour.
$60 – $70 / HourOpen to Albania, Austria and 38 more countriesSenior civil engineers build infrastructure evaluation tasks for AI: realistic design, permitting and construction scenarios, reference calculations and rubrics, on either a US (ASCE, ACI, AISC, AASHTO) or International (Eurocodes, ISO) standards track. 5+ years, PE or equivalent strongly preferred. Remote hourly contract at $70–80/hour.
$70 – $80 / HourWorldwideCertified explosives specialists, forensic analysts and licensee inspectors red-team frontier AI models: write benign, dual-use and adversarial prompts from casework, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideGenerate, structure and evaluate expert ALD and thin-film data for a frontier AI lab building semiconductor and physical-science models: solve hard problems, rate model reasoning, and turn recipes into model-ready data. US-based, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role through Cincinnatus LLC, embedded with a leading AI lab: senior board-certified physicians write golden solutions, instruction specs and clinical benchmarks for frontier models. Hybrid in the Bay Area, 40 hours a week for an initial 6 months, 4+ years post-residency and a US licence. $70–110/hour.
$70 – $110 / HourHybridOpen to United StatesInterpret dermatology cases from images and history, annotate lesions to a schema, grade AI assessments and write the criteria they are judged by. Non-clinical, shared expert pool, open worldwide. Board-certified or board-eligible, 3+ years post-residency, 15 hours a week minimum. A flat $270/hour.
$270 / HourWorldwideA short, well-paid sprint for very senior software engineers: help a leading foundation-model lab improve its models on hard SWE tasks. 10+ years at top US tech firms, about 20 hours a week for 2–3 weeks. $150–210/hour by geography and level; the 2-hour vetting exercise is paid $100.
$150 – $210 / HourWorldwideRegulatory counsel and compliance lawyers design AI evaluation scenarios, compliance memos, filings and rubrics across financial services, FDA, antitrust, privacy and energy regulation. US (APA, SEC, FDA, FTC) or EU/UK track. Remote hourly contract at $90–100/hour; 5+ years in regulatory practice.
$90 – $100 / HourWorldwidePatent attorneys, patent agents and IP counsel build AI evaluation scenarios on prosecution, licensing, FTO and IP litigation, with reference applications, office action responses and rubrics. US (35 U.S.C., MPEP) or EPC/PCT track. Remote hourly contract at $90–100/hour; 5+ years in IP practice.
$90 – $100 / HourWorldwideJoin a leading AI lab's research team as its accounting and audit specialist: review model outputs for misapplied standards, write instruction specs and golden solutions, and design benchmarks. Needs an active CPA, CA, CIA, CFE or CMA and 4+ years; bookkeeping-only roles do not count. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesBuild enterprise legal evaluation tasks for AI: realistic Fortune 500 scenarios, model-grade reference work and criterion-referenced rubrics. For in-house counsel or AmLaw 100 lawyers with an active bar admission, US-based, 20+ hours a week. $110–150/hour.
$110 – $150 / HourOpen to United StatesResidency-trained physicians in any specialty write grading criteria, evaluate multi-turn clinical dialogues and annotate clinical reasoning for AI systems. Shared pool with several workstreams. US only, 3+ years post-residency, 20 hours a week minimum, $150/hour.
$150 / HourOpen to United StatesTwo to three paid one-hour video interviews with Mercor about how enterprise security work is done and judged, feeding a benchmark for AI cyber defense agents. $125–175/hour, US-based, no prep and nothing to label or submit.
$125 – $175 / HourOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.