Evaluate frontier AI responses on grey-area and policy-sensitive topics (misinformation, political persuasion, self-harm, violence, cyber, biosecurity), apply safety rubrics and write structured feedback. 5+ years in trust and safety, journalism, policy, research or security. US, UK and most of Europe. $60–70/hour.
Remote AI training and data labeling jobs
Filter jobs
Location: United States
- Worldwide83 jobs
- United States45 jobs, applied. Activate to remove
- United Kingdom6 jobs
- Canada2 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering6 jobs
- Business & Finance13 jobs
- Software & IT7 jobs
- Health & Medicine8 jobs
- Law, Policy & Security6 jobs
- General & Data Collection1 job
- Science & Math5 jobs
- AI Safety & Evaluation2 jobs
- Video, Image & Design2 jobs
- Data, AI & ML3 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Newest
45 open roles matching these filters · page 1 of 2
- $60 – $70 / HourOpen to Albania, Austria and 38 more countries
Generate, structure and evaluate expert ALD and thin-film data for a frontier AI lab building semiconductor and physical-science models: solve hard problems, rate model reasoning, and turn recipes into model-ready data. US-based, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role through Cincinnatus LLC, embedded with a leading AI lab: senior board-certified physicians write golden solutions, instruction specs and clinical benchmarks for frontier models. Hybrid in the Bay Area, 40 hours a week for an initial 6 months, 4+ years post-residency and a US licence. $70–110/hour.
$70 – $110 / HourHybridOpen to United StatesJoin a leading AI lab's research team as its accounting and audit specialist: review model outputs for misapplied standards, write instruction specs and golden solutions, and design benchmarks. Needs an active CPA, CA, CIA, CFE or CMA and 4+ years; bookkeeping-only roles do not count. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesBuild enterprise legal evaluation tasks for AI: realistic Fortune 500 scenarios, model-grade reference work and criterion-referenced rubrics. For in-house counsel or AmLaw 100 lawyers with an active bar admission, US-based, 20+ hours a week. $110–150/hour.
$110 – $150 / HourOpen to United StatesResidency-trained physicians in any specialty write grading criteria, evaluate multi-turn clinical dialogues and annotate clinical reasoning for AI systems. Shared pool with several workstreams. US only, 3+ years post-residency, 20 hours a week minimum, $150/hour.
$150 / HourOpen to United StatesTwo to three paid one-hour video interviews with Mercor about how enterprise security work is done and judged, feeding a benchmark for AI cyber defense agents. $125–175/hour, US-based, no prep and nothing to label or submit.
$125 – $175 / HourOpen to United StatesPractising US hospitalists and inpatient internists review H&Ps, progress notes and discharge summaries, and judge AI-written inpatient documentation against what a hospitalist would chart. Needs C1 or better in one of 25 listed languages. 2+ years post-residency, 10 hours a week, a flat $170/hour.
$170 / HourOpen to United StatesBuild and evaluate training data for a frontier lab's materials science models: DFT, AIMD, classical MD, surface and adsorption modeling, reaction energetics. For US-based computational PhDs fluent in VASP, Quantum ESPRESSO, CP2K, LAMMPS or ASE. Long-term, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role (via Cincinnatus LLC) for counsel and senior associates with 8 to 15 years' practice, embedded with a leading AI lab in the Bay Area to review legal model outputs, write golden solutions and build benchmarks. Hybrid, 6-month initial term, $85–120/hour. US bar admission required.
$85 – $120 / HourHybridOpen to United StatesWrite point-in-time forecasts on specific swing-state Senate, governor and statewide races, and grade AI political analyses against your own. For state pollsters, campaign analysts and political scientists with live-race experience. US or Canada residents, $150–250/hour.
$150 – $250 / HourOpen to Canada, United StatesThe top tier of Mercor's embedded legal expert role: a full-time W-2 job (via Cincinnatus LLC) with a leading AI lab in the Bay Area, reviewing legal model outputs, writing golden solutions and building benchmarks. For partners and general counsel. Hybrid, 6-month initial term, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesA hybrid, Bay Area-based W-2 role embedded with a leading AI lab: senior software engineers vet model outputs, write instruction specs and golden solutions, and build engineering benchmarks. 4+ years, senior-level progression and a CS or engineering degree. 40 hours a week for an initial 6 months, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesFull-time Bay Area hybrid role with an AI lab's research team: review model reasoning on life sciences tasks, write golden solutions and instruction specs, and design benchmarks. For life sciences PhDs with 4+ years of substantive research experience. W-2 via Cincinnatus, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesAuthor AI evaluation tasks from real fire and life safety review work: egress plan checks, sprinkler and alarm review, hazmat control areas, firestop photo verification. For US fire marshals, fire protection engineers and NICET III+ designer-reviewers with 3+ years in the seat. Remote hourly contract at $45–60/hour.
$45 – $60 / HourOpen to United StatesAudit repository-level software engineering benchmark tasks for a frontier AI lab: reference patches, test harnesses, Docker isolation, and signs of answer leakage or reward hacking. For US engineers with 3+ years and real open-source contributor or maintainer history. $70–90/hour.
$70 – $90 / HourOpen to United StatesA full-time W-2 placement at a leading AI lab through Cincinnatus LLC: US-based mechanical engineers vet model outputs, write instruction specs and reference solutions, and build benchmarks. 5+ years in industry and a mechanical engineering degree required. 40 hours a week for an initial 2–3 months, $60–90/hour.
$60 – $90 / HourOpen to United StatesJudge AI-enhanced and upscaled video and stills at pixel level for a leading AI lab's GenAI team: artifacts, noise, aliasing, banding, grain, sharpening. For VFX and rendering supervisors, colorists, DPs, lighting artists and high-end photographers with 5+ years. US, 20 hours a week, $60–90/hour.
$60 – $90 / HourOpen to United StatesAct as ground truth for an AI lab teaching models real enterprise sales work: audit workflows, build golden reference trajectories in a mock sales stack, and refine rubrics. For sellers with around 10 years in enterprise sales. US only, $60–90/hour, placed via Cincinnatus.
$60 – $90 / HourOpen to United StatesPractising US primary care, family medicine or general internal medicine physicians review outpatient notes and judge AI-written documentation against what a PCP would actually chart. Needs C1 or better in one of 25 listed languages. 2+ years post-residency, 10 hours a week, $170–190/hour.
$170 – $190 / HourOpen to United StatesExperimental scientists create and review training data for a frontier lab's materials science models: inorganic synthesis, superconductors, semiconductors and advanced packaging, characterization (XRD, SEM, TEM) and fabrication. PhD, MS or equivalent hands-on experience. US-based, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role through Cincinnatus LLC, embedded with a leading AI lab: senior drug development scientists write golden solutions, instruction specs and benchmarks for pharma R&D reasoning. Hybrid in the Bay Area, 40 hours a week for an initial 6 months. PhD, PharmD or MD with 4+ years in industry R&D. $75–115/hour.
$75 – $115 / HourHybridOpen to United StatesA full-time W-2 role (via Cincinnatus LLC) embedded with a leading AI lab in the Bay Area: review legal model outputs, write instruction specs and golden solutions, and build legal benchmarks. Hybrid, on-site several days a week, 6-month initial term. $60–100/hour; JD, 5+ years' practice, US bar.
$60 – $100 / HourHybridOpen to United StatesFull-time, on-site-hybrid role in the Bay Area: review AI output on business and sales operations tasks, write instruction specs and golden solutions, and build benchmarks with an AI lab's research team. For senior ops leaders with 4+ years. W-2 via Cincinnatus, $60–100/hour.
$60 – $100 / HourHybridOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.