Build HR evaluation tasks for AI on a US employment-law track, an international (UK/EU) track, or both: scenarios, reference policies and investigation reports, and rubrics. For HR leaders and CHROs with 5+ years at large companies. Remote hourly contract at $70–80/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide230 jobs
- United States63 jobs
- United Kingdom7 jobs
- Canada5 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language: English
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering55 jobs
- Business & Finance45 jobs
- Software & IT38 jobs
- Health & Medicine60 jobs
- Law, Policy & Security51 jobs
- General & Data Collection2 jobs
- Science & Math46 jobs
- AI Safety & Evaluation22 jobs
- Video, Image & Design6 jobs
- Data, AI & ML17 jobs
- Writing & Education4 jobs
- Other fields5 jobs
Newest
294 open roles matching these filters · page 5 of 13
- $70 – $80 / HourWorldwide
Engineers with propulsion, initiation or effects test experience red-team frontier AI models: write benign, dual-use and adversarial prompts, judge the replies against a policy standard, and write reference answers with the reasoning. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's nuclear red-team panel: fuel-cycle engineers, safeguards inspectors, nuclear security, forensics and nonproliferation specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 21 hired this month.
$65 – $75 / TaskWorldwideBigLaw immigration attorneys review immigration documents and legal scenarios, draft and grade memos and briefs, and correct AI-generated legal output. Ten openings, contractor, fully remote, $140–400/hour. An active licence and complex US immigration casework at a large firm are the stated profile.
$140 – $400 / HourWorldwidePractising US hospitalists and inpatient internists review H&Ps, progress notes and discharge summaries, and judge AI-written inpatient documentation against what a hospitalist would chart. Needs C1 or better in one of 25 listed languages. 2+ years post-residency, 10 hours a week, a flat $170/hour.
$170 / HourOpen to United StatesBuild and evaluate training data for a frontier lab's materials science models: DFT, AIMD, classical MD, surface and adsorption modeling, reaction energetics. For US-based computational PhDs fluent in VASP, Quantum ESPRESSO, CP2K, LAMMPS or ASE. Long-term, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role (via Cincinnatus LLC) for counsel and senior associates with 8 to 15 years' practice, embedded with a leading AI lab in the Bay Area to review legal model outputs, write golden solutions and build benchmarks. Hybrid, 6-month initial term, $85–120/hour. US bar admission required.
$85 – $120 / HourHybridOpen to United StatesWrite point-in-time forecasts on specific swing-state Senate, governor and statewide races, and grade AI political analyses against your own. For state pollsters, campaign analysts and political scientists with live-race experience. US or Canada residents, $150–250/hour.
$150 – $250 / HourOpen to Canada, United StatesBuild AI evaluation tasks set inside Fortune 500 insurance: commercial underwriting, claims adjudication, reserving, reinsurance and NAIC compliance scenarios, with reference guidelines and rubrics. Needs 5+ years at a major carrier or reinsurer. $50–60/hour, remote contract.
$50 – $60 / HourWorldwideMaterials science PhDs author original, executable research problems for a scientific-computing AI benchmark, with depth in both semiconductor materials and molecular modeling. Tasks ship only when frontier models fail them more often than not. 6 weeks, 20+ hours a week, Git and Docker workflow. $70/hour, 212 hired this month.
$70 / HourWorldwideThe top tier of Mercor's embedded legal expert role: a full-time W-2 job (via Cincinnatus LLC) with a leading AI lab in the Bay Area, reviewing legal model outputs, writing golden solutions and building benchmarks. For partners and general counsel. Hybrid, 6-month initial term, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesBuild realistic production-management tasks for AI benchmarks (budget reconciliation, schedule changes, vendor and crew coordination) from authentic budgets, call sheets and contracts, then write 35+ criterion rubrics to grade the answers. Five years of credited production experience; paid per accepted task at $45–85/hour. Fifty openings.
$45 – $85 / HourWorldwidePart-time legal AI work for in-house technology lawyers: run simulated negotiations and redlines on MSAs, NDAs and DPAs, grade AI responses and write the evaluation criteria. Requires three years in-house on tech transactions; no bar admission is listed. 50 openings at $85–105/hour.
$85 – $105 / HourWorldwideThe reviewer seat on micro1's physics work: critique derivations and arguments written by researchers or AI, find the errors and unjustified steps, and write precise feedback, cross-checking in SymPy and Python. Postdocs and junior faculty, US, Canada and UK focused. 30 openings, contractor, $80–150/hour.
$80 – $150 / HourWorldwideRadiological emergency planners, field monitoring teams and consequence modellers red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideA hybrid, Bay Area-based W-2 role embedded with a leading AI lab: senior software engineers vet model outputs, write instruction specs and golden solutions, and build engineering benchmarks. 4+ years, senior-level progression and a CS or engineering degree. 40 hours a week for an initial 6 months, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesCreate, solve and review hard full-stack engineering tasks, debug across browser, API and database layers, judge AI-generated solutions, and port whole builds from one language to another. Global and fully remote, about 15 flexible hours a week, paid per task on a $50–100/hour band, 300 openings.
$50 – $100 / HourWorldwideProduce real design work in Figma (layouts, wireframes, prototypes, design systems) and critique design iterations in writing, as reference material for training AI on visual and UX quality. US-based designers with 5+ years and a strong portfolio. $30–80/hour, ten openings, contractor.
$30 – $80 / HourOpen to United StatesA paid expert conversation for engineers who build and run LLM agents in production: a short AI screening interview (no coding), then, if selected, a 30-minute live call on agent reliability, evaluation and internal adoption, paid $100–500 depending on depth of experience.
$100 – $500 / TaskWorldwidePaid pilot for a biotech and pharma research team: create and critique rubrics that assess commercial drugs and development programs, judge investment-style theses, and assess 5–10 companies end to end. For specialist-fund biotech analysts with 5+ years. About 10–20 hours over 1–2 weeks, $120–200/hour.
$120 – $200 / HourWorldwideExplain physics to an AI and grade what it writes back: write clear explanations of hard concepts, build physics training material, and review model output for conceptual accuracy. A physics PhD is the preferred bar. 100 openings, contractor, $70–90/hour, fully remote.
$70 – $90 / HourWorldwideRead US sales and use tax statutes subsection by subsection, write the rubric an AI's formal translation of the law must satisfy, then grade that translation pass/fail and write test scenarios for what it missed. CPA, CA or US tax attorney background required. Ten openings, contract, $20–30/hour.
$20 – $30 / HourWorldwideResearch-level AI training work on magnetic order: magnetic structure factors, propagation vectors, magnetic space groups in BNS notation checked against MAGNDATA, AFM/FM classification, neutron scattering and MOKE signatures. Ten openings, contractor, $80–160/hour, remote.
$80 – $160 / HourWorldwideLend lab chemistry expertise (reactions, synthesis, separations, analytical methods) to AI training data, judging experimental reasoning and explaining it plainly. Contractor, remote, 100 openings, $70–90/hour. A PhD is preferred and hands-on wet-lab experience is expected.
$70 – $90 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.