Build HR evaluation tasks for AI on a US employment-law track, an international (UK/EU) track, or both: scenarios, reference policies and investigation reports, and rubrics. For HR leaders and CHROs with 5+ years at large companies. Remote hourly contract at $70–80/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide83 jobs
- United States45 jobs
- United Kingdom6 jobs
- Canada2 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering29 jobs
- Business & Finance29 jobs
- Software & IT14 jobs
- Health & Medicine20 jobs
- Law, Policy & Security13 jobs
- General & Data Collection1 job
- Science & Math28 jobs
- AI Safety & Evaluation20 jobs
- Video, Image & Design3 jobs
- Data, AI & ML11 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Newest
131 open roles matching these filters · page 2 of 6
- $70 – $80 / HourWorldwide
Engineers with propulsion, initiation or effects test experience red-team frontier AI models: write benign, dual-use and adversarial prompts, judge the replies against a policy standard, and write reference answers with the reasoning. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's nuclear red-team panel: fuel-cycle engineers, safeguards inspectors, nuclear security, forensics and nonproliferation specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 21 hired this month.
$65 – $75 / TaskWorldwidePractising US hospitalists and inpatient internists review H&Ps, progress notes and discharge summaries, and judge AI-written inpatient documentation against what a hospitalist would chart. Needs C1 or better in one of 25 listed languages. 2+ years post-residency, 10 hours a week, a flat $170/hour.
$170 / HourOpen to United StatesBuild and evaluate training data for a frontier lab's materials science models: DFT, AIMD, classical MD, surface and adsorption modeling, reaction energetics. For US-based computational PhDs fluent in VASP, Quantum ESPRESSO, CP2K, LAMMPS or ASE. Long-term, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role (via Cincinnatus LLC) for counsel and senior associates with 8 to 15 years' practice, embedded with a leading AI lab in the Bay Area to review legal model outputs, write golden solutions and build benchmarks. Hybrid, 6-month initial term, $85–120/hour. US bar admission required.
$85 – $120 / HourHybridOpen to United StatesWrite point-in-time forecasts on specific swing-state Senate, governor and statewide races, and grade AI political analyses against your own. For state pollsters, campaign analysts and political scientists with live-race experience. US or Canada residents, $150–250/hour.
$150 – $250 / HourOpen to Canada, United StatesBuild AI evaluation tasks set inside Fortune 500 insurance: commercial underwriting, claims adjudication, reserving, reinsurance and NAIC compliance scenarios, with reference guidelines and rubrics. Needs 5+ years at a major carrier or reinsurer. $50–60/hour, remote contract.
$50 – $60 / HourWorldwideMaterials science PhDs author original, executable research problems for a scientific-computing AI benchmark, with depth in both semiconductor materials and molecular modeling. Tasks ship only when frontier models fail them more often than not. 6 weeks, 20+ hours a week, Git and Docker workflow. $70/hour, 212 hired this month.
$70 / HourWorldwideThe top tier of Mercor's embedded legal expert role: a full-time W-2 job (via Cincinnatus LLC) with a leading AI lab in the Bay Area, reviewing legal model outputs, writing golden solutions and building benchmarks. For partners and general counsel. Hybrid, 6-month initial term, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesRadiological emergency planners, field monitoring teams and consequence modellers red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideA hybrid, Bay Area-based W-2 role embedded with a leading AI lab: senior software engineers vet model outputs, write instruction specs and golden solutions, and build engineering benchmarks. 4+ years, senior-level progression and a CS or engineering degree. 40 hours a week for an initial 6 months, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesA paid expert conversation for engineers who build and run LLM agents in production: a short AI screening interview (no coding), then, if selected, a 30-minute live call on agent reliability, evaluation and internal adoption, paid $100–500 depending on depth of experience.
$100 – $500 / TaskWorldwidePaid pilot for a biotech and pharma research team: create and critique rubrics that assess commercial drugs and development programs, judge investment-style theses, and assess 5–10 companies end to end. For specialist-fund biotech analysts with 5+ years. About 10–20 hours over 1–2 weeks, $120–200/hour.
$120 – $200 / HourWorldwideRed-team frontier AI models on chemical safety: write benign, dual-use and adversarial prompts from exposure and process-hazard work, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideFull-time Bay Area hybrid role with an AI lab's research team: review model reasoning on life sciences tasks, write golden solutions and instruction specs, and design benchmarks. For life sciences PhDs with 4+ years of substantive research experience. W-2 via Cincinnatus, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesUmbrella listing for Mercor's chemistry red-team panel: synthetic, analytical, forensic, defence and process safety chemists write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 5 hired this month.
$65 – $75 / TaskWorldwideRed-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from trace analysis and method validation, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideAuthor AI evaluation tasks from real fire and life safety review work: egress plan checks, sprinkler and alarm review, hazmat control areas, firestop photo verification. For US fire marshals, fire protection engineers and NICET III+ designer-reviewers with 3+ years in the seat. Remote hourly contract at $45–60/hour.
$45 – $60 / HourOpen to United StatesAuthor AI evaluation tasks in budgeting, forecasting, variance analysis and capital allocation: realistic FP&A scenarios, reference models and decks, and rubrics that reward real planning judgment over template work. For FP&A directors and CFOs with 5+ years. $80–90/hour, remote contract.
$80 – $90 / HourWorldwideAudit repository-level software engineering benchmark tasks for a frontier AI lab: reference patches, test harnesses, Docker isolation, and signs of answer leakage or reward hacking. For US engineers with 3+ years and real open-source contributor or maintainer history. $70–90/hour.
$70 – $90 / HourOpen to United StatesPaid pilot for board-certified, practising physicians with deep therapeutic-area expertise: interpret trial endpoints, judge whether results would change prescribing and real-world uptake, and write rubrics that evaluate AI analysis of drugs. 10–20 hours over 1–2 weeks. $150–230/hour.
$150 – $230 / HourWorldwideGrade AI-generated slides, spreadsheets and documents to consulting standard, flagging factual, visual and presentation errors with structured written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideRed-team frontier AI models on chemical misuse: write benign, dual-use and adversarial prompts from route design and scale-up experience, judge responses against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.