Build retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide83 jobs
- United States45 jobs
- United Kingdom6 jobs
- Canada2 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering29 jobs
- Business & Finance29 jobs
- Software & IT14 jobs
- Health & Medicine20 jobs
- Law, Policy & Security13 jobs
- General & Data Collection1 job
- Science & Math28 jobs
- AI Safety & Evaluation20 jobs
- Video, Image & Design3 jobs
- Data, AI & ML11 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Newest
131 open roles matching these filters · page 5 of 6
- $60 – $80 / HourOpen to United States
Senior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.
$90 – $110 / HourOpen to United StatesEvaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.
$70 – $90 / HourOpen to United StatesPaid pilot for US market access and pricing professionals: set and pressure-test gross-to-net and formulary-tier assumptions by drug class, judge whether coverage and rebating assumptions match real payer behaviour, and write rubrics for AI drug analysis. 5+ years, 10–20 hours. $175–200/hour.
$175 – $200 / HourWorldwideDesign adversarial prompts, find jailbreaks and policy failures, and document vulnerabilities in frontier AI models across cyber, biosecurity, fraud, misinformation and political content. Hourly remote contract at $70–84/hour for residents of Europe, the UK and the US; 5+ years' relevant experience required.
$70 – $84 / HourOpen to Albania, Austria and 38 more countriesFull-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesGrade AI-generated slides, spreadsheets and documents for real-world data science quality, flagging factual, visual and presentation errors in structured written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwidePaid pilot for US epidemiologists: size patient populations for drugs and indications (prevalence, incidence, diagnosed to treated to addressable), judge whether estimates are sound and well sourced, and write rubrics for AI drug analysis. 5+ years and a graduate degree, 10–20 hours. $150–175/hour.
$150 – $175 / HourWorldwideCoordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesAn expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.
$80 – $150 / TaskWorldwideGrade AI-generated slides, spreadsheets and documents for real-world software engineering quality, flagging factual, visual and presentation errors with written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideEvaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesUmbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's radiological safety red-team panel: RSOs, health physicists, source security, emergency response and nuclear medicine specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 8 hired this month.
$65 – $75 / TaskWorldwideEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesRF, microwave, antenna and electromagnetics engineers solve and critique hard technical problems for a short-term expert evaluation project: analysing systems and trade-offs, checking calculations and assumptions, and writing rigorous explanations. Hands-on industry experience and an EE-family degree. Remote hourly contract at $80–95/hour.
$80 – $95 / HourWorldwideQA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.
$70 – $110 / HourWorldwideWrite or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.
$94 – $119 / HourWorldwideWrite or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.
$61 – $77 / HourWorldwideAuthor executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.
$70 / HourWorldwideSenior materials scientists, and electrical or mechanical engineers, author realistic tasks with a prompt, a data room and a grading method, run them against an AI model and tighten them until the model can no longer reason through cleanly. Daily onboarding and office hours. Remote hourly contract at $60–90/hour.
$60 – $90 / HourWorldwideFull-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesGrade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideFull-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.
$110 – $150 / HourHybridOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.