Experienced administrators design benchmark tasks for AI agents: realistic calendar, travel, expense and records scenarios with mock emails and receipts, the expert solution path, and a rubric of 35+ criteria. Five years of admin work in a mid-to-large organisation. Paid per accepted task, $20–40/hour band.
Remote AI training and data labeling jobs
Filter jobs
Location: Worldwide
Language: English
Field
- Languages & Linguistics76 jobs
- Audio & Voice65 jobs
- Engineering91 jobs
- Business & Finance72 jobs
- Software & IT59 jobs
- Health & Medicine58 jobs
- Law, Policy & Security51 jobs
- General & Data Collection43 jobs
- Science & Math56 jobs
- AI Safety & Evaluation59 jobs
- Video, Image & Design31 jobs
- Data, AI & ML27 jobs
- Writing & Education11 jobs
- Other fields6 jobs
Newest
508 open roles matching these filters · page 16 of 22
- $20 – $40 / HourWorldwide
Turn your favourite open repositories into reinforcement learning environments for frontier coding models: subtle bugs, non-trivial features or performance problems, each with a reproducible setup, robust verifiers and a reference solution. About 15 hours a week, paid per task on a $50–100/hour band.
$50 – $100 / HourWorldwideAudit Italian speech data for Amazon's Sonic collections: judge annotators' transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Italian as spoken in Italy is a hard requirement; rationales are written in English. Remote hourly contract at $39.50/hour.
$39.50 / HourWorldwideAudit German speech data for Amazon's Sonic collections: verify annotators' transcripts against audio using fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native German as spoken in Germany is a hard requirement; rationales are in English. Joint top rate of the Sonic set at $41.50/hour, remote.
$41.50 / HourWorldwideTranscribe Spanish and English audio with precision, add metadata and contextual notes, and proofread for an AI training dataset. Native or near-native Spanish and professional transcription experience preferred. 15 openings, contractor, $20–36/hour.
$20 – $36 / HourWorldwideTest AI chat models in Marathi and English for safety failures (jailbreaks, bias, harmful answers) and judge whether their Marathi is accurate and appropriate rather than Hindi in disguise. Evaluation judgment is the core ask. Remote hourly contract, $16–22/hour, weekly pay.
$16 – $22 / HourWorldwideRemote hourly contract for colorists and video editors who already work in DaVinci Resolve on their own Mac. $30–40/hour, paid weekly via Stripe or Wise. The free version of Resolve is widely used; you also need a display above 2.5 megapixels. Tasks are not described.
$30 – $40 / HourWorldwideJoin Mercor's bench of management consultants for future AI evaluation projects: writing grading criteria for consulting deliverables and scoring AI and human work. No project is open yet. For consultants with 1+ year at MBB or equivalent. Remote, $100–150/hour.
$100 – $150 / HourWorldwideOperational health physicists from DOE sites, national labs, reactors and decommissioning projects red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideRed-team frontier AI models from the chemical defence side: write benign, dual-use and adversarial prompts drawn from countermeasures, protection and detection work, judge how models respond against a policy standard, and write the reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideRemote hourly contract for game developers, technical artists and virtual production or archviz specialists who already run Unreal Engine on their own Windows PC. $30–40/hour, paid weekly via Stripe or Wise. The engine is free to download; a display above 2.5 megapixels is required.
$30 – $40 / HourWorldwideKorean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.
$63 – $67 / HourWorldwideNative Korean speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models stay safe in Korean. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; South Korea or East Asia preferred, not required; 7 hires this month.
$48 – $52 / HourWorldwideSenior M&A, securities and governance lawyers design corporate law scenarios, reference memos and rubrics that test AI on deal work. US (Delaware, SEC) or international (UK Companies Act, EU) track. Remote hourly contract at $90–100/hour; 5+ years at a major firm, bank or large company.
$90 – $100 / HourWorldwideHindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.
$23 – $27 / HourWorldwideAudit Spanish (Spain) speech data for Amazon's Sonic project: check annotators' transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Spanish as spoken in Spain is a hard requirement; rationales are in English. Remote hourly contract at $39.50/hour.
$39.50 / HourWorldwidePortuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.
$50 – $54 / HourWorldwideNamed RSOs and radiation protection managers on broad-scope, hospital, university or industrial licences red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideBoard-certified radiologists label findings, write reference reports, grade AI-generated reads and author rubrics for medical imaging AI. Non-clinical, remote worldwide, hourly contract at $200–400/hour with a 15-hour weekly minimum. Weekly pay via Stripe or Wise.
$200 – $400 / HourWorldwideAI safety testing in Urdu and English: provoke jailbreaks, bias and harmful output from chat models and judge whether their Urdu answers are accurate and appropriate, in Nastaliq script or Roman Urdu. Evaluation judgment is the core ask. Remote hourly contract at $16–22/hour, weekly pay.
$16 – $22 / HourWorldwideEvaluate generative music AI for a leading AI lab: compare AI-made songs head to head on musicality, prompt adherence, vocals and mix, label genre and structure, and check lyrics and vocals. For Thai-speaking producers or engineers with 2+ years' experience. Remote, flexible hours, up to 6 months, $18/hour.
$18 / HourWorldwideRate AI-generated music for a leading AI lab: head-to-head song comparisons on musicality, prompt adherence, vocals and mix, genre and structure labelling, and lyric and vocal checks. For Russian-speaking producers and mix engineers with 2+ years' experience. Remote, flexible hours, up to 6 months, $35–49/hour.
$35 – $49 / HourWorldwideRemote hourly contract for video editors and content producers who already cut in Adobe Premiere on their own Windows PC. $30–40/hour, paid weekly via Stripe or Wise. Your own Adobe subscription and a display above 2.5 megapixels are required; the tasks are not described.
$30 – $40 / HourWorldwidePaid pilot for US market access and pricing professionals: set and pressure-test gross-to-net and formulary-tier assumptions by drug class, judge whether coverage and rebating assumptions match real payer behaviour, and write rubrics for AI drug analysis. 5+ years, 10–20 hours. $175–200/hour.
$175 – $200 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.