Pediatric acute care RNs evaluate AI outputs built from nursing flowsheet documentation, annotate pediatric assessment data and help write clinical benchmarks. US licence outside California, inpatient bedside work within the past 10 years. $55–65/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
Language: English
Field
- Languages & Linguistics78 jobs
- Audio & Voice74 jobs
- Engineering107 jobs
- Business & Finance104 jobs
- Software & IT76 jobs
- Health & Medicine74 jobs
- Law, Policy & Security70 jobs
- General & Data Collection53 jobs
- Science & Math63 jobs
- AI Safety & Evaluation63 jobs
- Video, Image & Design37 jobs
- Data, AI & ML34 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
631 open roles matching these filters · page 19 of 27
- $55 – $65 / HourOpen to United States
Belgium-based PhD chemists and biologists who write Belgian Dutch: author specialised science prompts and grade how AI models handle accuracy and dual-use safety. Belgium residence required. Part-time remote at $61–65/hour, $13 above the Belgian Dutch generalist role; 9 hires this month.
$61 – $65 / HourOpen to BelgiumNative Danish speakers write sensitive-topic prompts, classify prompts and conversations, and flag adversarial phrasing to make AI models safer in Danish. Business English and a bachelor's (in progress counts). Part-time remote at $48–52/hour; Denmark or Western Europe preferred, not required.
$48 – $52 / HourWorldwideAustralia-based native English speakers with an authentic Australian accent record one session of about four hours, which a Mercor client uses to clone the voice for its internal customer-experience AI agent. $50–100/hour in USD, no acting experience required. You are licensing your voice, so check the terms before recording.
$50 – $100 / HourOpen to AustraliaRead a specific non-G7 central bank's statements, minutes and speeches in the source language, score them dovish to hawkish from a fixed evidence cutoff, and grade AI macro analyses. For former central bank economists or senior local rates and FX strategists, typically 8+ years. $150–250/hour, remote.
$150 – $250 / HourWorldwideAudit French speech data for Amazon's Sonic project: check annotators' transcripts against audio with fixed error codes and pass/fail calls, and correct word-level timestamp alignment. Native French as spoken in France is a hard requirement; rationales are in English. Joint top rate of the Sonic set at $41.50/hour, remote contract.
$41.50 / HourWorldwideRadioactive source security and vulnerability assessment specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from Category 1 and 2 source security work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideUS-based cloud and DevOps engineers design and grade AI training tasks on Kubernetes failure diagnosis, AWS service integration, Terraform or CDK design and CI/CD, and write the rubrics behind them. 4+ years at a top-tier organisation. Full-time W-2 through Cincinnatus LLC at a leading AI lab, $75–110/hour.
$75 – $110 / HourOpen to United StatesInpatient registered nurses with broad clinical exposure support ongoing clinical AI product work: annotation, clinical review, model evaluation, user-feedback investigation and guideline development, alongside engineers and clinicians. US only, 10 hours a week minimum, $55–65/hour.
$55 – $65 / HourOpen to United StatesExperienced administrators design benchmark tasks for AI agents: realistic calendar, travel, expense and records scenarios with mock emails and receipts, the expert solution path, and a rubric of 35+ criteria. Five years of admin work in a mid-to-large organisation. Paid per accepted task, $20–40/hour band.
$20 – $40 / HourWorldwideDesign finance Excel tasks from your own work (three-statement models, LBO and DCF valuation, forecasting), write model solutions, and evaluate AI attempts for a leading tech company's GenAI team. US only, W-2 through Cincinnatus LLC, 40 hours a week to end of September then 20. $70–100/hour.
$70 – $100 / HourOpen to United StatesTurn your favourite open repositories into reinforcement learning environments for frontier coding models: subtle bugs, non-trivial features or performance problems, each with a reproducible setup, robust verifiers and a reference solution. About 15 hours a week, paid per task on a $50–100/hour band.
$50 – $100 / HourWorldwideSouth Africa-based native English speakers with an authentic South African accent record one session of about four hours, which a Mercor client uses to clone the voice for its internal customer-experience AI agent. $50–100/hour in USD, a strong rate locally. No acting experience required. You are licensing your voice.
$50 – $100 / HourOpen to South AfricaAudit end-to-end AI-assisted coding sessions (traces from tools like Cursor, Copilot or Claude Code) used to train and evaluate a frontier lab's models, judging correctness, workflow and reasoning with rubric-based feedback. 3+ years of software development plus hands-on agentic coding. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesAudit Italian speech data for Amazon's Sonic collections: judge annotators' transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Italian as spoken in Italy is a hard requirement; rationales are written in English. Remote hourly contract at $39.50/hour.
$39.50 / HourWorldwideFull-time Bay Area hybrid role embedded with an AI lab: review model reasoning on materials problems, write golden solutions and specs, and build benchmarks. For materials PhDs (or master's with exceptional industrial depth) with 4+ years of R&D. W-2 via Cincinnatus, $70–110/hour.
$70 – $110 / HourHybridOpen to United StatesML systems engineers write and evaluate training tasks for a frontier lab across GPU kernels, performance profiling, distributed debugging and LLM inference serving, plus the rubrics that grade them. 2+ years of hands-on ML infrastructure work. Canada, UK or US; 40 hours a week. $90–120/hour.
$90 – $120 / HourOpen to Canada, United Kingdom and 1 more countryAudit German speech data for Amazon's Sonic collections: verify annotators' transcripts against audio using fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native German as spoken in Germany is a hard requirement; rationales are in English. Joint top rate of the Sonic set at $41.50/hour, remote.
$41.50 / HourWorldwideTranscribe Spanish and English audio with precision, add metadata and contextual notes, and proofread for an AI training dataset. Native or near-native Spanish and professional transcription experience preferred. 15 openings, contractor, $20–36/hour.
$20 – $36 / HourWorldwideTest AI chat models in Marathi and English for safety failures (jailbreaks, bias, harmful answers) and judge whether their Marathi is accurate and appropriate rather than Hindi in disguise. Evaluation judgment is the core ask. Remote hourly contract, $16–22/hour, weekly pay.
$16 – $22 / HourWorldwideRemote hourly contract for colorists and video editors who already work in DaVinci Resolve on their own Mac. $30–40/hour, paid weekly via Stripe or Wise. The free version of Resolve is widely used; you also need a display above 2.5 megapixels. Tasks are not described.
$30 – $40 / HourWorldwideHelp an AI lab evaluate a performance transfer model: define what a faithfully captured performance looks like, curate easy and hard benchmark examples, and shape the reviewer pool. For character animators with 4+ years and 2+ feature or AAA credits. US, sessions in LA or NYC, $60–90/hour.
$60 – $90 / HourHybridOpen to United StatesJoin Mercor's bench of management consultants for future AI evaluation projects: writing grading criteria for consulting deliverables and scoring AI and human work. No project is open yet. For consultants with 1+ year at MBB or equivalent. Remote, $100–150/hour.
$100 – $150 / HourWorldwideOperational health physicists from DOE sites, national labs, reactors and decommissioning projects red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.