Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.
  • Radioactive source security and vulnerability assessment specialists red-team frontier AI models: write benign, dual-use and adversarial prompts from Category 1 and 2 source security work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • US-based cloud and DevOps engineers design and grade AI training tasks on Kubernetes failure diagnosis, AWS service integration, Terraform or CDK design and CI/CD, and write the rubrics behind them. 4+ years at a top-tier organisation. Full-time W-2 through Cincinnatus LLC at a leading AI lab, $75–110/hour.

    $75 – $110 / HourOpen to United States
  • Inpatient registered nurses with broad clinical exposure support ongoing clinical AI product work: annotation, clinical review, model evaluation, user-feedback investigation and guideline development, alongside engineers and clinicians. US only, 10 hours a week minimum, $55–65/hour.

    $55 – $65 / HourOpen to United States
  • Design finance Excel tasks from your own work (three-statement models, LBO and DCF valuation, forecasting), write model solutions, and evaluate AI attempts for a leading tech company's GenAI team. US only, W-2 through Cincinnatus LLC, 40 hours a week to end of September then 20. $70–100/hour.

    $70 – $100 / HourOpen to United States
  • South Africa-based native English speakers with an authentic South African accent record one session of about four hours, which a Mercor client uses to clone the voice for its internal customer-experience AI agent. $50–100/hour in USD, a strong rate locally. No acting experience required. You are licensing your voice.

    $50 – $100 / HourOpen to South Africa
  • Audit end-to-end AI-assisted coding sessions (traces from tools like Cursor, Copilot or Claude Code) used to train and evaluate a frontier lab's models, judging correctness, workflow and reasoning with rubric-based feedback. 3+ years of software development plus hands-on agentic coding. US-only, $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Audit Italian speech data for Amazon's Sonic collections: judge annotators' transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Italian as spoken in Italy is a hard requirement; rationales are written in English. Remote hourly contract at $39.50/hour.

    $39.50 / HourWorldwide
  • Full-time Bay Area hybrid role embedded with an AI lab: review model reasoning on materials problems, write golden solutions and specs, and build benchmarks. For materials PhDs (or master's with exceptional industrial depth) with 4+ years of R&D. W-2 via Cincinnatus, $70–110/hour.

    $70 – $110 / HourHybridOpen to United States
  • ML systems engineers write and evaluate training tasks for a frontier lab across GPU kernels, performance profiling, distributed debugging and LLM inference serving, plus the rubrics that grade them. 2+ years of hands-on ML infrastructure work. Canada, UK or US; 40 hours a week. $90–120/hour.

    $90 – $120 / HourOpen to Canada, United Kingdom and 1 more country
  • Audit German speech data for Amazon's Sonic collections: verify annotators' transcripts against audio using fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native German as spoken in Germany is a hard requirement; rationales are in English. Joint top rate of the Sonic set at $41.50/hour, remote.

    $41.50 / HourWorldwide
  • Test AI chat models in Marathi and English for safety failures (jailbreaks, bias, harmful answers) and judge whether their Marathi is accurate and appropriate rather than Hindi in disguise. Evaluation judgment is the core ask. Remote hourly contract, $16–22/hour, weekly pay.

    $16 – $22 / HourWorldwide
  • Remote hourly contract for colorists and video editors who already work in DaVinci Resolve on their own Mac. $30–40/hour, paid weekly via Stripe or Wise. The free version of Resolve is widely used; you also need a display above 2.5 megapixels. Tasks are not described.

    $30 – $40 / HourWorldwide
  • Help an AI lab evaluate a performance transfer model: define what a faithfully captured performance looks like, curate easy and hard benchmark examples, and shape the reviewer pool. For character animators with 4+ years and 2+ feature or AAA credits. US, sessions in LA or NYC, $60–90/hour.

    $60 – $90 / HourHybridOpen to United States
  • Join Mercor's bench of management consultants for future AI evaluation projects: writing grading criteria for consulting deliverables and scoring AI and human work. No project is open yet. For consultants with 1+ year at MBB or equivalent. Remote, $100–150/hour.

    $100 – $150 / HourWorldwide
  • Operational health physicists from DOE sites, national labs, reactors and decommissioning projects red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Audit applied machine-learning tasks used to train and evaluate a frontier AI lab's models: experiment design, model selection, evaluation methodology, leakage and metric gaming. For US practitioners with 3+ years of hands-on experimental ML in PyTorch, TensorFlow, scikit-learn or XGBoost. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Red-team frontier AI models from the chemical defence side: write benign, dual-use and adversarial prompts drawn from countermeasures, protection and detection work, judge how models respond against a policy standard, and write the reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Design Excel tasks from cross-industry experience (data models, dashboards, Power Query, VBA automation), write the solutions, and evaluate AI attempts for a tech company's GenAI team. Breadth across functions, teaching or a PhD help. US only, W-2 via Cincinnatus LLC, 40 then 20 hours a week. $70–100/hour.

    $70 – $100 / HourOpen to United States
  • Remote hourly contract for game developers, technical artists and virtual production or archviz specialists who already run Unreal Engine on their own Windows PC. $30–40/hour, paid weekly via Stripe or Wise. The engine is free to download; a display above 2.5 megapixels is required.

    $30 – $40 / HourWorldwide
  • Korean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.

    $63 – $67 / HourWorldwide
  • Native Korean speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models stay safe in Korean. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; South Korea or East Asia preferred, not required; 7 hires this month.

    $48 – $52 / HourWorldwide
  • Senior M&A, securities and governance lawyers design corporate law scenarios, reference memos and rubrics that test AI on deal work. US (Delaware, SEC) or international (UK Companies Act, EU) track. Remote hourly contract at $90–100/hour; 5+ years at a major firm, bank or large company.

    $90 – $100 / HourWorldwide
  • Hindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.

    $23 – $27 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.