Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.
  • ML systems engineers write and evaluate training tasks for a frontier lab across GPU kernels, performance profiling, distributed debugging and LLM inference serving, plus the rubrics that grade them. 2+ years of hands-on ML infrastructure work. Canada, UK or US; 40 hours a week. $90–120/hour.

    $90 – $120 / HourOpen to Canada, United Kingdom and 1 more country
  • Audit German speech data for Amazon's Sonic collections: verify annotators' transcripts against audio using fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native German as spoken in Germany is a hard requirement; rationales are in English. Joint top rate of the Sonic set at $41.50/hour, remote.

    $41.50 / HourWorldwide
  • Native Bengali speakers based in India map the layout of real Bengali PDF pages and transcribe every text region character for character in Bengali script, handwriting included, to train document AI. Remote hourly contract at $12.68/hour, with a second-expert review of every task.

    $12.68 / HourOpen to India
  • Test AI chat models in Marathi and English for safety failures (jailbreaks, bias, harmful answers) and judge whether their Marathi is accurate and appropriate rather than Hindi in disguise. Evaluation judgment is the core ask. Remote hourly contract, $16–22/hour, weekly pay.

    $16 – $22 / HourWorldwide
  • Remote hourly contract for colorists and video editors who already work in DaVinci Resolve on their own Mac. $30–40/hour, paid weekly via Stripe or Wise. The free version of Resolve is widely used; you also need a display above 2.5 megapixels. Tasks are not described.

    $30 – $40 / HourWorldwide
  • Help an AI lab evaluate a performance transfer model: define what a faithfully captured performance looks like, curate easy and hard benchmark examples, and shape the reviewer pool. For character animators with 4+ years and 2+ feature or AAA credits. US, sessions in LA or NYC, $60–90/hour.

    $60 – $90 / HourHybridOpen to United States
  • Join Mercor's bench of management consultants for future AI evaluation projects: writing grading criteria for consulting deliverables and scoring AI and human work. No project is open yet. For consultants with 1+ year at MBB or equivalent. Remote, $100–150/hour.

    $100 – $150 / HourWorldwide
  • Operational health physicists from DOE sites, national labs, reactors and decommissioning projects red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Audit applied machine-learning tasks used to train and evaluate a frontier AI lab's models: experiment design, model selection, evaluation methodology, leakage and metric gaming. For US practitioners with 3+ years of hands-on experimental ML in PyTorch, TensorFlow, scikit-learn or XGBoost. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Red-team frontier AI models from the chemical defence side: write benign, dual-use and adversarial prompts drawn from countermeasures, protection and detection work, judge how models respond against a policy standard, and write the reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • Design Excel tasks from cross-industry experience (data models, dashboards, Power Query, VBA automation), write the solutions, and evaluate AI attempts for a tech company's GenAI team. Breadth across functions, teaching or a PhD help. US only, W-2 via Cincinnatus LLC, 40 then 20 hours a week. $70–100/hour.

    $70 – $100 / HourOpen to United States
  • Remote hourly contract for game developers, technical artists and virtual production or archviz specialists who already run Unreal Engine on their own Windows PC. $30–40/hour, paid weekly via Stripe or Wise. The engine is free to download; a display above 2.5 megapixels is required.

    $30 – $40 / HourWorldwide
  • Korean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.

    $63 – $67 / HourWorldwide
  • Native Korean speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models stay safe in Korean. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; South Korea or East Asia preferred, not required; 7 hires this month.

    $48 – $52 / HourWorldwide
  • Senior M&A, securities and governance lawyers design corporate law scenarios, reference memos and rubrics that test AI on deal work. US (Delaware, SEC) or international (UK Companies Act, EU) track. Remote hourly contract at $90–100/hour; 5+ years at a major firm, bank or large company.

    $90 – $100 / HourWorldwide
  • Hindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.

    $23 – $27 / HourWorldwide
  • Audit Spanish (Spain) speech data for Amazon's Sonic project: check annotators' transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Spanish as spoken in Spain is a hard requirement; rationales are in English. Remote hourly contract at $39.50/hour.

    $39.50 / HourWorldwide
  • Portuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.

    $50 – $54 / HourWorldwide
  • Named RSOs and radiation protection managers on broad-scope, hospital, university or industrial licences red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.

    $65 – $75 / TaskWorldwide
  • India-based full-stack engineers build real applications on top of a leading AI lab's pre-release models, wiring in tool interfaces, evaluation harnesses and telemetry, and writing up the model failures they hit. 3+ years at a top-tier organisation and two programming languages. Full-time, 40 hours a week, $25–30/hour.

    $25 – $30 / HourOpen to India
  • Paid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.

    $130 – $210 / HourOpen to United States
  • Evaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Board-certified radiologists label findings, write reference reports, grade AI-generated reads and author rubrics for medical imaging AI. Non-clinical, remote worldwide, hourly contract at $200–400/hour with a 15-hour weekly minimum. Weekly pay via Stripe or Wise.

    $200 – $400 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.