Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.

Filter jobs

Location
Show all 72 location options
Language: English
  • English14 jobs, applied. Activate to remove
  • Germanno roles alongside your other filters
  • Spanishno roles alongside your other filters
  • Frenchno roles alongside your other filters
  • Japaneseno roles alongside your other filters
  • Portugueseno roles alongside your other filters
Show all 49 language options
Field: Data, AI & ML
Level

14 open roles matching these filters

  • Remote hourly contract for biostatisticians, epidemiologists and applied economists who already run Stata SE on their own Windows PC. $45–55/hour, paid weekly via Stripe or Wise. Your own Stata SE licence and a display above 2.5 megapixels are required. Tasks are not described.

    $45 – $55 / HourWorldwide
  • A talent pool, not an open project: Mercor is collecting data scientists for future work evaluating how well AI does real data science, from writing grading criteria for analyses, models and A/B write-ups to scoring and justifying. 1+ year of experience at a top tech, AI or quant firm. $100–150/hour when work exists.

    $100 – $150 / HourWorldwide
  • Write point-in-time forecasts on specific swing-state Senate, governor and statewide races, and grade AI political analyses against your own. For state pollsters, campaign analysts and political scientists with live-race experience. US or Canada residents, $150–250/hour.

    $150 – $250 / HourOpen to Canada, United States
  • A paid expert conversation for engineers who build and run LLM agents in production: a short AI screening interview (no coding), then, if selected, a 30-minute live call on agent reliability, evaluation and internal adoption, paid $100–500 depending on depth of experience.

    $100 – $500 / TaskWorldwide
  • Hands-on ML researchers take on scoped, open-ended empirical problems: training image classifiers and generators from scratch, fine-tuning open-weight LLMs, adversarial robustness, compression under hard budgets, and multilingual pre-training. 3+ years of ML research (PhD counts). Remote hourly contract at $100–120/hour.

    $100 – $120 / HourWorldwide
  • Author point-in-time election and political-risk forecasts, document calls on polling and win probabilities, and grade AI analyses. For senior national forecasters, campaign analytics leads, pollsters and political-risk analysts with 8+ years. Remote, $150–250/hour.

    $150 – $250 / HourWorldwide
  • Map how an upcoming catalyst should ripple through suppliers, customers, competitors and substitutes, estimate direction and magnitude from primary filings, and grade AI analyses. For senior sector PMs and lead equity analysts with 8+ years. Remote, $150–250/hour.

    $150 – $250 / HourWorldwide
  • ML systems engineers write and evaluate training tasks for a frontier lab across GPU kernels, performance profiling, distributed debugging and LLM inference serving, plus the rubrics that grade them. 2+ years of hands-on ML infrastructure work. Canada, UK or US; 40 hours a week. $90–120/hour.

    $90 – $120 / HourOpen to Canada, United Kingdom and 1 more country
  • Audit applied machine-learning tasks used to train and evaluate a frontier AI lab's models: experiment design, model selection, evaluation methodology, leakage and metric gaming. For US practitioners with 3+ years of hands-on experimental ML in PyTorch, TensorFlow, scikit-learn or XGBoost. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Evaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Evaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Grade AI-generated slides, spreadsheets and documents for real-world data science quality, flagging factual, visual and presentation errors in structured written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.

    $100 – $150 / HourWorldwide
  • An expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.

    $80 – $150 / TaskWorldwide
  • Author executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.

    $70 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.