Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.
  • Native French speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models behave safely in French. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; France or Western Europe preferred, not required; 7 hires this month.

    $48 – $52 / HourWorldwide
  • Arabic-speaking PhD chemists and biologists write specialised science prompts in Arabic and grade AI answers for accuracy and dual-use safety. Part-time remote at $38–42/hour; Saudi Arabia or MENA preferred, not required. 13 hires this month, the most active listing in the series.

    $38 – $42 / HourWorldwide
  • Evaluate AI-written Punjabi song lyrics for a leading AI lab: check them against published songs for similarity, rate quality, creativity, prompt adherence and originality, and judge whether slang and regional expressions ring true. For Punjabi songwriters, performers or music writers. Remote, flexible, up to 6 months, $15/hour.

    $15 / HourWorldwide
  • Assess AI-generated Norwegian song lyrics for a leading AI lab: compare with published songs for similarity, rate quality, creativity, prompt adherence and originality, and judge whether word choice and dialect sound natural. For Norwegian songwriters, performers or music journalists. Remote, flexible, up to 6 months, $42–78/hour.

    $42 – $78 / HourWorldwide
  • Write original, research-frontier questions in your own discipline that current AI models cannot answer, with sourced and cited reference answers, then test and harden them. Open to any field. Paid per accepted task, $40–90/hour band, five openings, remote contractor.

    $40 – $90 / HourWorldwide
  • Native or near-native German editors, critics and linguists evaluate LLM-generated German text, rewrite it to literary and cultural standards, annotate linguistic features and give feedback to researchers at a leading AI lab. Remote hourly contract at $50/hour.

    $50 / HourWorldwide
  • Remote hourly contract for Python developers who already run PyCharm on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You supply the licence, the Mac and a display above 2.5 megapixels. Tasks are not described in the ad.

    $55 – $65 / HourWorldwide
  • Native Thai speakers use a test build of an AI voice assistant on their own Apple Silicon Mac to do everyday tasks by voice, then transcribe their queries, rate the responses and log the data. Paid $12.6 per task, 4-week contract starting immediately. Remote.

    $12.60 / TaskWorldwide
  • Umbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.

    $65 – $75 / TaskWorldwide
  • Umbrella listing for Mercor's radiological safety red-team panel: RSOs, health physicists, source security, emergency response and nuclear medicine specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 8 hired this month.

    $65 – $75 / TaskWorldwide
  • Native or near-native Spanish translators, editors and annotators evaluate LLM-generated Spanish, correct or rewrite it, annotate grammatical and semantic features, and flag model errors for researchers at a leading AI lab. Remote hourly contract at $50/hour.

    $50 / HourWorldwide
  • RF, microwave, antenna and electromagnetics engineers solve and critique hard technical problems for a short-term expert evaluation project: analysing systems and trade-offs, checking calculations and assumptions, and writing rigorous explanations. Hands-on industry experience and an EE-family degree. Remote hourly contract at $80–95/hour.

    $80 – $95 / HourWorldwide
  • Build hard codebase-exploration tasks for an RL environment that trains AI agents: rewrite engineering questions about production Go repos (etcd, Traefik, Helm, gRPC-Go, Temporal and more) so frontier agents fail them, tune rubrics and foils, and pass a validation loop. $130 per approved task, remote.

    $130 / TaskWorldwide
  • Transcribe Italian audio, and English when required, edit transcripts for completeness, annotate data for AI training and flag unclear or low-quality recordings. Native Italian and professional transcription experience. 15 openings, contractor, $20–36/hour.

    $20 – $36 / HourWorldwide
  • QA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.

    $70 – $110 / HourWorldwide
  • Write or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.

    $94 – $119 / HourWorldwide
  • Write or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.

    $61 – $77 / HourWorldwide
  • Author executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.

    $70 / HourWorldwide
  • Senior materials scientists, and electrical or mechanical engineers, author realistic tasks with a prompt, a data room and a grading method, run them against an AI model and tighten them until the model can no longer reason through cleanly. Daily onboarding and office hours. Remote hourly contract at $60–90/hour.

    $60 – $90 / HourWorldwide
  • Record yourself from a trailing third-person camera while walking, moving through changing scenery and driving, to build video datasets for AI. Unlike micro1's 360° Video Recorder listing, you must already own both the camera and the mount. $20/hour, 100 openings.

    $20 / HourWorldwide
  • Remote hourly contract for developers on any stack who already use Visual Studio Code on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You need your own Mac and a display above 2.5 megapixels. Tasks are not described in the ad.

    $55 – $65 / HourWorldwide
  • Grade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.

    $100 – $150 / HourWorldwide
  • Turn real lab and test-bench experience into AI benchmark tasks: build test logs, calibration records and failure reports, define the right diagnosis, and write 35+ point rubrics. Remote contractor, $30–70/hour paid per accepted task, 50 openings, 4+ years in test or production.

    $30 – $70 / HourWorldwide
  • Annotate and score robotics video footage against a scoring guide: label actions, objects and events, rate clips for quality and relevance, and QA your own annotations. Mid-level annotation experience preferred, robotics footage experience heavily preferred. Flat $7/hour, 20 openings, remote contract.

    $7 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.