Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.
  • Korean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.

    $63 – $67 / HourWorldwide
  • Hindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.

    $23 – $27 / HourWorldwide
  • Portuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.

    $50 – $54 / HourWorldwide
  • India-based full-stack engineers build real applications on top of a leading AI lab's pre-release models, wiring in tool interfaces, evaluation harnesses and telemetry, and writing up the model failures they hit. 3+ years at a top-tier organisation and two programming languages. Full-time, 40 hours a week, $25–30/hour.

    $25 – $30 / HourOpen to India
  • Evaluate generative music AI for a leading AI lab: compare AI-made songs head to head on musicality, prompt adherence, vocals and mix, label genre and structure, and check lyrics and vocals. For Thai-speaking producers or engineers with 2+ years' experience. Remote, flexible hours, up to 6 months, $18/hour.

    $18 / HourWorldwide
  • Rate AI-generated music for a leading AI lab: head-to-head song comparisons on musicality, prompt adherence, vocals and mix, genre and structure labelling, and lyric and vocal checks. For Russian-speaking producers and mix engineers with 2+ years' experience. Remote, flexible hours, up to 6 months, $35–49/hour.

    $35 – $49 / HourWorldwide
  • Turn ambiguous AI program requirements into clear, contradiction-free rater guidelines and rubrics across finance, retail, insurance, legal and sports. For linguists, instructional designers and technical writers with 3+ years and GenAI/RLHF guideline experience. US, 35+ hours a week, $45–65/hour.

    $45 – $65 / HourOpen to United States
  • Audit Kubernetes tasks used to train and evaluate a frontier AI lab's models: cluster-operations scenarios, manifest and Helm correctness, and failure-mode troubleshooting (CrashLoopBackOff, OOMKilled, eviction). For US engineers with 3+ years of production Kubernetes and Go, Python or TypeScript. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Experienced Spanish tutors lead recorded one-on-one lessons with native English-speaking learners so a language-learning platform can capture authentic tutoring conversations. 3+ years' teaching and a degree required. About 7 hours a week for roughly three weeks, $60–65/hour, remote.

    $60 – $65 / HourWorldwide
  • Arabic-speaking PhD chemists and biologists write specialised science prompts in Arabic and grade AI answers for accuracy and dual-use safety. Part-time remote at $38–42/hour; Saudi Arabia or MENA preferred, not required. 13 hires this month, the most active listing in the series.

    $38 – $42 / HourWorldwide
  • Evaluate vulnerability-reproduction and remediation tasks for a frontier AI lab: faithful CVE reproductions in Docker labs, sound fixes, and two-part verification (functionality plus vulnerability tests). For US AppSec engineers, pentesters and vulnerability researchers with 3+ years. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Author AI evaluation tasks from real drawing sets, documents and site photos, with the correct RFI response, coordination comments or markup as the answer. For licensed architects, project architects and job captains with 3+ years. US only, $45–60/hour.

    $45 – $60 / HourOpen to United States
  • Remote hourly contract for Python developers who already run PyCharm on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You supply the licence, the Mac and a display above 2.5 megapixels. Tasks are not described in the ad.

    $55 – $65 / HourWorldwide
  • Build hard codebase-exploration tasks for an RL environment that trains AI agents: rewrite engineering questions about production Go repos (etcd, Traefik, Helm, gRPC-Go, Temporal and more) so frontier agents fail them, tune rubrics and foils, and pass a validation loop. $130 per approved task, remote.

    $130 / TaskWorldwide
  • Complete self-contained fund-ops exercises from mock ledgers, statements and notices: reconciliations, NAV variance attribution, trade-break investigations, corporate actions and payment exceptions, each graded against a rubric. 3+ years in fund admin or investment ops. US only, about 15 hours a week, $75–110/hour.

    $75 – $110 / HourOpen to United States
  • Remote hourly contract for developers on any stack who already use Visual Studio Code on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You need your own Mac and a display above 2.5 megapixels. Tasks are not described in the ad.

    $55 – $65 / HourWorldwide
  • Build a real estate deal model in Excel from a synthetic document pack, with live formulas and sourced assumptions, then score two other contractors' models against fixed criteria. For associates at real-estate-focused mid-market PE funds with ~2 years of IB first. US only, about 20–25 hours over 1.5–2 weeks, $80–100/hour.

    $80 – $100 / HourOpen to United States
  • Turn real lab and test-bench experience into AI benchmark tasks: build test logs, calibration records and failure reports, define the right diagnosis, and write 35+ point rubrics. Remote contractor, $30–70/hour paid per accepted task, 50 openings, 4+ years in test or production.

    $30 – $70 / HourWorldwide
  • Mechanical designers build and validate parametric FreeCAD models of parts and assemblies, applying GD&T, to create a CAD dataset for AI training. Optional Python scripting. Remote contractor role, 50 openings, $40–100/hour.

    $40 – $100 / HourWorldwide
  • Turn everyday insurance judgment into AI training data: design realistic underwriting, claims, actuarial and compliance scenarios, review AI outputs and write feedback. Open to every insurance specialty, 3+ years, US-based. $75/hour.

    $75 / HourOpen to United States
  • Review AI-generated and human-written text for grammar, usage, meaning and instruction-following, then write short, objective feedback against detailed guidelines. Native U.S. English is required. Fifty openings, contractor, $40–50/hour, at least 20 hours a week on a four-week initial sprint.

    $40 – $50 / HourWorldwide
  • Open multi-part assemblies in FreeCAD, work out their design intent, and write several exact-answer technical questions per assembly that a model cannot guess its way through. Fifty openings, flagged high demand, contractor at $40–80/hour. FreeCAD experience helps but is not required.

    $40 – $80 / HourWorldwide
  • Grade and rank answers from AI chat and search tools against detailed guidelines, write prompts that probe edge cases, and reconcile ratings with the rubric. Generalist contractor role for Northern America and Europe, $14–36/hour, 50 openings, posted this week.

    $14 – $36 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.