Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.
  • QA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.

    $70 – $110 / HourWorldwide
  • Write or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.

    $94 – $119 / HourWorldwide
  • Write or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.

    $61 – $77 / HourWorldwide
  • Author executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.

    $70 / HourWorldwide
  • Senior materials scientists, and electrical or mechanical engineers, author realistic tasks with a prompt, a data room and a grading method, run them against an AI model and tighten them until the model can no longer reason through cleanly. Daily onboarding and office hours. Remote hourly contract at $60–90/hour.

    $60 – $90 / HourWorldwide
  • Record yourself from a trailing third-person camera while walking, moving through changing scenery and driving, to build video datasets for AI. Unlike micro1's 360° Video Recorder listing, you must already own both the camera and the mount. $20/hour, 100 openings.

    $20 / HourWorldwide
  • Complete self-contained fund-ops exercises from mock ledgers, statements and notices: reconciliations, NAV variance attribution, trade-break investigations, corporate actions and payment exceptions, each graded against a rubric. 3+ years in fund admin or investment ops. US only, about 15 hours a week, $75–110/hour.

    $75 – $110 / HourOpen to United States
  • Full-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.

    $100 – $150 / HourHybridOpen to United States
  • Remote hourly contract for developers on any stack who already use Visual Studio Code on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You need your own Mac and a display above 2.5 megapixels. Tasks are not described in the ad.

    $55 – $65 / HourWorldwide
  • Grade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.

    $100 – $150 / HourWorldwide
  • Build a real estate deal model in Excel from a synthetic document pack, with live formulas and sourced assumptions, then score two other contractors' models against fixed criteria. For associates at real-estate-focused mid-market PE funds with ~2 years of IB first. US only, about 20–25 hours over 1.5–2 weeks, $80–100/hour.

    $80 – $100 / HourOpen to United States
  • Full-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.

    $110 – $150 / HourHybridOpen to United States
  • Turn real lab and test-bench experience into AI benchmark tasks: build test logs, calibration records and failure reports, define the right diagnosis, and write 35+ point rubrics. Remote contractor, $30–70/hour paid per accepted task, 50 openings, 4+ years in test or production.

    $30 – $70 / HourWorldwide
  • Annotate and score robotics video footage against a scoring guide: label actions, objects and events, rate clips for quality and relevance, and QA your own annotations. Mid-level annotation experience preferred, robotics footage experience heavily preferred. Flat $7/hour, 20 openings, remote contract.

    $7 / HourWorldwide
  • Tag, segment and mark keyframes in video using Final Cut Pro on macOS to build AI training datasets. Professional or academic Final Cut experience counts; aimed at entry to mid-level editors. One opening, contractor, $15–80/hour.

    $15 – $80 / HourWorldwide
  • Mechanical designers build and validate parametric FreeCAD models of parts and assemblies, applying GD&T, to create a CAD dataset for AI training. Optional Python scripting. Remote contractor role, 50 openings, $40–100/hour.

    $40 – $100 / HourWorldwide
  • Turn everyday insurance judgment into AI training data: design realistic underwriting, claims, actuarial and compliance scenarios, review AI outputs and write feedback. Open to every insurance specialty, 3+ years, US-based. $75/hour.

    $75 / HourOpen to United States
  • Judge Vietnamese audio clips, from human speakers and AI speech models, for tones, pronunciation and naturalness, and back each rating with a written English explanation. Native Vietnamese and B2+ English preferred. 100 openings, contractor, $30–65/hour.

    $30 – $65 / HourWorldwide
  • A full-time, salaried research role at micro1 designing benchmarks, rubrics, datasets and evaluation pipelines for frontier coding agents. Base salary $200,000–260,000 plus equity and benefits, remote, one opening. Three years in software engineering, ML or evaluation.

    $200000 – $260000 / YearWorldwide
  • Listen to Bengali audio clips and rate how native and fluent the speaker or AI model sounds, then justify each rating in written English. Native Bengali plus B2 English; no AI experience needed. 100 openings, contractor, $30–65/hour.

    $30 – $65 / HourWorldwide
  • Native or near-native Telugu speakers transcribe audio and video, segment and label speech data, and proofread existing Telugu transcripts for an AI language data project, flagging dialect and cultural nuance. No AI experience needed. Ten openings, contractor, $10–24/hour.

    $10 – $24 / HourWorldwide
  • Clinicians whose main practice is young abuse survivors review AI mental-health guidance on abuse cases, build case scenarios, and write trauma-informed best-practice content so models respond safely to vulnerable youth. MD or PhD (psychology or social work) preferred. $100–200/hour, 8 openings, contractor, remote.

    $100 – $200 / HourWorldwide
  • A part-time fellowship for federal civil litigators: draft and evaluate motions, judge where AI-written advocacy falls short of persuasive, and build the grading criteria that measure it. Requires at least three documented federal motion wins. 50 openings, $150–300/hour, remote contractor.

    $150 – $300 / HourWorldwide
  • Review AI-generated and human-written text for grammar, usage, meaning and instruction-following, then write short, objective feedback against detailed guidelines. Native U.S. English is required. Fifty openings, contractor, $40–50/hour, at least 20 hours a week on a four-week initial sprint.

    $40 – $50 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.