Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.
  • Design marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.

    $60 – $80 / HourOpen to United States
  • Build retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.

    $60 – $80 / HourOpen to United States
  • Senior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.

    $90 – $110 / HourOpen to United States
  • Evaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Design adversarial prompts, find jailbreaks and policy failures, and document vulnerabilities in frontier AI models across cyber, biosecurity, fraud, misinformation and political content. Hourly remote contract at $70–84/hour for residents of Europe, the UK and the US; 5+ years' relevant experience required.

    $70 – $84 / HourOpen to Albania, Austria and 38 more countries
  • Full-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.

    $60 – $100 / HourHybridOpen to United States
  • Coordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.

    $40 – $60 / HourOpen to United States
  • Evaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.

    $60 – $80 / HourOpen to United States
  • Evaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.

    $65 – $90 / HourOpen to United States
  • Full-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.

    $100 – $150 / HourHybridOpen to United States
  • Full-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.

    $110 – $150 / HourHybridOpen to United States
  • Read clinical images of skin lesions, describe them in precise terms, judge how well AI assessments hold up, and write the criteria that define a high-quality dermatological read. Non-clinical, no patient care. Active US licence and five years post-residency. A flat $270/hour.

    $270 / HourOpen to United States
  • Design the primers, plasmids, gRNAs, mRNA constructs and repair templates that become ground truth for a frontier model, and write the rubrics that judge sequence design quality. PhD strongly preferred, first-author record expected, 20 hours a week minimum. US only, $70–105/hour.

    $70 – $105 / HourOpen to United States
  • Write difficult problems in your own field for a top AI lab's language models. Open to PhDs (or equivalent industry experience) in medicine, statistics, AI/ML, computer science, game development or mechanical and aerospace engineering, living in the US, UK, Canada or Australia. Remote hourly contract at $55–80/hour, paid weekly.

    $55 – $80 / HourOpen to Australia, Canada and 2 more countries

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.