Skip to content
Labeling Jobs

Remote AI training and data labeling jobs

Every role here has been checked against the platform that posted it. Pay is shown as reported, and marked when it is an estimate rather than a firm rate.

Filter jobs

Location
Show all 72 location options
Language
  • English38 jobs
  • Germanno roles alongside your other filters
  • Spanishno roles alongside your other filters
  • Frenchno roles alongside your other filters
  • Japaneseno roles alongside your other filters
  • Portugueseno roles alongside your other filters
Show all 49 language options
Field: Software & IT
Level: Senior

42 open roles matching these filters · page 2 of 2

  • Evaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.

    $70 – $90 / HourOpen to United States
  • Senior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.

    $90 – $110 / HourOpen to United States
  • Evaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.

    $70 – $90 / HourOpen to United States
  • An expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.

    $80 – $150 / TaskWorldwide
  • Grade AI-generated slides, spreadsheets and documents for real-world software engineering quality, flagging factual, visual and presentation errors with written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.

    $100 – $150 / HourWorldwide
  • A full-time, salaried research role at micro1 designing benchmarks, rubrics, datasets and evaluation pipelines for frontier coding agents. Base salary $200,000–260,000 plus equity and benefits, remote, one opening. Three years in software engineering, ML or evaluation.

    $200000 – $260000 / YearWorldwide
  • Build reinforcement learning environments that test whether an AI model can deploy, secure, scale and recover production cloud infrastructure: realistic scenarios, deterministic tests, golden solutions and deliberately broken variants. 20 hours a week, paid per accepted task, $50–100/hour band.

    $50 – $100 / HourWorldwide
  • Write hard feature and bug-fix tasks for AI coding models, plus deterministic verifiers that accept any valid solution and reject the rest. About 15 flexible hours a week, paid per accepted task on a $30–100/hour band, 300 openings; security backgrounds are a plus.

    $30 – $100 / HourWorldwide
  • Experienced IT leaders plan and run a small technology initiative with an external vendor, SLAs and stakeholder reporting, generating realistic management scenarios to train AI. Contractor, remote, 50 openings, $70–130/hour. Six years leading multi-quarter tech initiatives is the bar.

    $70 – $130 / HourWorldwide
  • Build features, fix hard bugs and refactor legacy code inside existing open source projects, with clear write-ups of each solution, as training input for AI coding models. Expert level in two languages plus open source and competitive programming experience. 37 openings, paid per task on a $50–100/hour band.

    $50 – $100 / HourWorldwide
  • Build, tune and evaluate ML models in Python with MongoDB-backed data pipelines for a customer AI training project, and document every experiment. Remote contractor, $80–140/hour, 35 openings. scikit-learn, TensorFlow or PyTorch plus hands-on MongoDB is the core stack.

    $80 – $140 / HourWorldwide
  • Technical experts audit the tasks used to train and evaluate AI systems, checking each is accurate, realistic, solvable, reproducible and properly tested. You are matched to one of eight specialties, from GPU kernels and Kubernetes to CVE security and SWE-Bench, and take an assessment in it. $60/hour, remote contractor.

    $60 / HourWorldwide
  • Build and document the production pipelines you'd work on during a normal week, alongside experts from two other disciplines, so the outputs can be turned into the tasks and rubrics that train frontier models. $140–200/hour for 40 hours a week from 14 September to 10 October 2026. You must be based in the UK with the right to work there.

    $140 – $200 / HourOpen to United Kingdom
  • Build reinforcement learning environments that test whether an AI model can fix bugs, add features, refactor code and tune performance in real repositories, then write the golden reference solution it is graded against. Public open-source contributions are the requirement, not AI experience. Remote contractor work, roughly 15 hours a week.

    $50 – $150 / HourWorldwide
  • Stress-test autonomous coding agents: review the code they write, find the edge cases where they break down, and build the rubrics and harnesses used to score them. Task-based, fully remote, 10–40 hours a week at your own pace.

    $80 – $120 / HourWorldwide
  • Design the coding benchmarks that decide whether a frontier AI model can actually reason, debug and ship production code, and build the data pipelines behind those evaluations. A three-month remote hourly contract, full-time availability preferred.

    $80 – $100 / HourWorldwide
  • Build and maintain the backend services, APIs and data pipelines behind AI training, evaluation and deployment. A fully remote hourly contract at 20–40 hours a week, aimed at engineers with five or more years of production experience.

    $60 – $100 / HourWorldwide
  • Build reinforcement learning environments that test whether an AI agent can fix bugs, implement features, refactor and optimise code while discovering and reasoning over information from real MCP servers. You design the reproducible environment, the deterministic verification and the golden reference solution. Remote contractor work, roughly 15 hours a week, hours of your choosing.

    $60 – $120 / HourWorldwide

Nothing that fits today?

New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.