Write realistic feature and bug-fix tasks for AI coding models, plus deterministic verifiers that accept valid solutions and reject the rest. About 15 flexible hours a week, paid per accepted task on a $30–100/hour band, 100 openings; a strong testing background is the stated plus.
Remote AI training and data labeling jobs
Filter jobs
Location: Worldwide
- Worldwide32 jobs, applied. Activate to remove
- United States9 jobs
- United Kingdom2 jobs
- Canada1 job
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English29 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japaneseno roles alongside your other filters
- Portugueseno roles alongside your other filters
Field: Software & IT
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering49 jobs
- Business & Finance30 jobs
- Software & IT32 jobs, applied. Activate to remove
- Health & Medicine54 jobs
- Law, Policy & Security35 jobs
- General & Data Collection1 job
- Science & Math41 jobs
- AI Safety & Evaluation20 jobs
- Video, Image & Design3 jobs
- Data, AI & ML13 jobs
- Writing & Education2 jobs
- Other fields5 jobs
Newest
32 open roles matching these filters · page 1 of 2
- $30 – $100 / HourWorldwide
Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.
$90 – $110 / HourWorldwideEmbedded firmware, FPGA/RTL, flight software and hardware test automation engineers review and write hard problems about software that runs on real hardware (timing, race conditions, interfaces, bring-up) for an AI research initiative. 5+ years, 3 recent hands-on. Remote contract, $100–120/hour.
$100 – $120 / HourWorldwideMaintainers of real open source repositories review LLM-generated patches against their own codebase, justify each verdict in writing, and peer-review other experts. Open globally, about 15 hours a week, paid per accepted task at a stated $150–300/hour equivalent. 10 openings.
$150 – $300 / HourWorldwideExperienced IT and systems managers: critique IT strategy, architecture, vendor and security-compliance scenarios and write the expert input that trains AI tools for IT professionals. 5+ years leading technical teams or enterprise infrastructure. 5 openings, contractor, $60–75/hour.
$60 – $75 / HourWorldwideA short, well-paid sprint for very senior software engineers: help a leading foundation-model lab improve its models on hard SWE tasks. 10+ years at top US tech firms, about 20 hours a week for 2–3 weeks. $150–210/hour by geography and level; the 2-hour vetting exercise is paid $100.
$150 – $210 / HourWorldwideDesign GPU programming tasks in CUDA, WebGPU or GLSL for training LLMs on performance and architecture, including kernel profiling and C++ host code. Remote contractor, $60–95/hour paid per accepted task, 25 openings. Graphics, HPC and ML acceleration backgrounds all qualify.
$60 – $95 / HourWorldwideCreate, solve and review hard full-stack engineering tasks, debug across browser, API and database layers, judge AI-generated solutions, and port whole builds from one language to another. Global and fully remote, about 15 flexible hours a week, paid per task on a $50–100/hour band, 300 openings.
$50 – $100 / HourWorldwideA paid expert conversation for engineers who build and run LLM agents in production: a short AI screening interview (no coding), then, if selected, a 30-minute live call on agent reliability, evaluation and internal adoption, paid $100–500 depending on depth of experience.
$100 – $500 / TaskWorldwideExpert C# and .NET developers with strong MySQL skills build and maintain backend services and APIs on a micro1 project framed as AI training. Remote contract, only 2 openings, $60–110/hour.
$60 – $110 / HourWorldwideBuild reinforcement learning environments where an AI agent must fix bugs, add features or optimise code using real MCP servers, with deterministic verifiers and golden solutions. About 15 hours a week, your own schedule, paid per accepted task against a $60–120/hour band. 100 openings.
$60 – $120 / HourWorldwideContribute real-world code, explanations and code reviews as AI training data, drawing on five-plus years in Python, Java, JavaScript/Node.js or full-stack work. A narrow $60–75/hour band, 15 openings, remote contract with documentation and architecture writing at its core.
$60 – $75 / HourWorldwideDesign reinforcement learning environments where an AI agent must use real Model Context Protocol (MCP) servers to fix bugs, build features, refactor or optimise code, with deterministic verification and golden reference solutions. About 15 hours a week, paid per accepted task on an $80–120/hour band, 100 openings.
$80 – $120 / HourWorldwideCreate, solve and review hard software-engineering tasks in real repositories, including porting whole builds between languages and reviewing AI-generated patches. Global remote contractor, $100–150/hour paid per accepted task, about 15 hours a week. A verifiable record of substantial open-source contributions is required.
$100 – $150 / HourWorldwideGPU specialists profile and optimise CUDA kernels, refactor C++ and CUDA code, and write GLSL and WebGPU shaders on a project run with a leading AI lab. Remote contractor role, 50 openings, $60–100/hour.
$60 – $100 / HourWorldwideBuild reinforcement learning environments that test whether AI models can fix bugs, add features, refactor and optimise real code, each with a reproducible setup and a golden reference solution. Around 15 hours a week, paid per accepted task on a $100–150/hour band, 100 openings, no AI experience needed.
$100 – $150 / HourWorldwidePackage real bug fixes, features, refactors and performance problems as reproducible reinforcement learning environments with golden reference solutions for AI coding models. About 15 hours a week, paid per accepted task on a $50–100/hour band, 100 openings. Same text as a higher-paying sibling posting.
$50 – $100 / HourWorldwideTurn your favourite open repositories into reinforcement learning environments for frontier coding models: subtle bugs, non-trivial features or performance problems, each with a reproducible setup, robust verifiers and a reference solution. About 15 hours a week, paid per task on a $50–100/hour band.
$50 – $100 / HourWorldwideAn expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.
$80 – $150 / TaskWorldwideGrade AI-generated slides, spreadsheets and documents for real-world software engineering quality, flagging factual, visual and presentation errors with written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideA full-time, salaried research role at micro1 designing benchmarks, rubrics, datasets and evaluation pipelines for frontier coding agents. Base salary $200,000–260,000 plus equity and benefits, remote, one opening. Three years in software engineering, ML or evaluation.
$200000 – $260000 / YearWorldwideBuild reinforcement learning environments that test whether an AI model can deploy, secure, scale and recover production cloud infrastructure: realistic scenarios, deterministic tests, golden solutions and deliberately broken variants. 20 hours a week, paid per accepted task, $50–100/hour band.
$50 – $100 / HourWorldwideWrite hard feature and bug-fix tasks for AI coding models, plus deterministic verifiers that accept any valid solution and reject the rest. About 15 flexible hours a week, paid per accepted task on a $30–100/hour band, 300 openings; security backgrounds are a plus.
$30 – $100 / HourWorldwideExperienced IT leaders plan and run a small technology initiative with an external vendor, SLAs and stakeholder reporting, generating realistic management scenarios to train AI. Contractor, remote, 50 openings, $70–130/hour. Six years leading multi-quarter tech initiatives is the bar.
$70 – $130 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.