AI training work for software engineers: the five kinds of job and what each pays
Platform comparison · 2 weeks ago
Engineering is where AI training pays best, but the listings describe very different jobs under similar titles. Using more than a hundred engineering roles on this site, here are the five kinds of work (auditing, task authoring, environment building, reviewing agent output, building on pre-release models), what each pays, who gets in, and how to pick.
Why engineers are the best-paid part of this market
AI labs are now training models to act as software engineers: to fix bugs in real repositories, operate a terminal, write kernels, configure clusters. You cannot grade that work, or write the tests for it, without being able to do it yourself. That is why engineering listings sit at the top of this board. Of the more than 100 roles on this site that ask for real programming or infrastructure skills, most advertise $50 an hour or more, and a handful reach $150 to $300.
The titles are unhelpful, though. "Senior Software Engineer" on one platform means writing benchmark tasks; on another it means a full-time job building apps on a pre-release model. Read by what you actually do, the listings fall into five kinds of work.
1. Auditing tasks someone else wrote
What it is: a lab has a pile of training or benchmark tasks (a Kubernetes incident to diagnose, a CVE to reproduce, a repository bug with a reference patch) and needs an expert to check each one is correct, realistic, solvable and fairly tested. You write rubric-based feedback; you do not write the task.
Examples and pay: Mercor runs a family of US-only auditor listings at $70 to $90 an hour: SWE-Bench Task Auditor, Kubernetes Task Auditor, CVE Vulnerability Expert, ML Challenge Task Auditor, AI Developer Trace Task Auditor, GPU Kernel Expert, Trainium NKI Kernel Expert and AWS Serverless and IaC Task Auditor.
Invisible Technologies' Technical Task Auditor covers the same eight specialty areas (GPU kernels, AWS Trainium and NKI, CVE and application security, machine learning, AI-assisted developer workflows, AWS serverless and IaC, Kubernetes, SWE-bench) in one listing at $60 an hour, with an assessment in your chosen specialty. Neither platform names the client, so we cannot say whether it is the same project. What we can say is that the Mercor listings are US-only while the Invisible listing states no country requirement, which makes it the route for engineers outside the US.
Who gets in: typically three or more years in the specialty, and production depth rather than tutorial familiarity. The SWE-bench auditor asks for real open-source contributor or maintainer history; the kernel roles want kernel work in at least two frameworks.
Why choose it: the most predictable engineering work here. Hourly pay, a clear bar, and it uses judgment you already have. The trade-off is repetition: auditing the fiftieth Helm chart is not the same as building one.
2. Writing tasks and benchmarks
What it is: you create the problems. A realistic bug or feature request, the starting repository or environment, a reference solution, and a test or verifier that accepts every correct answer and rejects the rest. The verifier is the hardest part of the job and the one that matters most.
Examples and pay: micro1's Senior Software Engineer (Coding Tasks and Verifiers) pays per accepted task on a $30 to $100 hourly band. Its Competitive Coder listing wants original problems with C++ checkers at $45 to $65. Mercor's Go codebase Q&A pays $130 per approved task for questions about production Go repositories that frontier agents fail. Invisible's terminal and infrastructure scenarios project pays $35 an hour to write Terminal-Bench-style tasks with an automated verifier. Alignerr's AI evaluation benchmarks role pays $80 to $100 for a three-month contract designing coding benchmarks.
Who gets in: strong engineers who enjoy adversarial thinking. The task must defeat a frontier model without being unfair, and a verifier that can be gamed gets rejected.
Why choose it: the most intellectually interesting work of the five. But look closely at the pay model. Most of these are paid per accepted task, so the hourly band describes a fast contributor whose tasks pass review. Your first few tasks will almost certainly take longer and be rejected more often. The existing guide on how you actually get paid covers what to ask.
3. Building RL environments
What it is: a step beyond single tasks. You build a reproducible environment (a repository, tools, sometimes real servers) in which an agent must fix bugs, add features or refactor, with deterministic verification and a golden solution.
Examples and pay: micro1's RL environments listing pays per accepted task on a $100 to $150 band, around 15 hours a week, and its MCP environments listing ($80 to $120) has agents working through real Model Context Protocol servers. Alignerr's agentic coding role ($80 to $120) mixes reviewing agent code with building the harnesses that score it.
Who gets in: senior engineers comfortable with Docker, test infrastructure and making things reproducible. Several of these listings state no AI experience is needed.
Why choose it: high bands and real engineering. The same per-task caveat applies, and environment work has more ways to fail review than a single task does.
4. Reviewing what agents produce
What it is: a model has written a patch, a pull request or a whole session of AI-assisted coding, and you judge it as a senior reviewer would.
Examples and pay: micro1's Open Source GitHub Maintainer is the standout: maintainers review LLM-generated patches for their own repositories at a stated $150 to $300 an hour, paid per task, about 15 hours a week, open globally. As the title says, you must actually maintain an open-source project. micro1's developer workflow evaluation ($50 to $70) runs pull requests, code review and CI through AI tools and grades the results.
Who gets in: for the top of this category, proof that others trust your code review, usually a public maintainer record.
Why choose it: if you maintain something, it is the best-paid use of knowledge nobody else has. If you do not, look at category 1 instead.
5. Building on pre-release models
What it is: closer to a normal engineering job. You build real applications on a lab's unreleased models and document where they fail.
Examples and pay: Mercor's full-stack family shows how location drives pay for identical work: Software Engineer, Full Stack (US) at $50 to $65, the senior US tier at $90 to $110 for 6+ years, and the India tier at $25 to $30 for the same three-year bar. The US roles are full-time W-2 through an employer of record. At the far end, Mercor's Expert Senior SWE is a two-to-three-week sprint at $150 to $210 for engineers with 10+ years at top US tech firms.
Who gets in: engineers who can commit 40 hours a week (for the full-time roles) and show a strong employer history.
Why choose it: stability and, in the US, benefits. It is a job rather than a side engagement, which is either the point or a dealbreaker.
How to choose
- Outside the US? Most of Mercor's auditor and full-stack roles are US-only. Invisible's Technical Task Auditor, micro1's per-task roles and Alignerr's roles state no country requirement. Check each listing.
- Want predictable pay? Prefer hourly listings (categories 1 and 5). Per-task bands (categories 2 to 4) reward speed and experience with the review process.
- Maintain an open-source project? Category 4 is built for you and pays the most.
- Deep in one specialty (kernels, Kubernetes, AppSec)? Auditing uses it directly and the assessments are specialty-specific.
- Enjoy breaking things? Task and environment authoring, but expect a learning curve on the verifier.
Two cautions that apply everywhere. Several listings require H-1B and STEM OPT holders to look elsewhere; Mercor and Invisible both state it. And much of this work involves real repositories and your own machine, so read our privacy checks for screen recording and personal accounts before you record or connect anything.
Sources
Pay, eligibility and requirements are taken from the listings on this site as read between 24 and 26 September 2026, which quote each platform's own ad. The counts come from the 632 open roles published here on 26 September 2026. Platforms change these listings often, so the listing page always outranks this guide. We are not affiliated with any of these platforms and we do not process applications.
Questions
- What is the highest-paying AI training work for software engineers?
- On this site, reviewing LLM-generated patches for an open-source repository you maintain (micro1, stated $150 to $300 an hour, paid per task) and short sprints for very senior engineers (Mercor Expert Senior SWE, $150 to $210). Both have narrow entry requirements. The broadest well-paid category is task auditing at $60 to $90 an hour.
- Can I do engineering AI training work from outside the United States?
- Yes, but check each listing. Mercor's task-auditor family and its full-stack roles are US-only, with a separate India tier for full-stack work. Invisible Technologies' Technical Task Auditor, micro1's per-task engineering roles and Alignerr's engineering roles state no country requirement.
- Is engineering AI training work paid hourly or per task?
- Both, and the listing title does not tell you which. Auditing and full-stack roles are usually hourly. Task writing, RL environment building and patch review are usually paid per accepted task, with the hourly band describing a fast contributor whose work passes review. Ask for the per-task fee and typical task time before accepting.
- Do I need AI or machine learning experience?
- Usually not. Most of these listings want software engineering depth in a specialty (Kubernetes, kernels, security, a language ecosystem) and several say explicitly that no AI experience is needed. The ML Challenge auditor is the exception, since the tasks themselves are about machine learning.
Platforms covered here
Put this into practice
Every listing shows its pay and who it is open to.