What AI training work actually is
Overview · 8 September 2026
"AI training" covers three different jobs: labeling data, judging what a model produced, and doing your own professional work so a model can learn from it. They pay differently, screen differently, and want different people. Here is what separates them, and which one the listings on this board are actually offering.
Three jobs share one name
Search for "AI training jobs" and you get results that have almost nothing in common. Some are hourly annotation work with no qualifications asked. Some want a completed medical residency. The phrase covers at least three distinct kinds of work, and knowing which one a listing is offering tells you most of what you need: what it pays, how hard it is to get, and whether your background is relevant at all.
- Data labeling. You add structure to raw material so a model can learn from it.
- Evaluation. A model has already produced something. You judge it.
- Expert work. You do your own professional job, and your output becomes the training material.
The three are often mixed inside a single listing. But they sit on a ladder, and the rungs are far apart.
Data labeling: adding the structure
This is the original version of the work and the one most people picture. You take unstructured material and mark it up: drawing boxes around objects in an image, tagging which speaker is talking in an audio clip, marking a sentence as a complaint or a compliment, transcribing speech.
The defining feature is that there is a correct answer and a written guideline telling you what it is. Your job is to apply the guideline consistently, at volume, faster than the person next to you. Quality is usually measured by agreement: how often you match the guideline and other annotators on the same item.
On this board, the Video Annotation Specialist role at micro1 is close to pure labeling, and the Audio Expert transcription role is a demanding technical version of it: reconstructing exact musical scores from mixed audio and converting them to MusicXML.
Evaluation: judging what the model produced
Here the model has already answered. You decide whether the answer is good, and increasingly you have to write down why.
That splits into a few recurring shapes:
- Rating. Score a single response against stated criteria.
- Ranking. Given two or more responses, say which is better and on what grounds.
- Error finding. Read a plausible-looking answer and locate the specific thing that is wrong with it.
- Rubric writing. Define the criteria other people will grade against.
Error finding is where the pay starts climbing, because it needs someone who would notice. micro1's STEM Expert listing describes it precisely: you identify "incorrect assumptions, reasoning errors, missing constraints, and subtle technical mistakes". Those are the failures where an answer looks right and is wrong for a structural reason, which is exactly what a non-specialist cannot catch.
Evaluation is why so many of these listings ask for a credential rather than experience with AI: the platform is paying for your ability to know that something is wrong.
Expert work: your job becomes the training data
The newest shape, and on this board the best paid, inverts the whole thing. Instead of reacting to model output, you produce the professional work you would produce anyway, and the platform captures it.
Mercor's UK data engineering and strategy consulting projects state it plainly. You are asked to "produce the kind of professional documents and deliverables you create day to day in your own field": working pipelines, specs, decks, operating-model maps. Those outputs are then logged and used to build the tasks and rubrics that train and evaluate frontier models.
micro1's Corporate Counsel and Attorney roles work the same way from the other direction: you run simulated contract negotiations and redlining exercises, then write the reasoning behind every markup. The written reasoning is what the platform is buying.
This is why a listing can offer $245 to $280 an hour for a management consultant or a senior journalist. The rate is buying a kind of judgement that is scarce and hard to synthesise, and it is buying the written explanation of that judgement, which is the part your employer never asked you for.
Where your task sits in the pipeline
Each of the three jobs feeds a different stage of how a model gets built. The stages overlap and platforms name them differently, but the order is roughly this:
- Data collection and preparation. Someone gathers the raw material (text, audio, images, video, code) and cleans it up: removing duplicates, stripping out personal details, sorting it into a format a training run can use. Transcription and simple classification tasks often sit here.
- Annotation and labeling. The material gets its structure. This is the data labeling described above, where a written guideline says what the correct answer is.
- Model training and iteration. Engineers train the model on that data and adjust it over repeated runs. Contractors almost never do this part themselves. Expert work feeds it from the side: the deliverables you produce become the tasks and rubrics a model is trained and tested against.
- Evaluation and human feedback. The model answers, and people judge the answers by rating them, ranking them or finding the error. This is the evaluation work above. Preference judgements made at this stage are what RLHF is built on.
Knowing the stage tells you what a task is for. A labeling guideline is asking for consistency; an evaluation rubric is asking you to notice what is wrong.
What this board is actually offering
Of the 32 listings live here on 8 September 2026:
- All 32 are remote contractor work, not employment. Every one is advertised as contract, and every one is remote.
- All 32 advertise an hourly rate, ranging from a $45 floor to a $400 ceiling. Whether you are paid by the hour is a separate question: see why these jobs advertise an hourly rate but pay per task.
- 22 are pitched at senior level, 9 at mid level, and 1 at entry level.
- Every rate is the platform's own published figure. None is our estimate.
The domain spread says the rest: 7 software engineering, 6 healthcare, 5 legal, 5 psychology, 3 consulting, 3 science, and single listings across design, product, audio, video, writing, marketing and transcription.
The board is overwhelmingly evaluation and expert work, aimed at people who already have a profession. If you are looking for entry-level annotation, this board is a poor match for you today, and we would rather say so than waste your evening.
How a task is usually structured
Whatever the shape, the unit of work is a task, and the size varies enormously.
micro1's Corporate Counsel listing is unusually specific: a task averages three and a half hours. That is half a working day per unit. At the other end, an annotation task may take seconds.
This matters because pay is frequently per task rather than per hour, even when the ad quotes an hourly band. micro1 states outright on several listings that pay is output-based, per task meeting the project specification, and that the time a task takes varies with the expert's experience and workflow. The hourly figure describes what a fast contributor reaches, not a rate you are owed for time spent.
What it is not
Worth stating, because the search results are full of it:
- It is not a salaried job. Every listing here is contractor work.
- It is not guaranteed volume. Passing a platform's assessment is not the same as having work: micro1 states that "certification does not guarantee immediate placement".
- It is not passive income. The work is genuinely difficult, and the reason the rates are high is that the easy version has largely been automated or moved to lower-cost markets.
- It is never something you pay for. See do I have to pay anything to get AI training work.
What we could not establish
- No platform here publishes what share of its work is labeling versus evaluation versus expert work. The split above is read off the listings we have collected, not off any platform's own breakdown.
- The boundary between "evaluation" and "expert work" is ours, not an industry standard. Platforms do not use consistent names for either.
- We do not know how long the current premium on expert work lasts. It exists because the supply of people who can do it is small, and that is not a permanent condition.
Sources
- The 32 listings live on this site on 8 September 2026, for every count, rate and quotation above
- The expert-facing pages of micro1, Mercor, Alignerr, Outlier and Ethos
Last updated 8 September 2026. Counts move as listings open and close. We are not affiliated with any of these platforms and we do not process applications.
Questions
- What is the difference between data labeling and AI evaluation?
- Labeling adds structure to raw material against a written guideline: tagging an image, transcribing audio, marking a sentence. There is a correct answer and you apply it consistently at volume. Evaluation starts after a model has produced something, and you judge it: rating it, ranking it against alternatives, or finding the specific error in an answer that looks plausible. Labeling rewards speed and consistency; evaluation rewards knowing enough to notice what is wrong.
- Do I need experience with AI to do this work?
- On most listings, no. Platforms are generally buying your existing professional judgement rather than any familiarity with machine learning, and several listings say so directly. micro1 states on role after role that no prior experience in AI is required and that your domain knowledge is what matters. The exceptions are the engineering roles that build evaluation infrastructure, where prior model evaluation work is sometimes asked for.
- Is AI training work entry level?
- Not on this board. Of the 32 listings live on 8 September 2026, 22 are pitched at senior level, 9 at mid level and 1 at entry level. The market has moved towards paying for scarce professional judgement rather than for annotation volume. Entry-level annotation work does exist elsewhere, but it is not what these platforms are mostly recruiting for.
- Why do some AI training jobs pay $45 an hour and others $400?
- Because they are different jobs. The floor is annotation and straightforward evaluation, where the platform is buying careful attention. The ceiling is expert work, where it is buying judgement that is scarce and the written reasoning behind it: on this board the top rates go to management consultants, senior journalists, BigLaw-experienced attorneys and senior software engineers. Every rate quoted here is the platform's own published figure.
- What is a task, and how long does one take?
- A task is the unit of work, and the size varies enormously by role. micro1's Corporate Counsel listing states an average of three and a half hours per task, which is half a working day. An annotation task may take seconds. This matters because pay is often per task rather than per hour even when the ad quotes an hourly band, so the size of a task decides what the advertised rate is actually worth to you.
Platforms covered here
Put this into practice
Every listing shows its pay and who it is open to.