What is RLHF? Reinforcement learning from human feedback explained for workers
Overview · 1 week ago
What RLHF is, what a worker actually does in an RLHF task (comparing, rating and rewriting model answers), who hires for it and what the 26 listings on this site that mention it pay.
RLHF in one paragraph
RLHF, reinforcement learning from human feedback, is the training step where people judge a language model's answers and the model is tuned towards what they preferred. The method was set out in OpenAI's 2022 InstructGPT paper, Training language models to follow instructions with human feedback. Labelers there did two things: wrote demonstrations of the answer they wanted, and ranked several model outputs to the same prompt (between four and nine at a time). The rankings trained a reward model that predicts which output a person would prefer, and the language model was then optimised against that reward with a reinforcement learning algorithm called PPO.
For a worker, only the first half of that is your job. You never touch the reward model or the reinforcement learning. You produce the human judgements it learns from. For how this fits alongside labeling and expert work generally, see what AI training work actually is.
What one RLHF task looks like
Listings rarely call the job "RLHF". It shows up as a skill or a line in the description, and the task underneath is usually one of these:
- Compare and rank. Read two or more answers to the same prompt, pick the better one and write why. micro1's Educator / Assessor / Writer listing describes exactly this: several AI responses to one prompt, scored against a rubric, each with a written rationale.
- Rate against a rubric. Score one answer on accuracy, clarity and instruction-following, and flag the error.
- Write the ideal answer. The demonstration half of RLHF: you write what the model should have said.
- Write the rules. Someone has to write the guidelines raters follow. Mercor's AI Rater Guidelines Writer hires linguists and instructional designers to do that, with success measured by higher agreement between raters.
Reviewers check whether your written reason holds up.
Who hires for it
Of the 699 listings on this site on 29 September 2026, 26 mention RLHF or human feedback in the job description, and six name RLHF as a listed skill. micro1 accounts for 18 of the 26, Mercor for 7 and Alignerr for one. Most are not billed as RLHF jobs at all: they are domain expert roles in data science, software, law, finance and product management where RLHF experience is listed as helpful rather than required.
What it pays
The 26 listings are all hourly bands, running from $14 to $200 an hour. The spread tracks the professional field:
- Generalist grading at the bottom: micro1's Senior AI Trainer pays $14–36.
- Rubric and guideline work in the middle: Mercor's guidelines writer pays $45–65, its AI Safety Practitioner $60–70.
- Professional fields at the top: micro1's data science, consulting and software engineering domain expert roles each list $100–200.
Several micro1 listings state that pay is per task meeting the specification rather than per hour, so a slow start earns less than the band suggests. Our payout guide covers that in detail.
Where to look
Most of this work sits under the AI Safety & Evaluation filter, with the technical versions under Data, AI & ML. Search the listing text for "rubric", "rank" or "preference" rather than for "RLHF"; the work is described far more often than it is named.
Questions
- What is RLHF in simple terms?
- Reinforcement learning from human feedback is a training step where people judge a language model's answers, usually by ranking several answers to the same prompt, and the model is tuned towards what they preferred. In the InstructGPT method the rankings train a reward model, and the language model is then optimised against it.
- What does an RLHF worker actually do?
- You produce the human judgements, not the training. A task is usually comparing and ranking several answers to one prompt, rating one answer against a rubric, writing the ideal answer yourself, or writing the guidelines other raters follow. Each judgement comes with a written reason.
- How much do RLHF jobs pay?
- On this site on 29 September 2026, 26 listings mentioned RLHF or human feedback, all quoting hourly bands from $14 to $200. Generalist grading sits at the bottom and professional fields such as data science, consulting and software engineering at the top.
Platforms covered here
Put this into practice
Every listing shows its pay and who it is open to.