Ranking and comparison tasks in AI training
Overview · 1 week ago
Side by side comparison tasks ask you to pick the better of two or more AI outputs and say why. What one task involves, the 24 listings on this site that describe it, and what they pay.
What a side by side comparison task is
A side by side comparison task shows you two or more AI outputs for the same prompt and asks which is better, and why. It is the most common form of preference data, the human judgement that techniques such as RLHF learn from. The broader picture of evaluation work is in what AI training work actually is; this page is about the comparison task specifically.
What you actually do
A single task usually runs like this:
- Read the prompt and every response in full. The differences are often small and buried halfway down.
- Judge each against the criteria. Listings name them: clarity, tone, helpfulness and instruction-following on micro1's Educator / Assessor / Writer role; musicality, creativity, prompt adherence, vocal quality and mix on Mercor's music projects.
- Pick the preferred output, sometimes with a strength rating, sometimes a full ranking of several.
- Write the reason. micro1's Excel Specialist listing gives the flavour: broken formulas hidden under correct-looking totals, or a deck whose charts do not match its tables. A preference without a sharp reason is of little use.
The format stretches well beyond chat text. Mercor's Music Production Expert (Russian) asks for pairwise judgements between two AI-generated songs. Invisible's Danish and Dutch speaker projects compare speech model samples on accuracy, fluency and naturalness. Invisible's AI Personalization Evaluation Expert project compares two assistants on how well each used your own context, and warns that the useful skill is noticing concrete differences, not just preferring one tone.
Who hires for it
By our reading of the task descriptions, 24 of the 699 listings on this site on 29 September 2026 describe comparing or ranking outputs as part of the job. Mercor posted 11, micro1 9, Invisible Technologies 3 and Alignerr 1 (where preference labelling is named as useful prior experience). Listings that describe the work only as "evaluation" may include comparisons too, so treat that number as a floor.
The subjects vary: office documents, legal and educational writing, product specs, accounting and insurance answers, materials science reasoning, music, lyrics and speech.
What it pays
22 of the 24 quote an hourly rate, from $14 to $140 an hour. The other two, Invisible's Danish and Dutch projects, pay $1.30 per unit.
- Generalist: micro1's Senior AI Trainer at $14–36, Invisible's personalization project at a flat $17, Mercor's Punjabi and Thai lyrics roles at $15 and $18.
- Office and business documents: micro1's Excel and PowerPoint specialists at $20–60, its business document expert at $35–50, Mercor's generalist experts at $50–70.
- Professional fields: $75–84 on Mercor's insurance, accounting and materials science roles; $90–140 on micro1's lawyer, educator, writer and product manager roles.
The rate depends on the subject matter and how much you know about it.
Doing it well
Decide on the criteria before you read the responses, or the first one you read anchors you. Check the longer answer for padding rather than rewarding it for length. Where two outputs are genuinely close, say so if the task allows it. Per-task pay, where it applies, is covered in our payout guide.
Current listings sit mostly under AI Safety & Evaluation, with the music and speech projects under Audio & Voice.
Questions
- What is a side by side comparison task?
- You see two or more AI outputs for the same prompt, judge each against the stated criteria, pick the preferred one or rank them, and write the reason for your choice. The outputs can be text, documents, songs, lyrics or speech samples.
- How much do comparison and ranking tasks pay?
- Of 24 listings on this site on 29 September 2026 that describe comparing or ranking outputs, 22 quoted hourly rates from $14 to $140, and two Invisible Technologies speech projects paid $1.30 per unit. The subject matter sets the rate.
- Why do comparison tasks ask for a written reason?
- Because a bare preference is of little use for training. Listings ask you to name the concrete difference, such as a broken formula under a correct-looking total, so the judgement can be checked and learned from.
Platforms covered here
Put this into practice
Every listing shows its pay and who it is open to.