AI Evaluation Specialist
micro1: pay and who it accepts · · Apply link checked
- Pay
- $30 – $90 / Hour
- Open to
United States
Canada
United Kingdom
Ireland
Australia
New Zealand
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- rubric-based evaluation
- quality assurance
- ai agents
- ai adoption
- process improvement
- written communication
What you'll do
This is rubric work at volume. You assess AI assistant outputs against detailed rubrics and defined quality standards, judging accuracy, relevance and whether the response followed the guidelines it was given, and you apply the same standard to the five hundredth example as to the first.
Beyond the score, you are looking for three specific things the ad names: reasoning gaps, tool-use failures and logic errors. Those are agent-era failure modes rather than writing problems, so the interesting cases are the ones where the assistant reached a defensible-looking answer through a broken path. You write up strengths as well as weaknesses, concisely, and maintain documentation that lets someone else trace how you arrived at a rating.
The work also goes beyond grading: contributors take part in deciding how ambiguous criteria should be read and how the standards evolve as the models do, which means the edge case you flag today can become the rule everyone applies next month.
What you need
- Experience in grading, quality assurance, editorial review, assessment, annotation or a similar field demanding careful analysis and detailed feedback
- Advanced, daily use of AI assistants such as ChatGPT or Claude as a real work tool
- The ability to synthesise complex information and communicate findings in writing
- A background in process improvement, rubric development or operational quality assessment, in an enterprise or educational setting
- Strong critical thinking, with consistency, integrity and fairness across evaluations
- Comfort working alone through large volumes of similar examples without the attention slipping
- A collaborative approach to ambiguous cases and evolving criteria
Eligibility is the US, Canada, UK, Ireland, Australia or New Zealand. Unlike micro1's AI Agent Power User listing on the same band, this one states no degree and no minimum years of experience.
What it pays
$30–90/hour, micro1's own published figure. The top of the band is three times the bottom with nothing published about what decides placement, which is a much wider spread than micro1's writing-evaluation roles carry.
Plan around $30 until you have an offer. No compensation structure is stated, so whether the rate buys your time or amounts to a per-task equivalent is unanswered, and on high-volume grading work that distinction matters more than usual: throughput becomes your effective rate.
Worth knowing
Good:
- No degree requirement and no stated minimum years, which makes this the more open of micro1's two $30–90 listings
- Teaching, marking, editorial review, QA and annotation backgrounds all qualify explicitly
- Daily AI assistant use is a credential here rather than a nice-to-have
- A real say in how rubrics get interpreted, not just applying them
- Six eligible countries
- Ten openings
Less good:
- A 3x spread with no published basis for placement, on a floor of $30
- High-volume repetitive grading is genuinely tiring, and the ad is honest that consistency across many similar examples is the job
- Effective hourly pay depends on throughput if the engagement turns out to be per task, which is not stated
- No project duration, weekly hours or start date stated
- Contractor engagement: no benefits, no notice period, your own tax to manage
- The enterprise customer is not named
About this listing
Posted by micro1 on its own jobs portal, confirmed open on 12 September 2026. The $30–90/hour band, the ten openings, the six-country eligibility list and every qualification above are its own published text for this listing.
micro1 runs a near-identical sibling listing, AI Agent Power User, on the same band, the same skills and the same countries. That one is the hands-on half, where you drive multi-step scenarios through connected tools; this one is the grading half. The Power User ad adds a bachelor's degree and a five-year experience floor that this one does not state, so if both appeal, this is the easier door.
No pay structure, project duration or weekly hours are published, and the customer is not named. Screening and payment run at the platform level. See micro1.
More roles at micro1
See all 380Similar roles at other platforms
Guides about micro1
See all 81- micro1 AI Research Scientist Interview: Model Evaluation, Tuning and Deep Learning at ScaleInterview prep
- micro1 Generative AI Specialist Interview: Attention Mechanisms, Bias and Scaling in ProductionInterview prep
- micro1 Finance Expert Interview: Accuracy, Risk Assessment and Personalised Advice LimitsInterview prep
- micro1 Accounting Manager Interview: Financial Reporting, Consolidation and Tax ComplianceInterview prep
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.