Evaluation Specialist/Recent Grad
- Pay
- $20 – $60 / Hour
- Open to
- Worldwide
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- research
- source triangulation
- question writing
- analytical thinking
- ai evaluation
What you'll do
You build hard questions. Each task is an original question-and-answer pair designed to test the limits of an advanced AI model, on any subject you can research well. The ad is specific about what "hard" means: multi-step questions that require synthesis and reasoning across sources, not something a single search result answers.
The loop goes like this. You draft a question, research the answer thoroughly and confirm it against multiple independent sources (the ad calls this source triangulation), then run the question against AI models. If the model gets it right, you make it harder; if the question turns out ambiguous, you tighten it. Every answer is documented with citations and a clear line of reasoning, and reviewers send feedback that you fold back in.
This is the benchmark-writing end of AI training, and it is closer to research than to annotation. People who enjoyed writing a thesis, running a quiz team or fact-checking will recognise the satisfactions and the frustrations.
Who fits
All listed as preferred:
- Thorough, independent research and critical evaluation of sources
- Precise written English (fluency required; it need not be your first language)
- A talent for well-structured, probing questions
- Self-direction and reliability on remote, independent work
- Genuine interest in testing what current AI can and cannot do
- AI training or content-creation experience is a plus, not required
The ad explicitly welcomes recent graduates, advanced-degree holders and people from any academic or professional background with research and writing ability.
What it pays
$20–60/hour as the published band, but compensation is output-based: you are paid per task that meets the project specification, minimum submission requirements apply, and there is a weekly minimum number of tasks (not quantified). Questions that the model answers easily, or that fail review, are time spent without pay, so early rates are likely to sit near the floor until you learn what passes.
Worth knowing
Good:
- One of the few micro1 listings aimed squarely at recent graduates
- Any subject background qualifies
- Non-native English speakers are explicitly welcome
- Intellectually engaging work that builds research skill
Less good:
- Only 5 openings
- Per-task pay on a $20 floor, with an unquantified weekly minimum
- Stumping frontier models is getting harder, which means more drafts per accepted task
- Fast start expected: roles fill within about 48 hours and first tasks begin 24–48 hours after onboarding
- Contractor engagement with no benefits; payment schedule and method are not stated
About this listing
Posted by micro1 on its own jobs portal and read on 24 September 2026. The $20–60/hour band, the 5 openings, the per-task compensation terms and the open-background eligibility are micro1's published text. See micro1.
More roles at micro1
See all 380Similar roles at other platforms
Guides about micro1
See all 81- micro1 AI Research Scientist Interview: Model Evaluation, Tuning and Deep Learning at ScaleInterview prep
- micro1 Academic Advisor Interview: Academic Planning, Conflict Resolution and Advising DataInterview prep
- micro1 Generative AI Specialist Interview: Attention Mechanisms, Bias and Scaling in ProductionInterview prep
- micro1 Clinical Research Scientist Interview: Trial Compliance, Protocol Amendments and DataInterview prep
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.