Research Evaluation Specialist (PhD / Researcher / Professor)
- Pay
- $40 – $90 / Hour
- Open to
- Worldwide
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- research
- source triangulation
- question writing
- benchmark design
- analytical thinking
- written precision
What you'll do
This is "hard question" writing: the kind of work behind expert-level AI benchmarks. You write original, high-difficulty question-and-answer pairs at the cutting edge of your own discipline, then try them against AI systems and make them harder when the model gets them right.
The cycle, as the ad describes it:
- Write a question that requires advanced reasoning, methodological nuance or synthesis across several sources, designed so it cannot be shortcut
- Source the answer from primary literature and authoritative references, with citations and reasoning
- Test it against AI systems, identify questions that are not challenging enough, and increase the difficulty without losing accuracy
- Polish the wording until the question has one defensible answer
- Revise based on reviewer feedback and project guidelines
The hard part is usually making a question difficult and unambiguous. A question that stumps the model only because it is vague will be rejected.
Who fits
- A completed PhD, active candidacy, or equivalent research experience as a specialist, researcher or professor
- Real depth in your field and familiarity with its primary literature
- Strong analytical thinking, attention to detail and written English
- Experience triangulating information across authoritative sources
- Self-direction: you will work alone and be judged on output
- Skill at writing challenging, original, methodologically sound questions
- AI training or evaluation experience helps but is not required
No discipline is specified, so this is open to researchers in any field, sciences or humanities.
What it pays
The band is $40–90/hour, but pay is output-based: per task that meets the project specification, with time per task depending on your experience and workflow, and a minimum number of tasks per week (the number is not stated).
For a PhD-level role, the floor is low. The deciding variable is how fast you can produce a question that survives both the model and the reviewer; the first few will take much longer than later ones.
Start
micro1 says roles like this are typically filled within 48 hours, with first tasks expected within 24–48 hours of onboarding.
Worth knowing
Good:
- Open to any academic discipline
- PhD candidates and non-PhD researchers with equivalent experience qualify
- Genuinely intellectual work that uses your literature knowledge
- Pay mechanics are published
- No country restriction stated
Less good:
- A $40 floor for PhD-level work, and per-task pay on top
- Questions the model can answer, or reviewers reject, are effectively unpaid time
- An unpublished weekly minimum
- Only five openings
- Contractor engagement: no benefits, no notice period, your own tax to manage
About this listing
Posted by micro1 on its own jobs portal and read on 24 September 2026. The band, the five openings, the qualifications, the output-based pay terms and the start timeline are from the listing. See micro1.
More roles at micro1
See all 380Similar roles at other platforms
Guides about micro1
See all 81- micro1 AI Research Scientist Interview: Model Evaluation, Tuning and Deep Learning at ScaleInterview prep
- micro1 Clinical Research Scientist Interview: Trial Compliance, Protocol Amendments and DataInterview prep
- micro1 Academic Advisor Interview: Academic Planning, Conflict Resolution and Advising DataInterview prep
- micro1 SEO Content Writer Interview: Keyword Research, Search Intent and Algorithm UpdatesInterview prep
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.