Senior Software Engineer: AI Evaluation & Benchmarks
Alignerr: pay and who it accepts · · Apply link checked
- Pay
- $80 – $100 / Hour
- Open to
- Worldwide
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- python
- llm benchmarking
- data pipelines
- code review
- unit testing
- git
What you'll do
You design the coding benchmarks used to measure frontier AI models against real programming tasks, and you build the data pipelines those evaluations run on.
That means constructing structured scenarios that test reasoning, debugging and code quality; analysing AI-generated code for correctness, reliability and edge-case failure; and writing detailed technical feedback on where and how a model breaks down. You work across large repositories and multiple languages. The contract runs three months, fully remote, with full-time availability preferred.
This is measurement work within AI training: the benchmarks you build are the yardstick other teams are held to.
What you need
- Four or more years of professional software engineering experience, stated as non-negotiable
- Experience at a high-growth tech company or a top-tier software organisation
- Expert Python: clean, performant, well-tested
- Hands-on work inside large, complex codebases and code repositories
- Proven experience designing and implementing LLM coding benchmarks and data pipelines
- Strong command of Git and modern development workflows
- Bilingual or native English, with strong written communication
Nice to have
- A senior or lead profile with a history of technical ownership
- A CS or ML degree, or equivalent professional experience
- A second language among JavaScript, Go or C++
- CI/CD experience and robust unit testing (pytest, Mocha, JUnit)
- Security engineering background or significant open-source contributions
- Familiarity with AI/ML evaluation methodology or model benchmarking
What it pays
$80–100/hour, stated on the listing. With full-time availability preferred, that works out at roughly $3,200 to $4,000 a week, or somewhere between $42,000 and $52,000 across the stated three months if the hours hold up.
The $80 floor is high for contract engineering work, which is consistent with how narrow the requirement is: prior experience designing LLM coding benchmarks is not a common line on a resume.
Worth knowing
Good:
- The rate is published up front, and the floor is high for contract engineering work
- A defined three-month scope, which is more certainty than most task-based AI work offers
- Benchmark design is portfolio-worthy work in a field where it is becoming a recognised specialism
- The ad states potential for extension as new evaluation challenges emerge
- Fully remote, with no country restriction stated
Less good:
- The prior-benchmark requirement is strict: "proven experience designing and implementing LLM coding benchmarks" rules out most senior engineers
- Pay is hourly with no stated minimum, so "full-time availability preferred" asks you to hold capacity nobody has committed to filling
- Three months is the stated length; extension is described as potential, not planned
- Hourly contract: no benefits, no notice period, your own tax to manage
About this listing
Posted by Alignerr on its own jobs board under the Coding category, confirmed open on 6 September 2026. The $80–100/hour rate is the ad's own published figure. The stated contract length is three months, but no application closing date is given, so none is set here, and no country restriction is stated either. Onboarding and payment run at the platform level. See Alignerr.
More roles at Alignerr
Similar roles at other platforms
- Software Engineer (Coding Tasks, Testing Focus)micro1 · $30 – $100 / Hour · 7d ago
- Lean Engineer, Formal Mathematics (Lean 4, Mathlib, Theorem Proving)Mercor · $90 – $110 / Hour · 10d ago
- Software & Firmware Engineering ExpertsMercor · $100 – $120 / Hour · 14d ago
- Open Source GitHub Maintainermicro1 · $150 – $300 / Hour · 15d ago
Guides about Alignerr
See all 21- Is Alignerr legit? Who runs it, what it pays and what workers reportSafety check
- Mercor vs Alignerr: which to apply to, and how they differPlatform comparison
- Alignerr onboarding: from sign-up to your first paid projectHow-to
- Mercor, micro1, Outlier, Alignerr and Handshake AI compared: which to apply to firstPlatform comparison
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.