Skip to content
Labeling Jobs

SWE-Bench Task Auditor

Pay
$70 – $90 / Hour
Open to
  • United States
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • python
  • open source
  • code review
  • docker
  • test design
  • benchmark auditing

What you'll do

SWE-Bench-style benchmarks give a model a real repository, a real issue and a hidden test suite, and check whether the model's patch makes the tests pass. The benchmark is only as good as each task, and this role is quality control for those tasks.

You evaluate:

  • Reference patches: is the "golden" fix actually correct and minimal?
  • Test harnesses and runners: do the tests verify the fix, or would a wrong patch also pass?
  • Isolation: does the Docker environment reproduce the repository state cleanly?
  • Grading integrity: could a model cheat, through answer leakage in the prompt or environment, or by gaming the tests (reward hacking)?

Then you write rubric-based feedback. The adversarial mindset is the core skill: you look at each task and ask how a clever model could score without solving it.

Who fits

Basic qualifications:

  • 3+ years of professional software engineering
  • Genuine open-source contribution or maintainer experience: merged PRs, committer or maintainer roles
  • The ability to audit reference patches, test runners and Docker isolation, and to detect leakage and reward hacking
  • Python plus at least one of Java, Go, TypeScript or C++

Preferred: familiarity with SWE-Bench (Verified) or similar repository benchmarks, maintainer history on major Python projects (Django, Flask, scikit-learn, sympy, pytest are named), and prior code review or task grading.

The open-source requirement is the gate. A strong engineer whose work has all been in private codebases will struggle to show it here.

What it pays

$70–90 per hour, weekly via Stripe or Wise. H-1B and STEM OPT candidates cannot be supported. Hours and duration are not published.

It shares its rate with the other auditor listings Mercor posted alongside it, including the AI Developer Trace Task Auditor, which asks for agent-tool experience instead of open-source history.

Worth knowing

Good:

  • Direct work on one of the most cited coding benchmarks' task format
  • Maintainers of the named Python projects have an obvious edge
  • Asynchronous, remote

Less good:

  • US-only, with the visa exclusion
  • The public open-source track record requirement excludes many experienced engineers
  • The ad does not say how many tasks or hours are available
  • Independent contractor; projects can be extended, shortened or ended early

About this listing

Posted by Mercor as a remote hourly contract for US residents, confirmed open on 24 September 2026. Qualifications and pay are the ad's own. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.