ML Challenge Task Auditor
- Pay
- $70 – $90 / Hour
- Open to
United States
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- machine learning
- experiment design
- model evaluation
- pytorch
- scikit-learn
- data leakage detection
What you'll do
Think of a Kaggle competition where you are not competing but checking the competition itself. A frontier AI lab uses applied ML tasks to train and test its models, and you audit those tasks for methodological rigor.
The questions you ask of each one:
- Is the experiment design sound, and is the model-selection reasoning defensible?
- Is evaluation methodology correct: the right metric, honest train/test/CV splits, no tuning on the test set?
- Is there data leakage, or a way to game the metric without learning anything?
- Do the claims hold up against the evidence, and can the results be reproduced?
Output is rubric-based written feedback. The ad adds a pointed clarification: this is about applied and experimental ML rigor, not building LLM applications and not MLOps. If your ML experience is mostly prompt pipelines or deployment infrastructure, this is not the listing.
Who fits
Basic qualifications:
- 3+ years of hands-on applied or experimental ML: experiment design, model selection, hyperparameter tuning, evaluation methodology
- A strong grasp of data-quality rigor, including leakage detection, metric gaming and split hygiene
- Proficiency with PyTorch, TensorFlow, scikit-learn or XGBoost
- The ability to critique ML claims against evidence and reproduce results
Preferred: Kaggle or other competition and benchmark experience, a graduate research or publication record in applied ML, and prior task grading or peer review. A seasoned conference reviewer already does most of this job.
What it pays
$70–90 per hour, weekly via Stripe or Wise. H-1B and STEM OPT candidates cannot be supported. Hours and project length are not published.
Among the $70–90 auditor listings Mercor posted together, this is the one for data scientists and ML researchers; engineers should compare the SWE-Bench Task Auditor.
Worth knowing
Good:
- Rewards the unglamorous rigor skills (leakage hunting, split hygiene) that are rarely paid for directly
- Kaggle track records count
- Asynchronous, remote
Less good:
- US-only, with the visa exclusion
- Reproducing results can take real compute and time; the ad does not say what environment is provided
- No hours or duration published
- Independent contractor; projects can be extended, shortened or ended early
About this listing
Posted by Mercor as a remote hourly contract for US residents, confirmed open on 24 September 2026. Qualifications, the scope note and pay are the ad's own. See Mercor.
More roles at Mercor
See all 304Similar roles at other platforms
Guides about Mercor
See all 52- How much does Mercor pay? Rates from 304 live listingsPay breakdown
- The Mercor AI interview: what happens, what it checks, and what comes afterInterview prep
- Mercor vs Alignerr: which to apply to, and how they differPlatform comparison
- Mercor, micro1, Outlier, Alignerr and Handshake AI compared: which to apply to firstPlatform comparison
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.