Umbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.
Remote AI training and data labeling jobs
Filter jobs
Location
Language: English
Field
- Languages & Linguistics78 jobs
- Audio & Voice74 jobs
- Engineering107 jobs
- Business & Finance104 jobs
- Software & IT76 jobs
- Health & Medicine74 jobs
- Law, Policy & Security70 jobs
- General & Data Collection53 jobs
- Science & Math63 jobs
- AI Safety & Evaluation63 jobs
- Video, Image & Design37 jobs
- Data, AI & ML34 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
631 open roles matching these filters · page 22 of 27
- $65 – $75 / TaskWorldwide
Umbrella listing for Mercor's radiological safety red-team panel: RSOs, health physicists, source security, emergency response and nuclear medicine specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 8 hired this month.
$65 – $75 / TaskWorldwideEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesCertified pharmacy technicians answer medication questions and review AI responses for an AI lab building prior authorization workflows: dosing, interactions, indications, PA requirements. US only, 30–40 hours a week during the project. A flat $35/hour.
$35 / HourOpen to United StatesRF, microwave, antenna and electromagnetics engineers solve and critique hard technical problems for a short-term expert evaluation project: analysing systems and trade-offs, checking calculations and assumptions, and writing rigorous explanations. Hands-on industry experience and an EE-family degree. Remote hourly contract at $80–95/hour.
$80 – $95 / HourWorldwideBuild hard codebase-exploration tasks for an RL environment that trains AI agents: rewrite engineering questions about production Go repos (etcd, Traefik, Helm, gRPC-Go, Temporal and more) so frontier agents fail them, tune rubrics and foils, and pass a validation loop. $130 per approved task, remote.
$130 / TaskWorldwideTranscribe Italian audio, and English when required, edit transcripts for completeness, annotate data for AI training and flag unclear or low-quality recordings. Native Italian and professional transcription experience. 15 openings, contractor, $20–36/hour.
$20 – $36 / HourWorldwideQA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.
$70 – $110 / HourWorldwideWrite or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.
$94 – $119 / HourWorldwideWrite or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.
$61 – $77 / HourWorldwideAuthor executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.
$70 / HourWorldwideSenior materials scientists, and electrical or mechanical engineers, author realistic tasks with a prompt, a data room and a grading method, run them against an AI model and tighten them until the model can no longer reason through cleanly. Daily onboarding and office hours. Remote hourly contract at $60–90/hour.
$60 – $90 / HourWorldwideComplete self-contained fund-ops exercises from mock ledgers, statements and notices: reconciliations, NAV variance attribution, trade-break investigations, corporate actions and payment exceptions, each graded against a rubric. 3+ years in fund admin or investment ops. US only, about 15 hours a week, $75–110/hour.
$75 – $110 / HourOpen to United StatesFull-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesRemote hourly contract for developers on any stack who already use Visual Studio Code on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You need your own Mac and a display above 2.5 megapixels. Tasks are not described in the ad.
$55 – $65 / HourWorldwideGrade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideBuild a real estate deal model in Excel from a synthetic document pack, with live formulas and sourced assumptions, then score two other contractors' models against fixed criteria. For associates at real-estate-focused mid-market PE funds with ~2 years of IB first. US only, about 20–25 hours over 1.5–2 weeks, $80–100/hour.
$80 – $100 / HourOpen to United StatesFull-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.
$110 – $150 / HourHybridOpen to United StatesTurn real lab and test-bench experience into AI benchmark tasks: build test logs, calibration records and failure reports, define the right diagnosis, and write 35+ point rubrics. Remote contractor, $30–70/hour paid per accepted task, 50 openings, 4+ years in test or production.
$30 – $70 / HourWorldwideAnnotate and score robotics video footage against a scoring guide: label actions, objects and events, rate clips for quality and relevance, and QA your own annotations. Mid-level annotation experience preferred, robotics footage experience heavily preferred. Flat $7/hour, 20 openings, remote contract.
$7 / HourWorldwideTag, segment and mark keyframes in video using Final Cut Pro on macOS to build AI training datasets. Professional or academic Final Cut experience counts; aimed at entry to mid-level editors. One opening, contractor, $15–80/hour.
$15 – $80 / HourWorldwideMechanical designers build and validate parametric FreeCAD models of parts and assemblies, applying GD&T, to create a CAD dataset for AI training. Optional Python scripting. Remote contractor role, 50 openings, $40–100/hour.
$40 – $100 / HourWorldwideTurn everyday insurance judgment into AI training data: design realistic underwriting, claims, actuarial and compliance scenarios, review AI outputs and write feedback. Open to every insurance specialty, 3+ years, US-based. $75/hour.
$75 / HourOpen to United StatesJudge Vietnamese audio clips, from human speakers and AI speech models, for tones, pronunciation and naturalness, and back each rating with a written English explanation. Native Vietnamese and B2+ English preferred. 100 openings, contractor, $30–65/hour.
$30 – $65 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.