Work inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide230 jobs
- United States63 jobs
- United Kingdom7 jobs
- Canada5 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language: English
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering55 jobs
- Business & Finance45 jobs
- Software & IT38 jobs
- Health & Medicine60 jobs
- Law, Policy & Security51 jobs
- General & Data Collection2 jobs
- Science & Math46 jobs
- AI Safety & Evaluation22 jobs
- Video, Image & Design6 jobs
- Data, AI & ML17 jobs
- Writing & Education4 jobs
- Other fields5 jobs
Newest
294 open roles matching these filters · page 10 of 13
- $60 – $100 / HourHybridOpen to United States
Design marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesBuild retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesSenior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.
$90 – $110 / HourOpen to United StatesEvaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.
$70 – $90 / HourOpen to United StatesPaid pilot for US market access and pricing professionals: set and pressure-test gross-to-net and formulary-tier assumptions by drug class, judge whether coverage and rebating assumptions match real payer behaviour, and write rubrics for AI drug analysis. 5+ years, 10–20 hours. $175–200/hour.
$175 – $200 / HourWorldwideDesign adversarial prompts, find jailbreaks and policy failures, and document vulnerabilities in frontier AI models across cyber, biosecurity, fraud, misinformation and political content. Hourly remote contract at $70–84/hour for residents of Europe, the UK and the US; 5+ years' relevant experience required.
$70 – $84 / HourOpen to Albania, Austria and 38 more countriesFull-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesGrade AI-generated slides, spreadsheets and documents for real-world data science quality, flagging factual, visual and presentation errors in structured written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwidePaid pilot for US epidemiologists: size patient populations for drugs and indications (prevalence, incidence, diagnosed to treated to addressable), judge whether estimates are sound and well sourced, and write rubrics for AI drug analysis. 5+ years and a graduate degree, 10–20 hours. $150–175/hour.
$150 – $175 / HourWorldwideCoordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesAn expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.
$80 – $150 / TaskWorldwideGrade AI-generated slides, spreadsheets and documents for real-world software engineering quality, flagging factual, visual and presentation errors with written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideWrite original, research-frontier questions in your own discipline that current AI models cannot answer, with sourced and cited reference answers, then test and harden them. Open to any field. Paid per accepted task, $40–90/hour band, five openings, remote contractor.
$40 – $90 / HourWorldwideEvaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesUmbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's radiological safety red-team panel: RSOs, health physicists, source security, emergency response and nuclear medicine specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 8 hired this month.
$65 – $75 / TaskWorldwideEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesRF, microwave, antenna and electromagnetics engineers solve and critique hard technical problems for a short-term expert evaluation project: analysing systems and trade-offs, checking calculations and assumptions, and writing rigorous explanations. Hands-on industry experience and an EE-family degree. Remote hourly contract at $80–95/hour.
$80 – $95 / HourWorldwideQA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.
$70 – $110 / HourWorldwideWrite or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.
$94 – $119 / HourWorldwideWrite or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.
$61 – $77 / HourWorldwideAuthor executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.
$70 / HourWorldwideSenior materials scientists, and electrical or mechanical engineers, author realistic tasks with a prompt, a data room and a grading method, run them against an AI model and tighten them until the model can no longer reason through cleanly. Daily onboarding and office hours. Remote hourly contract at $60–90/hour.
$60 – $90 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.