Paid pilot for US epidemiologists: size patient populations for drugs and indications (prevalence, incidence, diagnosed to treated to addressable), judge whether estimates are sound and well sourced, and write rubrics for AI drug analysis. 5+ years and a graduate degree, 10–20 hours. $150–175/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
Language
Field
- Languages & Linguistics119 jobs
- Audio & Voice114 jobs
- Engineering109 jobs
- Business & Finance107 jobs
- Software & IT88 jobs
- Health & Medicine78 jobs
- Law, Policy & Security70 jobs
- General & Data Collection68 jobs
- Science & Math65 jobs
- AI Safety & Evaluation63 jobs
- Video, Image & Design44 jobs
- Data, AI & ML35 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
716 open roles · page 23 of 30
- $150 – $175 / HourWorldwide
Coordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesAn expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.
$80 – $150 / TaskWorldwideGrade AI-generated slides, spreadsheets and documents for real-world software engineering quality, flagging factual, visual and presentation errors with written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideNative German speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models behave safely in German. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; Germany or Western Europe preferred, not required; 8 hires this month.
$48 – $52 / HourWorldwideNative French speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models behave safely in French. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; France or Western Europe preferred, not required; 7 hires this month.
$48 – $52 / HourWorldwideArabic-speaking PhD chemists and biologists write specialised science prompts in Arabic and grade AI answers for accuracy and dual-use safety. Part-time remote at $38–42/hour; Saudi Arabia or MENA preferred, not required. 13 hires this month, the most active listing in the series.
$38 – $42 / HourWorldwideEvaluate vulnerability-reproduction and remediation tasks for a frontier AI lab: faithful CVE reproductions in Docker labs, sound fixes, and two-part verification (functionality plus vulnerability tests). For US AppSec engineers, pentesters and vulnerability researchers with 3+ years. $70–90/hour.
$70 – $90 / HourOpen to United StatesEvaluate AI-written Punjabi song lyrics for a leading AI lab: check them against published songs for similarity, rate quality, creativity, prompt adherence and originality, and judge whether slang and regional expressions ring true. For Punjabi songwriters, performers or music writers. Remote, flexible, up to 6 months, $15/hour.
$15 / HourWorldwideAssess AI-generated Norwegian song lyrics for a leading AI lab: compare with published songs for similarity, rate quality, creativity, prompt adherence and originality, and judge whether word choice and dialect sound natural. For Norwegian songwriters, performers or music journalists. Remote, flexible, up to 6 months, $42–78/hour.
$42 – $78 / HourWorldwideAuthor AI evaluation tasks from real drawing sets, documents and site photos, with the correct RFI response, coordination comments or markup as the answer. For licensed architects, project architects and job captains with 3+ years. US only, $45–60/hour.
$45 – $60 / HourOpen to United StatesWrite original, research-frontier questions in your own discipline that current AI models cannot answer, with sourced and cited reference answers, then test and harden them. Open to any field. Paid per accepted task, $40–90/hour band, five openings, remote contractor.
$40 – $90 / HourWorldwideNative or near-native German editors, critics and linguists evaluate LLM-generated German text, rewrite it to literary and cultural standards, annotate linguistic features and give feedback to researchers at a leading AI lab. Remote hourly contract at $50/hour.
$50 / HourWorldwideEvaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesRemote hourly contract for Python developers who already run PyCharm on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You supply the licence, the Mac and a display above 2.5 megapixels. Tasks are not described in the ad.
$55 – $65 / HourWorldwideNative Thai speakers use a test build of an AI voice assistant on their own Apple Silicon Mac to do everyday tasks by voice, then transcribe their queries, rate the responses and log the data. Paid $12.6 per task, 4-week contract starting immediately. Remote.
$12.60 / TaskWorldwideUmbrella listing for Mercor's energetic materials red-team panel: chemists and engineers or operators write benign, dual-use and adversarial prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 16 hired this month.
$65 – $75 / TaskWorldwideUmbrella listing for Mercor's radiological safety red-team panel: RSOs, health physicists, source security, emergency response and nuclear medicine specialists write prompts, judge frontier AI replies against a policy standard, and write reference answers. Remote contract at $65–75 per task; 8 hired this month.
$65 – $75 / TaskWorldwideNative or near-native Spanish translators, editors and annotators evaluate LLM-generated Spanish, correct or rewrite it, annotate grammatical and semantic features, and flag model errors for researchers at a leading AI lab. Remote hourly contract at $50/hour.
$50 / HourWorldwideEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesCertified pharmacy technicians answer medication questions and review AI responses for an AI lab building prior authorization workflows: dosing, interactions, indications, PA requirements. US only, 30–40 hours a week during the project. A flat $35/hour.
$35 / HourOpen to United StatesRF, microwave, antenna and electromagnetics engineers solve and critique hard technical problems for a short-term expert evaluation project: analysing systems and trade-offs, checking calculations and assumptions, and writing rigorous explanations. Hands-on industry experience and an EE-family degree. Remote hourly contract at $80–95/hour.
$80 – $95 / HourWorldwideBuild hard codebase-exploration tasks for an RL environment that trains AI agents: rewrite engineering questions about production Go repos (etcd, Traefik, Helm, gRPC-Go, Temporal and more) so frontier agents fail them, tune rubrics and foils, and pass a validation loop. $130 per approved task, remote.
$130 / TaskWorldwideTranscribe Italian audio, and English when required, edit transcripts for completeness, annotate data for AI training and flag unclear or low-quality recordings. Native Italian and professional transcription experience. 15 openings, contractor, $20–36/hour.
$20 – $36 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.