A short, one-off evaluation task for mental health professionals: compare pairs of simulated clinical conversations and judge which is more realistic. The ad says it takes up to 30 minutes. Remote contract at $100/hour, for clinicians with a relevant degree and practical clinical experience.
Remote AI training and data labeling jobs
Filter jobs
Location: Worldwide
- Worldwide14 jobs, applied. Activate to remove
- United States14 jobs
- United Kingdomno roles alongside your other filters
- Canadano roles alongside your other filters
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English12 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japaneseno roles alongside your other filters
- Portugueseno roles alongside your other filters
Field: Health & Medicine
- Languages & Linguistics50 jobs
- Audio & Voice28 jobs
- Engineering31 jobs
- Business & Finance26 jobs
- Software & IT16 jobs
- Health & Medicine14 jobs, applied. Activate to remove
- Law, Policy & Security9 jobs
- General & Data Collection31 jobs
- Science & Math34 jobs
- AI Safety & Evaluation57 jobs
- Video, Image & Design6 jobs
- Data, AI & ML9 jobs
- Writing & Education1 job
- Other fieldsno roles alongside your other filters
Level
- Entryno roles alongside your other filters
- Juniorno roles alongside your other filters
- Medium2 jobs
- Senior12 jobs
Newest
14 open roles matching these filters
- $100 / HourWorldwide
Interpret dermatology cases from images and history, annotate lesions to a schema, grade AI assessments and write the criteria they are judged by. Non-clinical, shared expert pool, open worldwide. Board-certified or board-eligible, 3+ years post-residency, 15 hours a week minimum. A flat $270/hour.
$270 / HourWorldwideHealthcare back-office specialists (coding, prior authorization, denials and appeals, payer operations) design realistic operational scenarios, build the supporting records, and author and evaluate tasks that train AI agents. 2+ years and currently in role, 10 hours a week. $50–65/hour.
$50 – $65 / HourWorldwidePaid pilot for a biotech and pharma research team: create and critique rubrics that assess commercial drugs and development programs, judge investment-style theses, and assess 5–10 companies end to end. For specialist-fund biotech analysts with 5+ years. About 10–20 hours over 1–2 weeks, $120–200/hour.
$120 – $200 / HourWorldwidePaid pilot for board-certified, practising physicians with deep therapeutic-area expertise: interpret trial endpoints, judge whether results would change prescribing and real-world uptake, and write rubrics that evaluate AI analysis of drugs. 10–20 hours over 1–2 weeks. $150–230/hour.
$150 – $230 / HourWorldwideHands-on preclinical scientists from companies that develop their own drugs annotate R&D data and model outputs, and advise an AI lab building foundation models for drug discovery. Any modality, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwidePreclinical scientists with direct antibody-drug conjugate or bispecific antibody experience annotate R&D data and advise an AI lab building drug discovery foundation models. In-house asset developers only, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwideNuclear medicine physicists, radiopharmacy staff and cyclotron RSOs red-team frontier AI models: write benign, dual-use and adversarial prompts from medical isotope work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideBoard-certified radiologists label findings, write reference reports, grade AI-generated reads and author rubrics for medical imaging AI. Non-clinical, remote worldwide, hourly contract at $200–400/hour with a 15-hour weekly minimum. Weekly pay via Stripe or Wise.
$200 – $400 / HourWorldwidePaid pilot for US market access and pricing professionals: set and pressure-test gross-to-net and formulary-tier assumptions by drug class, judge whether coverage and rebating assumptions match real payer behaviour, and write rubrics for AI drug analysis. 5+ years, 10–20 hours. $175–200/hour.
$175 – $200 / HourWorldwidePaid pilot for US epidemiologists: size patient populations for drugs and indications (prevalence, incidence, diagnosed to treated to addressable), judge whether estimates are sound and well sourced, and write rubrics for AI drug analysis. 5+ years and a graduate degree, 10–20 hours. $150–175/hour.
$150 – $175 / HourWorldwideWrite or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.
$94 – $119 / HourWorldwideRead physician and patient research surveys and say whether the clinical terminology, treatment pathways and response options match how the disease is actually treated. A flat $200/hour, three years in a therapeutic area, and an MD is explicitly not required. Fully remote, on your own schedule.
$200 / HourWorldwideDesign and review physician and patient questionnaires (screeners, question wording, response scales, branching logic, respondent burden) and say whether an instrument will actually produce usable data. A flat $140/hour, and therapeutic-area specialism is welcome but not required. Fully remote, on your own schedule.
$140 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.