Audit applied machine-learning tasks used to train and evaluate a frontier AI lab's models: experiment design, model selection, evaluation methodology, leakage and metric gaming. For US practitioners with 3+ years of hands-on experimental ML in PyTorch, TensorFlow, scikit-learn or XGBoost. $70–90/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
Language: English
Field
- Languages & Linguistics78 jobs
- Audio & Voice74 jobs
- Engineering107 jobs
- Business & Finance104 jobs
- Software & IT76 jobs
- Health & Medicine74 jobs
- Law, Policy & Security70 jobs
- General & Data Collection53 jobs
- Science & Math63 jobs
- AI Safety & Evaluation63 jobs
- Video, Image & Design37 jobs
- Data, AI & ML34 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
631 open roles matching these filters · page 20 of 27
- $70 – $90 / HourOpen to United States
Red-team frontier AI models from the chemical defence side: write benign, dual-use and adversarial prompts drawn from countermeasures, protection and detection work, judge how models respond against a policy standard, and write the reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideDesign Excel tasks from cross-industry experience (data models, dashboards, Power Query, VBA automation), write the solutions, and evaluate AI attempts for a tech company's GenAI team. Breadth across functions, teaching or a PhD help. US only, W-2 via Cincinnatus LLC, 40 then 20 hours a week. $70–100/hour.
$70 – $100 / HourOpen to United StatesRemote hourly contract for game developers, technical artists and virtual production or archviz specialists who already run Unreal Engine on their own Windows PC. $30–40/hour, paid weekly via Stripe or Wise. The engine is free to download; a display above 2.5 megapixels is required.
$30 – $40 / HourWorldwideKorean-speaking PhD chemists and biologists write specialised science prompts in Korean and grade AI answers for accuracy and dual-use safety. Part-time remote at $63–67/hour, $15 above the Korean generalist role and third-highest in the series. East Asia preferred, not required.
$63 – $67 / HourWorldwideNative Korean speakers write prompts on sensitive subjects, classify conversations and flag adversarial phrasing so AI models stay safe in Korean. Business English and a bachelor's (in progress is fine). Part-time remote at $48–52/hour; South Korea or East Asia preferred, not required; 7 hires this month.
$48 – $52 / HourWorldwideCertified medical coders (CCDS, CHC, CCS or CPC) review clinical documentation and AI-generated notes for coding integrity, validate code-to-documentation alignment and annotate against guidelines. US-based, C1 in any second language, 10 hours a week. $45–65/hour.
$45 – $65 / HourOpen to United StatesSenior M&A, securities and governance lawyers design corporate law scenarios, reference memos and rubrics that test AI on deal work. US (Delaware, SEC) or international (UK Companies Act, EU) track. Remote hourly contract at $90–100/hour; 5+ years at a major firm, bank or large company.
$90 – $100 / HourWorldwideHindi-speaking PhD chemists and biologists write specialised science prompts in Hindi and grade AI answers for accuracy and safe handling of dual-use topics. Part-time remote at $23–27/hour; India or South Asia preferred, not required. 12 hires this month, among the busiest in the series.
$23 – $27 / HourWorldwideAudit Spanish (Spain) speech data for Amazon's Sonic project: check annotators' transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Spanish as spoken in Spain is a hard requirement; rationales are in English. Remote hourly contract at $39.50/hour.
$39.50 / HourWorldwidePortuguese-speaking PhD chemists and biologists write specialised science prompts in Portuguese and grade how AI models handle accuracy and dual-use safety. Part-time remote at $50–54/hour; Portugal or Western Europe preferred, not required. PhD candidates eligible; 9 hires this month.
$50 – $54 / HourWorldwideNamed RSOs and radiation protection managers on broad-scope, hospital, university or industrial licences red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideIndia-based full-stack engineers build real applications on top of a leading AI lab's pre-release models, wiring in tool interfaces, evaluation harnesses and telemetry, and writing up the model failures they hit. 3+ years at a top-tier organisation and two programming languages. Full-time, 40 hours a week, $25–30/hour.
$25 – $30 / HourOpen to IndiaPaid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.
$130 – $210 / HourOpen to United StatesEvaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesBoard-certified radiologists label findings, write reference reports, grade AI-generated reads and author rubrics for medical imaging AI. Non-clinical, remote worldwide, hourly contract at $200–400/hour with a 15-hour weekly minimum. Weekly pay via Stripe or Wise.
$200 – $400 / HourWorldwideWork inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesDesign marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesAI safety testing in Urdu and English: provoke jailbreaks, bias and harmful output from chat models and judge whether their Urdu answers are accurate and appropriate, in Nastaliq script or Roman Urdu. Evaluation judgment is the core ask. Remote hourly contract at $16–22/hour, weekly pay.
$16 – $22 / HourWorldwideBuild retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesEvaluate generative music AI for a leading AI lab: compare AI-made songs head to head on musicality, prompt adherence, vocals and mix, label genre and structure, and check lyrics and vocals. For Thai-speaking producers or engineers with 2+ years' experience. Remote, flexible hours, up to 6 months, $18/hour.
$18 / HourWorldwideRate AI-generated music for a leading AI lab: head-to-head song comparisons on musicality, prompt adherence, vocals and mix, genre and structure labelling, and lyric and vocal checks. For Russian-speaking producers and mix engineers with 2+ years' experience. Remote, flexible hours, up to 6 months, $35–49/hour.
$35 – $49 / HourWorldwideSenior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.
$90 – $110 / HourOpen to United StatesRemote hourly contract for video editors and content producers who already cut in Adobe Premiere on their own Windows PC. $30–40/hour, paid weekly via Stripe or Wise. Your own Adobe subscription and a display above 2.5 megapixels are required; the tasks are not described.
$30 – $40 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.