Probe AI models in Tamil and English for jailbreaks, bias and harmful output, and assess whether their Tamil answers are accurate and appropriate. Evaluation judgment is the core requirement, not prior red-teaming. Remote hourly contract, $16–22/hour, weekly pay; 57 hired this month.
Remote AI training and data labeling jobs
Filter jobs
Location: Worldwide
Language
Field
- Languages & Linguistics50 jobs
- Audio & Voice28 jobs
- Engineering31 jobs
- Business & Finance26 jobs
- Software & IT16 jobs
- Health & Medicine14 jobs
- Law, Policy & Security9 jobs
- General & Data Collection31 jobs
- Science & Math34 jobs
- AI Safety & Evaluation57 jobs
- Video, Image & Design6 jobs
- Data, AI & ML9 jobs
- Writing & Education1 job
- Other fieldsno roles alongside your other filters
Newest
197 open roles matching these filters · page 6 of 9
- $16 – $22 / HourWorldwide
Native Finnish speakers write prompts on sensitive topics, classify conversations and flag adversarial phrasing so AI models stay safe in Finnish. Needs business English and a bachelor's, finished or in progress. Part-time remote at $48–52/hour; Finland or Western Europe preferred, not required.
$48 – $52 / HourWorldwideRed-team AI models in Malay and English: jailbreaks, prompt injection, bias and multi-turn manipulation, captured as labelled, reproducible safety data. Native Malay required, along with prior adversarial, security or abuse-analysis experience. Remote hourly contract at $17–25/hour, paid weekly.
$17 – $25 / HourWorldwideWrite a Python simulation and a spec sheet with pass/fail thresholds; a frontier model probes your simulation a limited number of times, then submits a design that an agentic grader scores. For control, analog or RF circuit, power electronics or mechanical design experts with a PhD or equivalent industry record. Remote hourly contract, $60–90/hour.
$60 – $90 / HourWorldwideRemote hourly contract for lawyers, policy staff, consultants and technical writers who author long documents in Word on their own Mac. $60–70/hour, paid weekly via Stripe or Wise. Requires your own Microsoft licence and a screen above 2.5 megapixels. Tasks are not described.
$60 – $70 / HourWorldwideQuantum and computational chemistry PhDs write original, runnable research problems for a scientific-computing AI benchmark, with grading criteria, calibrated until frontier models fail them more often than they pass. 6 weeks at 20+ hours a week, Git and Docker workflow. Flat $70/hour; 182 hired this month.
$70 / HourWorldwideFuel-cycle and criticality safety engineers red-team frontier AI models: write benign, dual-use and adversarial prompts from enrichment, fabrication, reprocessing or criticality work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideHands-on preclinical scientists from companies that develop their own drugs annotate R&D data and model outputs, and advise an AI lab building foundation models for drug discovery. Any modality, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwideA paid research study for a major tech company building AI design tools: senior product, UX and brand designers take part in recorded remote sessions on how experts judge craft and decide what is ready to ship, either as interviewee or interviewer. About 4–5 hours per session. $150–250/hour, scope still being finalised.
$150 – $250 / HourWorldwideEvaluate AI-generated legal analyses of everyday civil problems (housing, family, consumer, foreclosure, debt) and write structured feedback on reasoning, procedure and access-to-justice issues. For US-admitted lawyers with pro bono program experience. Remote hourly contract at $170/hour.
$170 / HourWorldwideRemote hourly contract for consultants, analysts and marketers who build decks in PowerPoint on their own Windows PC. $60–70/hour, paid weekly via Stripe or Wise. Your own Microsoft licence and a screen setup above 2.5 megapixels are required; the ad does not describe the tasks.
$60 – $70 / HourWorldwideQA engineers, SDETs and test automation engineers review browser-based test workflows for AI-generated web apps, checking that each test is feasible, isolated, deterministic and gives a clean pass or fail before it joins a benchmark dataset. 3+ years and Playwright, Cypress or Selenium experience. Remote hourly contract at $30–60/hour.
$30 – $60 / HourWorldwideCurrent and former IAEA, Euratom and state-system safeguards inspectors and accountancy analysts red-team frontier AI models: write benign, dual-use and adversarial prompts, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwideAudit Mandarin Chinese speech data for two Amazon Sonic collections: judge annotators' transcripts against audio using fixed error codes, and fix word-level timestamp alignment. Native Mandarin as spoken in mainland China is a hard requirement; all rules and rationales are in English. Remote hourly contract at $21.50/hour.
$21.50 / HourWorldwideAudit Korean speech data for Amazon's Sonic project: check transcripts against audio with fixed error codes and pass/fail verdicts, and correct word-level timestamp alignment. Native Korean as spoken in South Korea is a hard requirement; rationales are written in English. Remote hourly contract at $25/hour.
$25 / HourWorldwideAudit Japanese speech data for Amazon's Sonic collections: verify annotators' transcripts against the audio with fixed error codes and pass/fail calls, and correct word-level timestamp alignment. Native Japanese as spoken in Japan is a hard requirement; rationales are in English. Remote hourly contract at $37.50/hour.
$37.50 / HourWorldwideTest AI chat models in Malayalam and English for safety failures and judge whether their answers are accurate, complete and appropriate. Native Malayalam plus careful evaluation skills; prior red-teaming is a plus rather than the core requirement. Remote hourly contract at $16–22/hour, paid weekly.
$16 – $22 / HourWorldwidePreclinical scientists with direct antibody-drug conjugate or bispecific antibody experience annotate R&D data and advise an AI lab building drug discovery foundation models. In-house asset developers only, 5+ years industry R&D, about 10 hours a week. $60–100/hour.
$60 – $100 / HourWorldwideNuclear medicine physicists, radiopharmacy staff and cyclotron RSOs red-team frontier AI models: write benign, dual-use and adversarial prompts from medical isotope work, judge model replies against a policy standard, and write reference answers. Remote contract at $65–75 per task.
$65 – $75 / TaskWorldwidePredict when a named sell-side analyst will publish after a catalyst and what the note will say, reasoning only from a fixed evidence cutoff, then grade AI analyses of the same call. For former lead sell-side analysts or very senior associates in the exact sector, typically 8+ years. $150–250/hour, remote.
$150 – $250 / HourWorldwideHelp run LLM training projects on browsing capabilities: track project and annotator performance in Google Sheets, review annotator quality and handle contributor communications. Ops or project management background preferred, but ownership matters more. Remote hourly contract at $50–60/hour, no location limit published.
$50 – $60 / HourWorldwideNative Danish speakers write sensitive-topic prompts, classify prompts and conversations, and flag adversarial phrasing to make AI models safer in Danish. Business English and a bachelor's (in progress counts). Part-time remote at $48–52/hour; Denmark or Western Europe preferred, not required.
$48 – $52 / HourWorldwideRead a specific non-G7 central bank's statements, minutes and speeches in the source language, score them dovish to hawkish from a fixed evidence cutoff, and grade AI macro analyses. For former central bank economists or senior local rates and FX strategists, typically 8+ years. $150–250/hour, remote.
$150 – $250 / HourWorldwideAudit French speech data for Amazon's Sonic project: check annotators' transcripts against audio with fixed error codes and pass/fail calls, and correct word-level timestamp alignment. Native French as spoken in France is a hard requirement; rationales are in English. Joint top rate of the Sonic set at $41.50/hour, remote contract.
$41.50 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.