Skip to content
Labeling Jobs

AI Safety Red Teamer

Pay
$70 – $84 / Hour
Open to

Albania, Austria, Belgium, Bosnia & Herzegovina, Bulgaria, Croatia, Czechia, Denmark

and 32 more countries

Estonia, Finland, France, Germany, Greece, Hungary, Iceland, Ireland, Italy, Kosovo, Latvia, Liechtenstein, Lithuania, Luxembourg, Malta, Moldova, Monaco, Netherlands, North Macedonia, Norway, Poland, Portugal, Romania, San Marino, Serbia, Slovakia, Slovenia, Spain, Sweden, Switzerland, United Kingdom, United States

Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • adversarial prompting
  • jailbreak testing
  • trust and safety
  • ai safety evaluation
  • technical writing

What you'll do

Unlike Mercor's domain-specific red-team panels (chemistry, nuclear, radiological, energetic materials), this is a generalist adversarial testing role. You stress-test frontier AI models across a broad spread of sensitive areas, looking for the places they break:

  • Design adversarial prompts and probe for jailbreaks, unsafe behaviour, hallucinations and policy failures
  • Test robustness across misinformation, cyber, biosecurity, fraud, political content and other grey-area topics
  • Document the vulnerabilities you find and contribute to safety benchmarks and red-teaming reports
  • Work with AI researchers on alignment, robustness and safety improvements

The emphasis is on attack craft rather than one field's expertise: finding the framing, sequence or ambiguity that gets a model to do something it should not, and writing it up so researchers can fix it. The ad does not say whether prompts are single-turn or multi-turn, or how tasks are structured.

Who qualifies

The bar is higher than most red-team listings:

  • A bachelor's degree or higher in computer science, cybersecurity, journalism, communications, psychology, biology, chemistry, public policy or a related field
  • 5+ years of professional experience in AI safety, AI red teaming, trust and safety, cybersecurity, investigative journalism, life sciences or a related field
  • Strong analytical reasoning, prompt design and written communication
  • Experience designing adversarial prompts or evaluating frontier AI systems

Preferred: hands-on work with AI red teaming, RLHF, SFT, alignment or trust and safety; familiarity with jailbreak testing and adversarial evaluation methods; and depth in at least one grey-area domain such as cyber, biosecurity, political content, misinformation or scientific safety.

The degree list is unusually broad. Investigative journalists and trust and safety staff qualify alongside security researchers, which suits people whose skill is thinking like a bad actor rather than a specific technical specialty.

Where you can work from

Residence is restricted to a published list of 40 countries: every EU member except Cyprus; Albania, Bosnia and Herzegovina, Iceland, Kosovo, Liechtenstein, Moldova, Monaco, North Macedonia, Norway, San Marino, Serbia and Switzerland; plus the United Kingdom and the United States. Canada and Australia are not on it.

What it pays

$70–84 per hour, Mercor's published range. That is an hourly rate, so it is not directly comparable with the $65–75 per task on Mercor's domain red-team panels, but it gives certainty those listings lack. The ad does not say what decides placement within the band.

Payment is weekly via Stripe or Wise. Mercor cannot support H-1B or STEM OPT candidates. You work as an independent contractor on your own schedule, and projects may be extended, shortened or ended early depending on needs and performance.

Worth knowing

Good:

  • Hourly pay with published weekly payment terms
  • Broad range of backgrounds accepted, including journalism and trust and safety
  • Covers many domains, so the work is varied
  • Work directly alongside AI researchers, per the ad

Less good:

  • Five years' professional experience plus prior adversarial prompt work is a real filter
  • Narrow band, so experienced red-teamers have little upside
  • Limited to Europe, the UK and the US
  • H-1B and STEM OPT holders excluded
  • No weekly hours or project length stated

About this listing

Posted by Mercor as an hourly remote contract, confirmed open on 24 September 2026. Mercor states the work involves no access to confidential or proprietary information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.