Skip to content
Labeling Jobs

AI Safety Experts: English & Dutch

Pay
$48 – $62 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • adversarial testing
  • prompt injection
  • bias detection
  • dutch
  • report writing

What you'll do

The Dutch version of Mercor's AI safety red-team listing is the only non-Nordic language paid at the top rate in this family. Dutch speakers are also, on the whole, strong English users, which suits a role built around switching between the two to see where a model's safety behaviour stops being consistent.

What the work involves:

  • Adversarial testing of chat models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, manipulation across multiple turns
  • Producing human data from the results: annotated failures, classified vulnerabilities, flagged systemic risks
  • Following taxonomies, benchmarks and playbooks for consistency
  • Writing reproducible reports, datasets and attack cases the client can act on

Native knowledge matters for judgment calls: whether a Dutch answer is harmful or just blunt (Dutch directness is easy for a model, and a non-native reviewer, to misread), whether a model carries stereotypes about Dutch or Flemish communities, and whether it handles Netherlands-specific topics accurately. The ad does not say whether Flemish (Belgian Dutch) speakers are equally welcome; worth asking if that is your variety.

Work is text-based. Sensitive topics are part of it (bias, misinformation, harmful behaviour); topics are communicated before exposure, and higher-sensitivity projects are optional with guidelines and wellness resources.

Who qualifies

  • Native fluency in Dutch and English
  • Prior red-teaming: AI adversarial work, cybersecurity or socio-technical probing
  • Structured testing using frameworks and benchmarks
  • Clear communication of risks to technical and non-technical audiences
  • Adaptability across projects and customers

Welcome extras: adversarial ML, penetration testing and reverse engineering, harassment and disinformation analysis, and psychology, acting or writing backgrounds. No degree or years of experience are stated.

What it pays

$48–62 per hour, published by Mercor, the same as the Norwegian, Danish, Swedish and Finnish variants and the top of this family. Payment is weekly via Stripe or Wise, per the ad. H-1B and STEM OPT candidates are not supported.

Worth knowing

Good:

  • Top-band pay, weekly payouts
  • No residence requirement published
  • Topics disclosed in advance; sensitive projects are opt-in
  • Self-scheduled remote work

Less good:

  • Unlike most of its sister listings, this one showed no hires count when we read it, so it may be newer or slower to staff; do not assume immediate work
  • Prior red-teaming experience is expected
  • Distressing content is part of the job
  • No weekly hours are stated, and projects can end early
  • Independent contractor terms; clients and models not named

About this listing

A Mercor remote hourly contract, read from the Mercor site on 24 September 2026. Pay, requirements, payment terms and the visa exclusion are as published. Mercor states the work involves no confidential information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.