Skip to content
Labeling Jobs

AI Safety Experts: English & Malay

Pay
$17 – $25 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • adversarial testing
  • bias detection
  • data annotation
  • malay
  • prompt injection

What you'll do

Malay has a close neighbour problem. It is mutually intelligible with Indonesian to a large degree, and because Indonesian dominates the combined training data, models often slip into Indonesian vocabulary or assumptions when they should be answering in Malay. Mercor is running this Malay red-team listing alongside a separate Indonesian one, which suggests the client wants the two tested as distinct languages.

The work:

  • Attack chat models and agents with jailbreaks, prompt injections, misuse scenarios, bias exploitation and manipulation built over several turns
  • Record outcomes as human data: annotated failures, classified vulnerabilities, flagged systemic risks
  • Stay consistent with taxonomies, benchmarks and playbooks
  • Produce reproducible reports, datasets and attack cases

A native speaker brings the judgment the job needs: spotting responses that are harmful or offensive in a Malaysian, Bruneian or Singaporean context, catching bias about the region's ethnic and religious communities, and noticing when mixing Malay and English (as many speakers do daily) weakens a refusal.

The work is text-only. Sensitive topics (bias, misinformation, harmful behaviour) are part of it; topics are shared before exposure, and higher-sensitivity projects are optional with guidelines and wellness resources.

Who fits

  • Native fluency in Malay and English
  • Prior red-teaming: AI adversarial work, cybersecurity or socio-technical probing
  • Framework- and benchmark-driven testing habits
  • Clear risk explanations for technical and non-technical readers
  • Comfort switching projects and clients

The ad also welcomes adversarial ML, pentesting and reverse engineering, harassment and disinformation analysis, and psychology, acting or writing experience. No degree or years of experience are stated.

What it pays

$17–25 per hour, Mercor's published band, identical to the Indonesian and Vietnamese variants. It is under half of the $48–62 top rate for this template.

Weekly payment via Stripe or Wise is stated in the ad. H-1B and STEM OPT candidates are not supported.

Worth knowing

Good:

  • A dedicated Malay listing rather than Malay folded into Indonesian
  • US-dollar pay, weekly
  • No residence requirement published
  • Sensitive work is opt-in, with topics disclosed first

Less good:

  • No hires count appeared on this listing when we read it, so volume is unclear
  • Prior red-teaming experience is expected even at this rate
  • Unpleasant content is part of red-teaming
  • Hours not stated; projects can be extended, shortened or ended early
  • Contractor terms, and clients are not named

About this listing

A Mercor remote hourly contract, read from the Mercor site on 24 September 2026. Pay, requirements, payment terms and the visa exclusion are as published. Mercor states the work involves no confidential information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.