Skip to content
Labeling Jobs

AI Safety Experts: English & Portuguese (global)

Pay
$29 – $45 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • jailbreak testing
  • bias detection
  • vulnerability classification
  • european portuguese
  • report writing

What you'll do

The important detail in this listing is the variety of Portuguese. Mercor wants native Portuguese "global, excluding Brazilian Portuguese", which in practice points to speakers of European Portuguese and the other non-Brazilian varieties (Angola, Mozambique and elsewhere). The ad does not list countries, but Brazilian speakers are ruled out by name.

That split makes sense for safety testing. Most Portuguese text on the web is Brazilian, so models tend to default to it, and their handling of European Portuguese vocabulary, idiom and social context is weaker, which is where safeguards tend to slip.

The work:

  • Probe chat models and agents with jailbreaks, prompt injections, misuse scenarios, bias exploitation and multi-turn manipulation
  • Record results as human data: annotated failures, classified vulnerabilities, flagged systemic risks
  • Keep testing consistent using the project's taxonomies, benchmarks and playbooks
  • Deliver reproducible reports, datasets and attack cases

Everything is text. Content includes sensitive topics such as bias, misinformation and harmful behaviour; topics are disclosed in advance, and higher-sensitivity projects are optional with guidelines and wellness support.

Who fits

  • Native fluency in English and non-Brazilian Portuguese
  • Prior red-teaming experience: AI adversarial work, cybersecurity or socio-technical probing
  • A structured approach, working from frameworks and benchmarks
  • Clear communication of risk to mixed audiences
  • Adaptability across projects and clients

Pluses: adversarial ML, penetration testing and reverse engineering, harassment and disinformation analysis, and backgrounds in psychology, acting or writing.

No degree or minimum experience in years is stated.

What it pays

$29–45 per hour, Mercor's published band. It sits in the middle of this listing family: below the $48–62 Nordic and Dutch rate, above Thai at $24–35 and well above the Southeast Asian and Indian-language variants. The ceiling here is the highest outside the top band.

Payments are weekly via Stripe or Wise, according to the ad. H-1B and STEM OPT candidates are not supported.

Worth knowing

Good:

  • A listing that specifically values European and African Portuguese, which are often an afterthought on AI data platforms
  • 172 hires this month, so the project is moving
  • Mid-to-upper pay within this family, paid weekly
  • No residence requirement published
  • Sensitive work is opt-in, with topics announced first

Less good:

  • If your Portuguese is Brazilian, this listing is closed to you by its own terms
  • Prior red-teaming experience is expected
  • Some content is distressing by nature
  • Hours are not stated, and projects can be extended, shortened or ended early
  • Contractor terms, and the clients are not named

About this listing

Posted by Mercor as a remote hourly contract, read from the Mercor site on 24 September 2026. Pay, the Portuguese variety requirement, payment terms and the visa exclusion are as published. Mercor notes the work does not involve confidential information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.