Skip to content
Labeling Jobs

AI Safety Experts: English & Vietnamese

Pay
$17 – $25 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • jailbreak testing
  • vulnerability classification
  • data annotation
  • vietnamese
  • prompt injection

What you'll do

Vietnamese has a feature that makes it interesting for safety testing: the same words can be typed with full diacritics, with none at all, or in the shorthand common in Vietnamese online chat, and models do not always treat those forms the same way. A native speaker knows how real users write, and can use that to see whether a model's safeguards hold up outside tidy, textbook input.

The role, per Mercor's brief:

  • Red-team conversational models and agents with jailbreaks, prompt injections, misuse cases, bias exploitation and multi-turn manipulation
  • Generate human data from what you find: annotate failures, classify vulnerabilities, flag systemic risks
  • Apply the project's taxonomies, benchmarks and playbooks
  • Document everything reproducibly in reports, datasets and attack cases

Judgment on Vietnamese-specific content (regional and generational language differences, politically and historically sensitive subjects, local misinformation) is where you add what automated testing lacks.

All work is text. You will see sensitive topics such as bias, misinformation and harmful behaviour; topics are communicated before exposure, and higher-sensitivity projects are opt-in with guidelines and wellness resources.

Who fits

  • Native fluency in Vietnamese and English, both required
  • Prior red-teaming experience: AI adversarial work, cybersecurity or socio-technical probing
  • A structured, benchmark-driven method
  • Clear communication with technical and non-technical stakeholders
  • Adaptability across projects and clients

The ad lists adversarial ML, pentesting and reverse engineering, harassment and disinformation probing, and psychology, acting or writing backgrounds as pluses. No degree or minimum years are published.

What it pays

$17–25 per hour, Mercor's published figure, the same band as Indonesian and Malay in this family. For comparison, Thai is $24–35 and the Nordic and Dutch versions pay $48–62 for the same template.

Payment is weekly via Stripe or Wise, per the ad. H-1B and STEM OPT candidates are not supported.

Worth knowing

Good:

  • Paid in US dollars, weekly
  • No residence requirement published, so the Vietnamese diaspora can apply
  • Topics announced in advance; sensitive projects are optional
  • Remote and self-scheduled

Less good:

  • No hires count was shown on this listing when we read it, unlike the Indonesian and Thai versions, so staffing may be slower here
  • The experience bar is the same as for listings paying up to three times more
  • Distressing material is part of red-team work
  • Hours are not stated, and projects can be extended, shortened or ended early
  • Contractor status; clients and models are not named

About this listing

Posted by Mercor as a remote hourly contract, read from the Mercor site on 24 September 2026. Pay, requirements, payment terms and the visa exclusion are the ad's own. Mercor states the work involves no confidential information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.