Skip to content
Labeling Jobs

AI Safety Experts: English & Malayalam

Pay
$16 – $22 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • ai response evaluation
  • content review
  • transliteration
  • malayalam
  • data annotation

What you'll do

Malayalam is one of the harder Indian languages for AI models to handle well. Its script is complex, words run long because of heavy agglutination, and a lot of everyday Malayalam online is written in "Manglish", Malayalam typed in Latin letters and mixed with English. A model that behaves safely on formal Malayalam may not on Manglish, and finding those gaps is part of what Mercor's AI safety project pays for.

The tasks in the ad are adversarial: try to get chat models and agents to break their rules with jailbreaks, prompt injections, misuse cases, bias exploitation and gradual multi-turn manipulation; then annotate the failures, classify them, flag systemic risks and write reproducible reports and attack cases, all following the project's taxonomies and playbooks.

The profile Mercor describes for this variant, though, is an evaluator's. Instead of prior red-teaming experience (which the European versions ask for), it wants people who can tell whether an AI answer is accurate, complete and appropriate and explain why, who catch small errors, and who apply guidelines consistently. So expect a blend: quality judgment on Malayalam output plus structured attempts to provoke bad output.

Text only. Sensitive topics (bias, misinformation, harmful behaviour) are included; topics are disclosed before exposure, and higher-sensitivity projects are optional with guidelines and wellness support.

Who qualifies

  • Native fluency in Malayalam and English
  • Strong judgment about language and content, with clear reasons for your calls
  • Rigour and consistency when working to quality standards
  • Clear written communication for technical and non-technical readers
  • Adaptability across task types and clients

Bonus backgrounds named in the ad: adversarial ML, cybersecurity, harassment or disinformation analysis, and psychology, acting or writing. No degree or minimum years are stated.

What it pays

$16–22 per hour, published by Mercor, the same as every other Indian-language variant in this family and Urdu. That is roughly a third of the $48–62 Nordic and Dutch rate for the same template.

Payments are weekly via Stripe or Wise, per the ad. H-1B and STEM OPT candidates are not supported.

Worth knowing

Good:

  • Accessible to careful native reviewers without a security background
  • Weekly payouts in US dollars
  • No residence requirement published, so the Malayali diaspora in the Gulf and elsewhere can apply
  • Sensitive work is opt-in

Less good:

  • Lowest pay band in the family
  • No hires count on the listing when we read it
  • Red-teaming content can be distressing
  • No weekly hours stated; projects can be extended, shortened or ended early
  • Contractor terms; clients and models not named

About this listing

A Mercor remote hourly contract, read from the Mercor site on 24 September 2026. Pay, requirements, payment terms and the visa exclusion come from the ad. Mercor states the work involves no confidential information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.