AI Safety Experts: English & Finnish
- Pay
- $48 – $62 / Hour
- Open to
- Worldwide
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- red teaming
- adversarial testing
- jailbreak testing
- data annotation
- finnish
- risk reporting
What you'll do
Finnish is a useful language for breaking AI models. It is morphologically rich, far less represented in training data than English, and unrelated to the Germanic and Romance languages that most safety tuning leans on. Mercor's red team wants native speakers who can exploit that gap deliberately and record what they find.
In practice you:
- Try to push chat models and agents past their safeguards with jailbreaks, prompt injections, misuse scenarios, bias exploitation and manipulation built over several turns
- Annotate each failure, classify the vulnerability, and flag anything that looks systemic
- Work to taxonomies, benchmarks and playbooks so that results are comparable across testers
- Produce reproducible reports, datasets and attack cases for the client
Judgment is a large part of it. Deciding whether a Finnish response is genuinely harmful, merely clumsy, or culturally off in a way an English reviewer would never catch is exactly the skill being bought.
The work is all text. Sensitive topics (bias, misinformation, harmful behaviour) come with the territory, but topics are disclosed before you see content, and higher-sensitivity projects are optional with guidelines and wellness resources.
Who fits
- Native fluency in Finnish and English, both required
- Prior red-teaming experience: AI adversarial work, cybersecurity, or socio-technical probing
- A systematic approach built on frameworks and benchmarks
- Clear communication of risk to technical and non-technical audiences
- Adaptability across projects and clients
Nice to have: adversarial ML (jailbreak datasets, RLHF/DPO attacks, model extraction), penetration testing and reverse engineering, harassment and disinformation probing, and creative backgrounds such as psychology, acting or writing.
The ad sets no degree or years-of-experience floor.
What it pays
$48–62 per hour, published by Mercor. Finnish shares the top band with Norwegian, Danish, Swedish and Dutch in this family; there is no Finnish premium or discount.
Payments are weekly via Stripe or Wise, per the ad. H-1B and STEM OPT candidates are not supported.
Worth knowing
Good:
- Top-band hourly rate for this template
- 137 hires on the listing this month, so it is actively staffing
- No residence requirement published
- Self-scheduled and fully remote
- Sensitive projects are opt-in, with advance notice of topics
Less good:
- Expect to need genuine adversarial experience, not only Finnish
- Some content will be unpleasant by design
- Weekly hours are not stated, and projects can end early
- Contractor terms, with no benefits
- The clients and models are not named
About this listing
Posted by Mercor as a remote hourly contract, read from the Mercor site on 24 September 2026. The pay, requirements, payment terms and visa exclusion are taken from the ad. Mercor says the work involves no confidential information from any employer, client or institution. See Mercor.
More roles at Mercor
See all 304Similar roles at other platforms
Guides about Mercor
See all 52- How much does Mercor pay? Rates from 304 live listingsPay breakdown
- The Mercor AI interview: what happens, what it checks, and what comes afterInterview prep
- Mercor vs Alignerr: which to apply to, and how they differPlatform comparison
- Mercor, micro1, Outlier, Alignerr and Handshake AI compared: which to apply to firstPlatform comparison
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.