Skip to content
Labeling Jobs

AI Safety Experts: English & Swedish

Pay
$48 – $62 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • red teaming
  • jailbreak testing
  • bias detection
  • vulnerability classification
  • swedish
  • report writing

What you'll do

Of all the language versions of Mercor's AI safety red-team listing, the Swedish one is hiring fastest: 243 people this month when we checked, well ahead of the Norwegian, Danish and Finnish versions. That suggests either a large Swedish workstream or several clients at once. The ad does not say which.

The work itself is the same adversarial brief. You test conversational models and AI agents by trying to get them to do what they should not: bypass their safety behaviour through jailbreaks, follow instructions hidden in content (prompt injection), reveal bias, or drift into misuse over a sequence of turns. Then you convert results into human data: annotated failures, vulnerability classifications, flags on systemic problems, and reports and datasets the client can act on. Taxonomies, benchmarks and playbooks keep the testing consistent.

A native Swedish tester adds what automated checks miss: whether a response is harmful in its Swedish context, whether a model repeats a stereotype about a Swedish group, whether a translation-style answer hides a safety gap, and whether switching language mid-conversation weakens a refusal.

Everything is text-based. You will review outputs on sensitive topics including bias, misinformation and harmful behaviour; topics are communicated before exposure, and higher-sensitivity projects are optional, backed by guidelines and wellness resources.

Who fits

  • Native-level Swedish and English (both required)
  • Prior red-teaming of some kind: AI adversarial work, cybersecurity, or socio-technical probing
  • Comfortable working in a structured way, inside frameworks and benchmarks
  • Able to explain a risk plainly to technical and non-technical people
  • Happy switching between projects and customers

Specialties the ad welcomes: adversarial ML, penetration testing and reverse engineering, harassment and disinformation analysis, and creative fields like psychology, acting and writing.

No degree or minimum years are published.

What it pays

$48–62 per hour, Mercor's own figure, the top band in this listing family and matched only by the other Nordic versions and Dutch. It is roughly double the Thai rate and close to three times the Indian-language rates for the same project template.

Weekly payment via Stripe or Wise is stated in the ad. H-1B and STEM OPT candidates cannot be supported.

Worth knowing

Good:

  • Highest hiring volume in this family, which usually means a faster path from application to work
  • Top-band pay and weekly payouts
  • No residence requirement in the ad
  • Opt-in for the most sensitive material, with topics announced first

Less good:

  • A busy listing also means competition; your screening answers need to show real adversarial experience
  • Red-teaming involves reading and eliciting harmful text by design
  • No stated hours, and projects can be extended, shortened or ended early
  • Independent contractor terms, no benefits
  • Clients and models are unnamed

About this listing

A Mercor remote hourly contract, read from the Mercor site on 24 September 2026. Pay, language and experience requirements, payment terms and the visa exclusion are as published. Mercor states the work will not involve confidential information from any employer, client or institution. See Mercor.

More roles at Mercor

See all 304

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.