Mercor
MercorVerified listing
RemoteAI & Machine LearningCybersecurity

AI Safety Experts — English & Dutch | $48-$62/hr Remote

48–62/hr
Remote
Posted August 3, 2026
hourly
5 openings

Overview

We're looking for bilingual English-Dutch AI red teamers to stress-test conversational AI models from every angle. In this remote, hourly role, you'll design and execute adversarial attacks—from jailbreaks to bias exploitation—and generate structured data that helps make AI systems safer. You'll work with clear guidelines and optional high-sensitivity projects, with wellness support available.

What You'll Do4

  • 1Probe conversational AI models and agents with adversarial techniques like jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
  • 2Generate high-quality human-labeled data by annotating model failures, classifying vulnerabilities, and identifying systemic risks.
  • 3Follow established taxonomies, benchmarks, and playbooks to keep red teaming consistent and reproducible.
  • 4Document findings thoroughly—produce reports, datasets, and attack examples that customers can use to strengthen their AI systems.

Requirements4

  • 1Prior experience in AI red teaming, adversarial machine learning, cybersecurity, or socio-technical probing of AI systems.
  • 2A naturally adversarial and curious mindset—you instinctively look for ways to push systems to their breaking points.
  • 3Ability to work within structured frameworks and benchmarks rather than relying on random hacking.
  • 4Strong communication skills to explain risks clearly to both technical and non-technical stakeholders.

Who Should Apply

This role is for a detail-oriented, adversarial thinker who combines deep language fluency in English and Dutch with hands-on experience breaking AI models. You're comfortable following structured taxonomies, documenting reproduceable vulnerabilities, and collaborating across diverse projects. Backgrounds in adversarial ML, cybersecurity, socio-technical risk analysis, or creative adversarial thinking (e.g., psychology, acting, writing) are a strong plus.

Salary Insight

The role pays between $48.00 and $62.00 per hour, based on experience and specialization.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

red teamingadversarial machine learningprompt injectionjailbreak testingcybersecuritypenetration testingvulnerability researchdata annotationtaxonomy developmentbenchmark testingrisk assessmentconversational ainlpenglishdutch

Application Tip

When applying, include a brief case study of a successful red teaming engagement where you uncovered a vulnerability that automated tests missed—specifically note any multilingual or cross-cultural aspect if possible.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Norwegian

This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.

48–62/hr
· 5 openings
EnglishNorwegianRed Teaming+14 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Malay

This remote role asks bilingual experts in English and Malay to stress-test AI systems by simulating attacks and probing for weak points. You'll generate high-quality red team data that helps customers build safer, more trustworthy models. All work is text-based, and sensitive topics are clearly disclosed with wellness support.

17–25/hr
· 20 openings
Red TeamingAdversarial Machine LearningPrompt Injection+14 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Finnish

We're assembling a specialized team to stress-test AI systems by simulating adversarial attacks. As an AI Safety Expert, you'll probe conversational models for vulnerabilities like bias, misinformation, and harmful behaviors — all while working remotely on an hourly basis. Native fluency in both English and Finnish is essential for this role.

48–62/hr
· 5 openings
Adversarial Machine LearningPrompt InjectionJailbreaking+8 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Danish

We're recruiting bilingual AI safety specialists (English & Danish) to join a red team that systematically probes conversational AI models for vulnerabilities. You'll generate adversarial data, surface hidden risks, and create actionable documentation to strengthen AI systems. This remote, hourly contract pays $48–$62/hr based on experience.

48–62/hr
· 5 openings
Red TeamingAdversarial MLPrompt Injection+12 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Indonesian

Neon is looking for bilingual AI safety experts fluent in English and Indonesian to join a human-driven red team. You'll probe advanced conversational AI models for vulnerabilities like bias, misinformation, and harmful behavior — all through text-based work. This remote, hourly role pays $17–$25/hr and gives you a direct hand in making AI systems safer before they reach the public.

17–25/hr
· 5 openings
Adversarial Machine LearningRed TeamingPrompt Injection+13 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Bengali

This role brings fluent English and Bengali speakers into a remote AI red team. You'll probe conversational AI systems with adversarial inputs to uncover biases, misinformation, and manipulation risks. All work is text-based, and you'll produce structured red team data that improves AI safety. You can opt out of higher-sensitivity topics, and clear guidelines plus wellness resources are provided.

20–22/hr
· 10 openings
EnglishBengaliConversational AI+14 more