Mercor
MercorVerified listing
RemoteAI & Machine LearningCybersecurity

AI Safety Experts — English & Finnish | $48-$62/hr Remote

48–62/hr
Remote
Posted August 3, 2026
hourly
5 openings

Overview

We're assembling a specialized team to stress-test AI systems by simulating adversarial attacks. As an AI Safety Expert, you'll probe conversational models for vulnerabilities like bias, misinformation, and harmful behaviors — all while working remotely on an hourly basis. Native fluency in both English and Finnish is essential for this role.

What You'll Do4

  • 1Conduct red teaming exercises on conversational AI models, including jailbreaks, prompt injections, misuse cases, and multi-turn manipulation.
  • 2Generate high-quality human data by annotating model failures, classifying vulnerability types, and flagging systemic risks.
  • 3Follow structured taxonomies, benchmarks, and playbooks to ensure consistent and repeatable testing.
  • 4Document findings in detailed reports and datasets that customers can use to strengthen their AI systems.

Requirements4

  • 1Native or bilingual proficiency in English and Finnish (both written and spoken).
  • 2Prior experience in AI red teaming, adversarial machine learning, or cybersecurity penetration testing.
  • 3Familiarity with frameworks like RLHF/DPO attacks, jailbreak datasets, and prompt injection techniques.
  • 4Strong analytical and documentation skills to produce clear vulnerability reports.

Who Should Apply

You're a curious and adversarial thinker who instinctively looks for ways to push systems to their limits. You enjoy structured probing and have a background in red teaming, cybersecurity, or socio-technical risk analysis. You're comfortable working independently on remote projects and can communicate complex risks to both technical and non-technical stakeholders.

Salary Insight

$48.00 - $62.00 per hour, depending on experience.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

adversarial machine learningprompt injectionjailbreakingrlhfdpopenetration testingexploit developmentreverse engineeringconversational aired teamingbenchmarking

Application Tip

Highlight specific examples of past red teaming work — such as a successful jailbreak or a creative prompt injection — in your application to demonstrate your adversarial mindset.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Swedish

This role involves stress-testing AI systems by simulating adversarial attacks in English and Swedish. You'll generate critical safety data by probing models for vulnerabilities like bias, misinformation, and security flaws. As part of a human red team, you'll help ensure AI behaves safely before it reaches users.

48–62/hr
· 5 openings
Red TeamingAdversarial Machine LearningPrompt Injection+12 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Norwegian

This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.

48–62/hr
· 5 openings
EnglishNorwegianRed Teaming+14 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Danish

We're recruiting bilingual AI safety specialists (English & Danish) to join a red team that systematically probes conversational AI models for vulnerabilities. You'll generate adversarial data, surface hidden risks, and create actionable documentation to strengthen AI systems. This remote, hourly contract pays $48–$62/hr based on experience.

48–62/hr
· 5 openings
Red TeamingAdversarial MLPrompt Injection+12 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Malay

This remote role asks bilingual experts in English and Malay to stress-test AI systems by simulating attacks and probing for weak points. You'll generate high-quality red team data that helps customers build safer, more trustworthy models. All work is text-based, and sensitive topics are clearly disclosed with wellness support.

17–25/hr
· 20 openings
Red TeamingAdversarial Machine LearningPrompt Injection+14 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Odia

We're assembling a red team of human experts to stress-test conversational AI models for safety vulnerabilities. This remote, text-based position requires native fluency in English and Odia. You'll attack AI systems with adversarial inputs, uncover weaknesses, and generate the high-quality data needed to make AI safer. The work covers sensitive topics like bias and misinformation, with clear guidelines and wellness support available.

20–22/hr
· 10 openings
Red TeamingPrompt InjectionJailbreaking+13 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Dutch

We're looking for bilingual English-Dutch AI red teamers to stress-test conversational AI models from every angle. In this remote, hourly role, you'll design and execute adversarial attacks—from jailbreaks to bias exploitation—and generate structured data that helps make AI systems safer. You'll work with clear guidelines and optional high-sensitivity projects, with wellness support available.

48–62/hr
· 5 openings
Red TeamingAdversarial Machine LearningPrompt Injection+12 more