Mercor
MercorVerified listing
RemoteAI & Machine LearningCybersecurity

AI Safety Red Teamer | $70-$84/hr Remote

70–84/hr
Remote
Posted August 4, 2026
hourly
4 openings

Overview

Join a team focused on strengthening the safety of advanced AI systems by uncovering their hidden weaknesses. As an AI Safety Red Teamer, you'll design clever prompts to stress-test models, spot dangerous behaviors, and help improve how these systems handle tricky real-world scenarios. This remote contractor role lets you work alongside top researchers while earning $70–$84 per hour.

What You'll Do5

  • 1Create adversarial prompts that push frontier AI models to their limits, revealing potential failure points.
  • 2Identify and document jailbreaks, unsafe outputs, hallucinations, and policy breaches in model behavior.
  • 3Test model performance across high-risk areas including misinformation, cybersecurity, biosecurity, and political content.
  • 4Collaborate with AI researchers to refine model alignment and boost overall robustness.
  • 5Compile clear vulnerability reports and contribute to safety benchmarks and red-teaming documentation.

Requirements4

  • 1A bachelor’s degree or higher in fields like Computer Science, Cybersecurity, Journalism, or a related discipline.
  • 2At least 5 years of professional experience in AI safety, red teaming, trust & safety, or investigative work.
  • 3Strong skills in analytical reasoning, prompt engineering, and written communication to document findings clearly.
  • 4Hands-on experience designing adversarial tests or evaluating frontier AI systems for vulnerabilities.

Who Should Apply

This role is ideal for someone who enjoys thinking like an attacker to make AI systems safer. You're curious, methodical, and have a knack for crafting prompts that reveal hidden flaws. If you have a background in safety research, cybersecurity, or investigative work and want to shape how cutting-edge models handle grey-area challenges, this is your opportunity.

Salary Insight

The position pays an hourly rate of $70–$84, reflecting the specialized skill set required for adversarial testing of advanced AI.

Location

Typeremote
LocationRemote
Eligible countriesUnited States, Denmark, Estonia, Finland, Iceland +35 more
This is a remote position

Application Tip

When applying, include a short portfolio or example of a prompt you designed that uncovered a realistic edge case in a language model — this shows your hands-on red-teaming approach.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

SME Careers

SME Careers

4d agoRemotecontract

Red-Teaming QA Lead for AI Safety and Evaluation

As a remote Red-Teaming QA Lead, you guide quality and consistency across AI red-teaming and safety-evaluation work performed by a distributed contractor team. You review ai red-teaming outputs, assess adversarial prompts and risk classifications, and provide precise written feedback to keep guidelines aligned. You’ll maintain project rubrics, coordinate updates for trainers and QAs, and manage documentation across a fast-moving remote workflow using Discord, Google Sheets, and dashboards. This hourly contract role supports SME Careers’ AI data services and helps improve safety training data for leading models.

Up to 100/hr
Adversarial TestingAI Red TeamingAI Safety+24 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Indonesian

Neon is looking for bilingual AI safety experts fluent in English and Indonesian to join a human-driven red team. You'll probe advanced conversational AI models for vulnerabilities like bias, misinformation, and harmful behavior — all through text-based work. This remote, hourly role pays $17–$25/hr and gives you a direct hand in making AI systems safer before they reach the public.

17–25/hr
· 5 openings
Adversarial Machine LearningRed TeamingPrompt Injection+13 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Punjabi

Join a remote red team that stress-tests AI models for safety vulnerabilities. You will probe conversational AI with adversarial inputs, uncover weak spots like jailbreaks or bias, and turn those findings into data that makes AI safer for customers. Native-level fluency in English and Punjabi is a must. The work is fully text-based and paid hourly, with optional exposure to sensitive topics supported by clear guidelines.

20–22/hr
· 10 openings
EnglishPunjabiRed Teaming+16 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Norwegian

This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.

48–62/hr
· 5 openings
EnglishNorwegianRed Teaming+14 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Bengali

This role brings fluent English and Bengali speakers into a remote AI red team. You'll probe conversational AI systems with adversarial inputs to uncover biases, misinformation, and manipulation risks. All work is text-based, and you'll produce structured red team data that improves AI safety. You can opt out of higher-sensitivity topics, and clear guidelines plus wellness resources are provided.

20–22/hr
· 10 openings
EnglishBengaliConversational AI+14 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Odia

We're assembling a red team of human experts to stress-test conversational AI models for safety vulnerabilities. This remote, text-based position requires native fluency in English and Odia. You'll attack AI systems with adversarial inputs, uncover weaknesses, and generate the high-quality data needed to make AI safer. The work covers sensitive topics like bias and misinformation, with clear guidelines and wellness support available.

20–22/hr
· 10 openings
Red TeamingPrompt InjectionJailbreaking+13 more