Mercor
MercorVerified listing
RemoteAI & Machine LearningCybersecurity

AI Safety Experts — English & Malay | $17-$25/hr Remote

17–25/hr
Remote
Posted August 4, 2026
hourly
20 openings

Overview

This remote role asks bilingual experts in English and Malay to stress-test AI systems by simulating attacks and probing for weak points. You'll generate high-quality red team data that helps customers build safer, more trustworthy models. All work is text-based, and sensitive topics are clearly disclosed with wellness support.

What You'll Do4

  • 1Probe conversational AI models and agents for security weaknesses, including jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation.
  • 2Produce high-quality human data by documenting model failures, classifying vulnerability types, and flagging systemic risks.
  • 3Follow predefined taxonomies, benchmarks, and testing playbooks to keep red teaming consistent and measurable.
  • 4Create reproducible reports, datasets, and attack examples that clients can use to strengthen their AI systems.

Requirements6

  • 1Native-level fluency in English and Malay (written and spoken).
  • 2Hands-on experience with red teaming in areas like adversarial machine learning, cybersecurity, or socio-technical probing.
  • 3A structured, adversarial mindset—you use frameworks and benchmarks, not just random testing.
  • 4Strong communication skills to explain risks to both technical and non-technical audiences.
  • 5Comfort working across multiple projects and adapting to new customer domains quickly.
  • 6Familiarity with jailbreak datasets, RLHF/DPO attacks, model extraction, penetration testing, exploit development, or disinformation/abuse analysis is a plus.

Who Should Apply

You're the kind of person who instinctively wants to break things to make them better. You enjoy adversarial thinking, work methodically, and can clearly document what you find. If you're bilingual in English and Malay and have any background in AI safety, security testing, or creative probing, this is a great fit.

Salary Insight

$17.00–$25.00 per hour, remote, paid hourly.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

red teamingadversarial machine learningprompt injectionjailbreakrlhfdpomodel extractionpenetration testingexploit developmentreverse engineeringconversational aidata annotationtaxonomybenchmarkenglishmalaycybersecurity

Application Tip

Highlight a specific red teaming example in your application—describe a vulnerability you found, how you tested it, and the reproducible steps you took. This will show you work with structure and can communicate technical findings clearly.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Bengali

This role brings fluent English and Bengali speakers into a remote AI red team. You'll probe conversational AI systems with adversarial inputs to uncover biases, misinformation, and manipulation risks. All work is text-based, and you'll produce structured red team data that improves AI safety. You can opt out of higher-sensitivity topics, and clear guidelines plus wellness resources are provided.

20–22/hr
· 10 openings
EnglishBengaliConversational AI+14 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Indonesian

Neon is looking for bilingual AI safety experts fluent in English and Indonesian to join a human-driven red team. You'll probe advanced conversational AI models for vulnerabilities like bias, misinformation, and harmful behavior — all through text-based work. This remote, hourly role pays $17–$25/hr and gives you a direct hand in making AI systems safer before they reach the public.

17–25/hr
· 5 openings
Adversarial Machine LearningRed TeamingPrompt Injection+13 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Norwegian

This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.

48–62/hr
· 5 openings
EnglishNorwegianRed Teaming+14 more
Mercor

Mercor

30d agoRemotehourly

AI Safety Experts — English & Thai

Neon is assembling a remote team of bilingual (English & Thai) AI safety specialists to stress-test conversational AI systems. You'll probe models for vulnerabilities like bias, misinformation, and prompt injection, then document and classify findings to help make AI safer. This text-based red teaming role offers flexible hourly work and the chance to shape AI safety at the frontier.

24–35/hr
· 5 openings
Red TeamingAdversarial MLPrompt Injection+10 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Dutch

We're looking for bilingual English-Dutch AI red teamers to stress-test conversational AI models from every angle. In this remote, hourly role, you'll design and execute adversarial attacks—from jailbreaks to bias exploitation—and generate structured data that helps make AI systems safer. You'll work with clear guidelines and optional high-sensitivity projects, with wellness support available.

48–62/hr
· 5 openings
Red TeamingAdversarial Machine LearningPrompt Injection+12 more
Mercor

Mercor

1mo agoRemotehourly

AI Safety Experts — English & Punjabi

Join a remote red team that stress-tests AI models for safety vulnerabilities. You will probe conversational AI with adversarial inputs, uncover weak spots like jailbreaks or bias, and turn those findings into data that makes AI safer for customers. Native-level fluency in English and Punjabi is a must. The work is fully text-based and paid hourly, with optional exposure to sensitive topics supported by clear guidelines.

20–22/hr
· 10 openings
EnglishPunjabiRed Teaming+16 more