
AI Safety Experts — English & Swedish | $48-$62/hr Remote
Overview
This role involves stress-testing AI systems by simulating adversarial attacks in English and Swedish. You'll generate critical safety data by probing models for vulnerabilities like bias, misinformation, and security flaws. As part of a human red team, you'll help ensure AI behaves safely before it reaches users.
What You'll Do4
- 1Probe conversational AI models with adversarial inputs to uncover weaknesses such as jailbreaks, prompt injections, and bias exploitation.
- 2Create high-quality human data by categorizing failures, classifying vulnerabilities, and flagging systemic risks.
- 3Follow structured taxonomies and playbooks to keep testing consistent and reproducible across projects.
- 4Produce detailed reports and datasets that engineers can use to strengthen AI safety and trustworthiness.
Requirements4
- 1Prior experience in AI red teaming, including adversarial testing or socio-technical probing.
- 2A curious, adversarial mindset with a knack for pushing systems to their limits.
- 3Strong ability to document and communicate risks clearly to both technical and non-technical stakeholders.
- 4Native or fluent proficiency in both English and Swedish.
Who Should Apply
This role is ideal for someone with a background in AI safety, cybersecurity, or creative adversarial thinking. You should enjoy systematically breaking things and have a structured approach to documenting vulnerabilities. If you're fluent in English and Swedish and want to make AI systems safer, this project offers a direct way to contribute.
Salary Insight
$48.00 - $62.00 per hour, depending on experience.
Location
Required Skills
Application Tip
Highlight specific examples of adversarial testing you've done — whether jailbreaking LLMs, finding prompt injection flaws, or uncovering bias in model outputs. Show how you document and reproduce your findings to make your application stand out.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedAI Safety Experts — English & Norwegian
This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.

Mercor
VerifiedAI Safety Experts — English & Finnish
We're assembling a specialized team to stress-test AI systems by simulating adversarial attacks. As an AI Safety Expert, you'll probe conversational models for vulnerabilities like bias, misinformation, and harmful behaviors — all while working remotely on an hourly basis. Native fluency in both English and Finnish is essential for this role.

Mercor
VerifiedAI Safety Experts — English & Bengali
This role brings fluent English and Bengali speakers into a remote AI red team. You'll probe conversational AI systems with adversarial inputs to uncover biases, misinformation, and manipulation risks. All work is text-based, and you'll produce structured red team data that improves AI safety. You can opt out of higher-sensitivity topics, and clear guidelines plus wellness resources are provided.

Mercor
VerifiedAI Safety Experts — English & Danish
We're recruiting bilingual AI safety specialists (English & Danish) to join a red team that systematically probes conversational AI models for vulnerabilities. You'll generate adversarial data, surface hidden risks, and create actionable documentation to strengthen AI systems. This remote, hourly contract pays $48–$62/hr based on experience.

Mercor
VerifiedAI Safety Experts — English & Odia
We're assembling a red team of human experts to stress-test conversational AI models for safety vulnerabilities. This remote, text-based position requires native fluency in English and Odia. You'll attack AI systems with adversarial inputs, uncover weaknesses, and generate the high-quality data needed to make AI safer. The work covers sensitive topics like bias and misinformation, with clear guidelines and wellness support available.

Mercor
VerifiedAI Safety Experts — English & Malay
This remote role asks bilingual experts in English and Malay to stress-test AI systems by simulating attacks and probing for weak points. You'll generate high-quality red team data that helps customers build safer, more trustworthy models. All work is text-based, and sensitive topics are clearly disclosed with wellness support.

