
AI Safety Experts — English & Norwegian | $48-$62/hr Remote
Overview
This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.
What You'll Do5
- 1Stress-test conversational AI through jailbreaks, prompt injections, misuse scenarios, and multi-turn manipulation to uncover hidden weaknesses.
- 2Annotate and classify failures, assign severity levels, and flag systemic risks in model outputs using structured taxonomies and benchmarks.
- 3Produce reproducible attack cases and datasets that customers can act on to harden their AI systems.
- 4Follow established playbooks to keep evaluation consistent across diverse projects and clients.
- 5Optionally engage in higher-sensitivity content probing with access to wellness resources and advance topic briefings.
Requirements4
- 1Native or bilingual fluency in English and Norwegian (both written) to assess nuanced language outputs.
- 2Proven experience in red teaming AI or cybersecurity, including adversarial testing, penetration testing, or socio-technical probing.
- 3Strong ability to apply structured frameworks and benchmarks rather than ad-hoc testing; familiarity with RLHF, DPO, or model extraction is a plus.
- 4Clear communication skills to explain technical risks to both engineers and non-technical stakeholders.
Who Should Apply
We're looking for curious, adversarial thinkers who instinctively push systems to their breaking points. You should be comfortable working with structured taxonomies and producing reproducible artifacts. Prior experience in cybersecurity, adversarial ML, or creative probing (e.g., psychology, writing) is highly valued. If you thrive on varied projects and care about making AI safer, this role is for you.
Salary Insight
The role pays $48.00–$62.00 per hour, based on experience and specialization.
Location
Required Skills
Application Tip
In your application, include a concrete example of a vulnerability you uncovered (e.g., a prompt injection or jailbreak) and explain how you documented it for reproducibility. This demonstrates the structured adversarial thinking we value.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedAI Safety Experts — English & Danish
We're recruiting bilingual AI safety specialists (English & Danish) to join a red team that systematically probes conversational AI models for vulnerabilities. You'll generate adversarial data, surface hidden risks, and create actionable documentation to strengthen AI systems. This remote, hourly contract pays $48–$62/hr based on experience.

Mercor
VerifiedAI Safety Experts — English & Swedish
This role involves stress-testing AI systems by simulating adversarial attacks in English and Swedish. You'll generate critical safety data by probing models for vulnerabilities like bias, misinformation, and security flaws. As part of a human red team, you'll help ensure AI behaves safely before it reaches users.

Mercor
VerifiedAI Safety Experts — English & Finnish
We're assembling a specialized team to stress-test AI systems by simulating adversarial attacks. As an AI Safety Expert, you'll probe conversational models for vulnerabilities like bias, misinformation, and harmful behaviors — all while working remotely on an hourly basis. Native fluency in both English and Finnish is essential for this role.

Mercor
VerifiedAI Safety Experts — English & Bengali
This role brings fluent English and Bengali speakers into a remote AI red team. You'll probe conversational AI systems with adversarial inputs to uncover biases, misinformation, and manipulation risks. All work is text-based, and you'll produce structured red team data that improves AI safety. You can opt out of higher-sensitivity topics, and clear guidelines plus wellness resources are provided.

Mercor
VerifiedAI Safety Experts — English & Malay
This remote role asks bilingual experts in English and Malay to stress-test AI systems by simulating attacks and probing for weak points. You'll generate high-quality red team data that helps customers build safer, more trustworthy models. All work is text-based, and sensitive topics are clearly disclosed with wellness support.

Mercor
VerifiedAI Safety Experts — English & Dutch
We're looking for bilingual English-Dutch AI red teamers to stress-test conversational AI models from every angle. In this remote, hourly role, you'll design and execute adversarial attacks—from jailbreaks to bias exploitation—and generate structured data that helps make AI systems safer. You'll work with clear guidelines and optional high-sensitivity projects, with wellness support available.

