
AI Jailbreak & Prompt-Injection Security Expert | $50-$90/hr Remote
Overview
micro1 is looking for a contractor to join a forward-thinking client project focused on hardening AI systems against exploitation. Your mission: design creative adversarial tests—think ethical jailbreaks, prompt injection, and tool-use abuse—to uncover weaknesses in modern LLMs. No deep AI background is required; your hands-on security expertise and adversarial mindset are what count.
What You'll Do6
- 1Build and execute advanced testing strategies for AI safety, including multi-turn jailbreak attempts, prompt injection attacks, and abuse of model tool-calling capabilities.
- 2Develop cross-domain elicitation methods that probe complex, chained adversarial behaviors across different contexts.
- 3Create and maintain regression test suites that systematically check models for known and novel jailbreak or injection vulnerabilities.
- 4Construct evaluation frameworks that put LLMs under realistic stress scenarios to measure and improve their robustness.
- 5Work with technical stakeholders to turn your findings into concrete safety improvements and mitigations.
- 6Document your processes, results, and recommended practices in clear reports and presentations for both engineering and non-technical audiences.
Requirements7
- 1At least 2 years of experience in adversarial machine learning, LLM red teaming, AI safety assessment, or a closely related security specialty.
- 2A proven track record of discovering or testing vulnerabilities linked to ethical jailbreaks, prompt injection, tool-use abuse, or other adversarial AI attacks.
- 3An advanced degree (PhD or MS) in computer science, cybersecurity, machine learning, or a similar field—or equivalent professional credentials.
- 4Credibility within the AI security community, evidenced by published research, open-source tools, talks at conferences, or recognized participation in bug bounty programs.
- 5Strong written and verbal communication skills, with an emphasis on precise documentation and collaborative problem-solving.
- 6Experience working on multi-disciplinary or cross-functional safety initiatives is a plus.
- 7Familiarity with current LLM architectures, prompt engineering techniques, and security testing tools is highly desirable.
Who Should Apply
This role is ideal for a security researcher or red-teamer who thrives on breaking things to make them safer. You have a hacker’s curiosity and a methodical approach to finding edge cases in AI behavior. You enjoy documenting your exploits clearly and collaborating with engineers to turn insights into defenses.
Salary Insight
$50–$90 per hour, based on experience and expertise.
Location
Required Skills
Application Tip
Include concrete examples of jailbreaks or prompt injections you’ve successfully executed—describe the technique, the model, and what you uncovered. A short write-up or a link to a published proof of concept will set you apart.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedAI Safety Experts — English & Malay
This remote role asks bilingual experts in English and Malay to stress-test AI systems by simulating attacks and probing for weak points. You'll generate high-quality red team data that helps customers build safer, more trustworthy models. All work is text-based, and sensitive topics are clearly disclosed with wellness support.

Mercor
VerifiedAI Safety Experts — English & Norwegian
This remote role invites fluent English and Norwegian speakers to join an elite red team probing AI models for vulnerabilities. You'll simulate adversarial attacks, document exploits, and generate high-quality human data that directly strengthens AI safety. The work is text-based and focuses on sensitive topics like bias and misinformation, with optional high-sensitivity projects supported by clear guidelines.

Mercor
VerifiedAI Safety Experts — English & Dutch
We're looking for bilingual English-Dutch AI red teamers to stress-test conversational AI models from every angle. In this remote, hourly role, you'll design and execute adversarial attacks—from jailbreaks to bias exploitation—and generate structured data that helps make AI systems safer. You'll work with clear guidelines and optional high-sensitivity projects, with wellness support available.

Mercor
VerifiedAI Safety Experts — English & Finnish
We're assembling a specialized team to stress-test AI systems by simulating adversarial attacks. As an AI Safety Expert, you'll probe conversational models for vulnerabilities like bias, misinformation, and harmful behaviors — all while working remotely on an hourly basis. Native fluency in both English and Finnish is essential for this role.

Mercor
VerifiedAI Safety Experts — English & Vietnamese
We're hiring bilingual (English & Vietnamese) red teamers to stress-test AI models by probing for vulnerabilities such as jailbreaks, bias, and misinformation. In this remote, contract role, you'll generate adversarial data and document findings to help make AI systems safer. It's a fit for anyone with a knack for pushing systems to their limits and a structured approach to testing.

Mercor
VerifiedAI Safety Experts — English & Assamese
Mercor is assembling a remote red team of human experts to attack AI models with adversarial inputs. This role focuses on conversational AI, using native English and Assamese to probe for jailbreaks, prompt injections, and bias exploitation. You'll generate red team data and vulnerability reports that make AI systems safer for customers. The work is entirely text-based, with optional participation in higher-sensitivity projects supported by clear guidelines and wellness resources. Hourly compensation is $20-$22.

