LLM Red-Teamer | $40-$65/hr Remote
Overview
micro1 is looking for sharp critical thinkers to help push the limits of frontier language models. In this role, you'll design tricky, adversarial multi-turn conversations and evaluate how well AI systems handle them. Your unique domain knowledge matters more than prior AI experience — we want people who can think like a hacker and write with precision.
What You'll Do7
- 1Craft complex, adversarial multi-turn dialogues and task-based scenarios that follow detailed project guidelines.
- 2Write clear, precise evaluation rubrics to judge model outputs against specific behavioral targets.
- 3Repeatedly test your scenarios against top-tier LLMs, increasing difficulty until you hit the quality bar.
- 4Deliver complete task packages including transcripts, target behaviors, binary rubrics, and supporting evidence.
- 5Review and document model strengths and failure modes relative to the project specs.
- 6Stay aligned with team leads and quality control as requirements shift over time.
- 7Work independently to produce consistent, high-quality deliverables at a steady pace.
Requirements7
- 1Exceptional written English skills with clarity, precision, and strong structure.
- 2Prior experience in AI human data environments such as RLHF, SFT, evaluations, annotation, or prompt engineering.
- 3Deep familiarity with large language models and ability to anticipate common failure patterns.
- 4Proven ability to work autonomously, interpreting and executing complex specs with minimal oversight.
- 5Strong critical thinking and meticulous attention to detail.
- 6Experience designing evaluation items or rubrics is a plus.
- 7Background in writing-intensive or analysis-focused fields like research, editorial, technical writing, or QA is beneficial.
Who Should Apply
This role is ideal if you enjoy stress-testing systems, have a knack for spotting subtle flaws, and can express your reasoning with precision. You're a self-starter who thrives on independent work and can consistently meet output targets. Whether you come from a writing, research, or QA background, your ability to create challenging scenarios and evaluate responses is key.
Salary Insight
Compensation ranges from $40 to $65 per hour, paid on a per-task basis (output-based). Minimum weekly submissions are required, and you should be ready to start within 24–48 hours of onboarding.
Required Skills
Application Tip
When applying, include a sample of an adversarial prompt or rubric you've designed to demonstrate your approach to testing AI models. That will set you apart from other candidates.
Similar open positions
Explore active roles that match your skills and interests.
Micro1
VerifiedAI Evaluation Analyst | $20-$30/hr Remote
As an AI Evaluation Analyst at micro1, you'll help train advanced language models by creating high-quality evaluation data and multi-turn conversations. This remote contract role is ideal for someone with strong analytical and writing skills who wants to directly influence how AI systems reason and interact. No previous AI experience is required—your domain expertise and attention to detail are what matter.
Micro1
VerifiedAI Jailbreak & Prompt-Injection Security Expert | $50-$90/hr Remote
micro1 is looking for a contractor to join a forward-thinking client project focused on hardening AI systems against exploitation. Your mission: design creative adversarial tests—think ethical jailbreaks, prompt injection, and tool-use abuse—to uncover weaknesses in modern LLMs. No deep AI background is required; your hands-on security expertise and adversarial mindset are what count.
Mercor
VerifiedMachine Learning & NLP Expert | $80-$110/hr Remote
Join a top-tier AI lab's generative AI team as a part-time remote Machine Learning and NLP Expert. You'll craft complex, real-world tasks to test and improve frontier models, author reference solutions, and analyze model performance. This role requires hands-on Python proficiency and deep expertise in modern ML and NLP methods, offering flexible remote work at approximately 20 hours per week.
Mercor
VerifiedLLM Research Scientist (Pre-training & Post-Training) | $100-$120/hr Remote
We're looking for experienced machine learning researchers who have hands-on expertise in training and improving large language models from end to end. In this role, you'll tackle clearly defined, open-ended empirical research problems related to LLM pre-training and post-training, working remotely on a flexible hourly basis. It's an opportunity to push the boundaries of foundation model research while collaborating with top AI researchers.
Micro1
VerifiedLitigator / Practicing Attorney (US) | $100-$150/hr Remote
We're seeking experienced litigators and practicing attorneys to apply their legal expertise to a pioneering project at the intersection of law and artificial intelligence. In this remote contract role, you'll help train next-generation AI systems by evaluating and refining AI-generated legal content — no prior AI knowledge is required. Your domain expertise and analytical rigor are the core assets here.
Mercor
VerifiedLegal contracts / diligence / redlines Evaluator | $80-$120/hr Remote
We're seeking seasoned legal professionals with deep expertise in contract review, due diligence, and redlining to evaluate AI-generated legal documents. In this remote, hourly role, you'll apply your subject-matter knowledge to assess the accuracy and quality of AI outputs, providing detailed feedback to help refine the model. It's a unique opportunity to shape cutting-edge AI while leveraging your legal skills.