AI Evaluation Analyst | $20-$30/hr Remote
Overview
As an AI Evaluation Analyst at micro1, you'll help train advanced language models by creating high-quality evaluation data and multi-turn conversations. This remote contract role is ideal for someone with strong analytical and writing skills who wants to directly influence how AI systems reason and interact. No previous AI experience is required—your domain expertise and attention to detail are what matter.
What You'll Do6
- 1Design complex, multi-turn conversation scenarios and scoring rubrics that align with project specifications.
- 2Run conversation drafts through state-of-the-art LLMs to test their quality, then iterate for higher difficulty and accuracy.
- 3Deliver complete evaluation packages including conversation transcripts, expected model behaviors, binary rubrics, and supporting evidence.
- 4Adhere closely to evolving project guidelines while maintaining high throughput and meticulous attention to detail.
- 5Collaborate with team leads and quality control to calibrate assessments as specifications change.
- 6Work autonomously and consistently, meeting weekly deliverable targets without constant supervision.
Requirements7
- 1Near-native command of written English with exceptional clarity, structure, and precision.
- 2Prior experience in data annotation, RLHF, SFT, evaluation, or prompt engineering for AI systems.
- 3Familiarity with how leading LLMs behave and their common failure patterns.
- 4Ability to independently follow complex, detailed guidelines with minimal oversight.
- 5Strong critical thinking and analytical skills, particularly in writing-heavy or analysis-intensive domains.
- 6Experience creating evaluation items, rubrics, or conducting deep analysis of technology outputs.
- 7A background in research, editing, technical writing, or quality assurance is a plus.
Who Should Apply
This role is perfect for detail-oriented individuals with excellent written communication skills—think editors, researchers, technical writers, or QA professionals looking to apply their expertise at the cutting edge of AI. You should enjoy working independently on structured tasks and be comfortable with iterative refinement. No AI background is needed; your ability to follow precise specs and produce clear, thoughtful outputs is what counts.
Salary Insight
Compensation is output-based, targeting $20-$30 per hour equivalent. You are paid per task that meets project specifications, with a minimum weekly submission requirement. The actual time per task varies based on experience and workflow.
Required Skills
Application Tip
To stand out, submit a short cover letter or note that demonstrates your ability to follow detailed instructions—for example, precisely format your application as requested or highlight a past project where you adhered to complex guidelines. This directly shows the spec-fidelity skills the role demands.
Similar open positions
Explore active roles that match your skills and interests.
Micro1
VerifiedAI Evaluation Specialist | $20-$35/hr Remote
As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.
Micro1
VerifiedLLM Red-Teamer | $40-$65/hr Remote
micro1 is looking for sharp critical thinkers to help push the limits of frontier language models. In this role, you'll design tricky, adversarial multi-turn conversations and evaluate how well AI systems handle them. Your unique domain knowledge matters more than prior AI experience — we want people who can think like a hacker and write with precision.
Micro1
VerifiedImage Evaluation Generalist | $20-$30/hr Remote
Micro1 is looking for sharp-eyed contractors to join a remote team evaluating images that will help train the next wave of AI systems. In this role, you'll apply your own expertise to spot inconsistencies and provide clear feedback — no AI background needed. Your assessments will directly influence how models learn to interpret visual data.
Micro1
VerifiedEnglish Speaking Generalist | $20-$30/hr Remote
Join a remote contract role at micro1, where you’ll help train next-generation AI systems by evaluating and comparing responses from models like ChatGPT and Claude. As an English Speaking Generalist, you’ll create diverse prompts, record your screen and voice, and provide detailed verbal feedback—all while using your domain expertise, not prior AI experience.
Mercor
VerifiedData analysis / quantitative readouts Evaluator | $80-$120/hr Remote
We're looking for seasoned data analysts to evaluate AI-generated reports, spreadsheets, and slide decks. You'll use your quantitative expertise to assess accuracy, rigor, and presentation quality, then deliver structured feedback. This is a remote, hourly role with a competitive rate.
Micro1
VerifiedAI Image & Video Evaluation Specialist (PT, MT and CT) | $30-$40/hr Remote
micro1 is looking for a sharp visual specialist to help train next-generation generative AI models. In this remote contract role, you'll generate and compare image and video outputs from multiple platforms like ChatGPT/Sora and Gemini, providing structured feedback on realism, composition, and coherence. Your expertise in visual media — not AI experience — is what matters most.