
AI Evaluation Specialist | $20-$35/hr Remote
Overview
As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.
What You'll Do5
- 1Create comprehensive evaluation exercises — including prompts, supporting files, and detailed grading rubrics — to test AI capabilities in real-world computer tasks.
- 2Define precise, unambiguous criteria that clearly separate successful task completion from failure across various administrative and workflow scenarios.
- 3Carefully observe and document AI agent behaviors, producing crisp, accurate summaries and reports in high-quality English.
- 4Refine tasks and rubrics based on feedback and team collaboration to ensure robust and fair benchmarking.
- 5Work across multiple domains and adapt evaluation frameworks as project needs evolve, sharing insights with the customer's team to drive continuous improvement.
Requirements6
- 1At least 3 years of experience in roles that demand written precision and structured thinking — for example, as a paralegal, executive assistant, analyst, librarian, technical writer, or QA specialist.
- 2Native or fluent English with proven ability to write succinct, specific, and unambiguous observations.
- 3Demonstrated experience designing or applying rubric-based evaluation, scoring against set criteria, or building structured assessment frameworks.
- 4Exceptional attention to detail, with a knack for catching subtle patterns or inconsistencies others might overlook.
- 5Strong computer literacy, including comfort with SaaS tools, file management, web browsers, and document editing platforms.
- 6Self-direction and the ability to take ownership of loosely defined projects and drive them to completion independently.
Who Should Apply
You're a precision-oriented professional who thrives on detail and clear communication. Whether you come from a background in legal, research, technical writing, or quality assurance, you have a talent for creating structured evaluations and documenting observations with clarity. You enjoy working autonomously on ambiguous challenges and are excited to contribute directly to improving AI systems.
Salary Insight
This is a contract role paying $20–$35 per hour, with flexibility for remote work.
Location
Required Skills
Application Tip
In your cover letter, briefly describe a specific rubric or scoring system you've designed or applied in a past role. Use an example that showcases your ability to define clear criteria and document nuanced findings — this directly mirrors the core work of the position.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedEvaluation Specialist for AI Training and Research
Remote contractor role for Evaluation Specialists or Recent Grads who will help train next‑generation AI systems. You will craft original QA pairs, perform rigorous source triangulation, and create multi‑step questions that require synthesis. The project emphasizes high-quality, real‑world input to influence how models learn, reason, and perform. Prior AI experience isn’t required; deep domain knowledge matters and it’s welcomed.

Micro1
VerifiedAi Consulting Domain Expert Focused on AI Output Evaluation
A remote contractor role focused on evaluating and enhancing AI outputs for real-world business use. You’ll help train next-generation systems by refining responses, writing and reviewing technical documentation, and improving prompts for large language models. Your domain knowledge in strategy and operations matters, even without prior AI experience. You’ll work with high-quality input to influence how models learn and perform.

Micro1
VerifiedAi Domain Expert Focused on Domain Knowledge and Evaluation
Remote, part-time contractor role focused on guiding AI systems through real-world input. You’ll review AI outputs for accuracy, craft prompts that test reasoning, and provide precise written feedback to improve model performance. Your domain knowledge matters most, even if you’re not required to have AI prior experience. You’ll join a global, distributed team and help shape how models learn and reason. Bolded areas reflect core competencies like data annotation, prompt engineering, and ethical awareness in AI.

Micro1
VerifiedAI Software Engineering Domain Expert
Remote, part-time contract work with micro1 where you apply deep software engineering knowledge to help train next-generation AI systems. You’ll review and polish AI-generated technical content, refine prompts, and judge model outputs using rubrics to ensure accuracy and quality. The role centers on drafting and editing technical docs, architecture plans, RFCs, and design specs that feed AI training, plus independent research and data annotation. Strong writing and a solid engineering background are essential, and you’ll collaborate asynchronously with project leads to meet deliverables.

Micro1
VerifiedData Science Domain Expert for AI Evaluation and Prompting
Remote contractor role focusing on AI data science with a domain expert lens. You’ll assess AI outputs, refine prompts, and annotate data to support high-quality model training. The work centers on applying deep domain knowledge to review research-style documents, technical reports, and experiment notes, shaping how models learn and reason. Strong writing, precise attention to detail, and independent research are essential as you contribute to rubric-based evaluations and content quality.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

