Mercor
MercorVerified listing
Remote

Generalist AI Evaluation Expert (Remote)

40–70/hr
Remote
Posted September 16, 2026
part-time
Not Sure? Upload your resume to see every role you match

Listing checked September 16, 2026 · pay as published by Mercor

Overview

Mercor runs a standing pool for people who want part-time work in AI training and model evaluation. The team connects generalists to projects that need careful human judgment, from writing prompts and reference answers to grading how well a model handles general reasoning. Contracts pay $40 to $70 per hour and open on short notice, so the listing stays active even when no project is live. You apply once, complete a short interview, and join a network that clients draw from when a matching task appears.

What You'll Do7

  • 1Write prompts and reference answers for everyday questions with a correct response, plus the context a strong answer needs.
  • 2Grade model output on general reasoning, checking whether the answer is right, follows the instruction given, and avoids filler.
  • 3Mark cases where a model sounds sure but states something false or clashes with an earlier claim in the same answer.
  • 4Flag responses that answer a question the user did not ask.
  • 5Explain your verdicts in writing so project teams can follow your logic.
  • 6Call out unclear instructions or ambiguous prompts instead of guessing at intent.
  • 7Apply each project's rubric to every prompt and output you review.

Requirements5

  • 1Careful reading habits and sound judgment, backed by a professional or academic background in any field.
  • 2Clear written communication, since most tasks ask you to justify your reasoning.
  • 3Comfort with ambiguous tasks and a willingness to flag unclear instructions.
  • 4Mercor does not require a specific credential. What matters is that you can say why an answer is wrong.
  • 5Ability to confirm your work location and complete a short AI interview, about 20 minutes.

Who Should Apply

Generalists who read with care, write clear explanations, and can defend a judgment about a model's answer fit this pool well. The work suits people from any professional or academic field who want part-time AI evaluation projects and are open to ambiguous prompts. Candidates who want a single fixed project, a set schedule, or a decision on this listing will find a poor fit, because Mercor treats it as a standing pool and reaches out when a client project matches. Applicants often score low when they summarize an answer without naming the exact flaw, or when they ignore the instruction and grade against their own preferences. A missing resume or an incomplete interview also keeps you out of the matching pool.

Salary Insight

Most contracts in this field pay between $40 and $70 per hour, and the project scope and depth set the exact rate. This listing does not promise a set number of hours or a start date. You see the rate, hours, and client when Mercor invites you to a specific project.

Pay and demand for Machine Learning & AI roles

Aggregated

Typical pay

$70/hour

This role

$40–$70/hr

Most Machine Learning & AI roles pay $48–$98 per hour. This role's pay falls inside that range.

Based on 722 similar roles that publish pay · 129 publish only a top rate; those count at the rate they gave

Typical rangeMedian payThis role

Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.

Live similar roles
814
Listed in last 30 days
456
Remote
98%

Hiring most right now: micro1 (296) · Mercor (132) · AfterQuery (90)

Most requested skills · share of roles

  • python
    14%
  • data annotation
    9%
  • technical writing
    8%
  • ai evaluation
    7%

Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.

Compare your resume against these roles

Location

Typeremote
LocationRemote
This is a remote position

Compensation

$40–70/hr

Required Skills

ai trainingmodel evaluationllm evaluationprompt writingreference answer creationrubric applicationgeneral reasoninganswer gradingdata annotationquality assuranceinstruction followingambiguity handling

Application Tip

In your resume and AI interview, show how you judge a model answer: name the instruction it missed, the claim it got wrong, or the filler it added. Use concrete examples from any field and state why a response fails, since that reasoning is what Mercor grades for pool fit.

Share:

See NearSkill jobs more often in your search

Application & verification flow

  1. 1Instant rubric match

    Your resume is scanned against this role’s requirements to check qualification fit.

  2. 2Screened before the employer sees it

    Only profiles that clear screening are passed on.

  3. 3Outcome by email

    We notify you at the address on your resume once the screening is reviewed.

Test your fit score before applying

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

6d agoRemotepart-time

Science Research Expert - AI Evaluation (Remote)

Mercor builds pools of scientific experts who help frontier AI teams judge model work in domains those teams cannot cover on their own. This listing is a standing application for part-time, remote research work, not a single opening. Accepted experts join a science pool, and Mercor invites matching people to specific projects that name the client, rate, hours, and hiring decision. Projects in life, physical, social sciences, math, and policy research have paid $60 to $120 per hour, with scope and depth setting the exact rate. Your role centers on research design, statistical reasoning, and written judgments that become grading rubrics for AI evaluation.

60–120/hr
BiologyBiotechnologyChemistry+20 more
Mercor

Mercor

6d agoRemotepart-time

Education & Training Expert (Remote, Part-Time)

Educators who want part-time remote work in AI evaluation can join Mercor's standing pool for education and training experts. The listing stays open even when no project is active, and Mercor draws from it when new work appears. Past projects include writing classroom or course problems with the level, misconception, and constraints, then writing the reference answer that sets the grading standard for a model's attempt. Other tasks involve reviewing model-written explanations, lesson material, and assessments for accuracy and learner fit. Pay ranges from $30 to $85 per hour depending on the project and your background.

30–85/hr
EducationCurriculum DesignInstructional Design+14 more
Mercor

Mercor

6d agoRemotepart-time

Legal Expert - AI Evaluation and Review (Remote)

Mercor keeps an open pool for legal professionals who want part-time, remote work on AI evaluation and training. Research teams behind major AI models turn to that pool for legal judgment they cannot generate on their own, and your standards shape the rubric that grades model output. Assignments can involve writing a legal problem from a matter you handled, complete with facts, forum, governing law, and a reference answer. You might also review model-drafted provisions, privilege calls, or citations for accuracy and persuasive force. Projects open on short notice, so the listing stays live even when no matching engagement runs this week.

60–150/hr
Legal PracticeLitigationCriminal Justice+16 more
Mercor

Mercor

6d agoRemotepart-time

Data Analysis Expert - Remote AI Evaluation

Mercor maintains a standing pool for data analysis experts who want to train and evaluate frontier AI models. Research teams bring their hardest data judgment calls here, and your standards help build the rubrics that grade model answers. The pool spans business intelligence, data engineering, statistics, and machine learning, with part-time project work that starts on short notice. A matching project may arrive within a week or take several months, depending on client needs. The listing stays open even when no project is active.

70–120/hr
Data AnalysisBusiness IntelligenceData Engineering+17 more
Mercor

Mercor

6d agoRemotepart-time

Management Consultant - Remote AI Evaluation

Mercor keeps an open roster of management consultants who help grade and train frontier AI models. The work draws on your own engagement history: you write case prompts from real client situations, then create the reference answers models are measured against. You also review model-generated market sizings, profitability analyses, and driver trees, checking whether the logic holds up to partner-level scrutiny. Assignments arrive on short notice, so the pool stays active whether a project starts this week or later. Rates for past projects landed between $90 and $120 per hour, set by scope and depth.

90–120/hr
Management ConsultingStrategy ConsultingOperations Consulting+17 more
Mercor

Mercor

6d agoRemotepart-time

Medical Expert - Remote AI Evaluation & Review

Mercor recruits health professionals who can supply the medical judgment that top AI research teams cannot generate on their own. The medicine pool is an open call, not a single opening, so it remains live while projects start on short notice. When a project launches, you may write a case from your practice with the presentation, workup, constraints, and reference answer that sets the standard for grading a model's attempt. You may also review model output for clinical reasoning, checking whether the differential follows a sensible order, whether a guideline applies to the patient, and whether the care plan is safe to carry out. Mercor sets rates per project, and most work is remote and part-time.

60–180/hr
Internal MedicineSurgeryDiagnostic Medicine+20 more