
Science Research Expert - AI Evaluation (Remote)
Listing checked September 16, 2026 · pay as published by Mercor
Overview
Mercor builds pools of scientific experts who help frontier AI teams judge model work in domains those teams cannot cover on their own. This listing is a standing application for part-time, remote research work, not a single opening. Accepted experts join a science pool, and Mercor invites matching people to specific projects that name the client, rate, hours, and hiring decision. Projects in life, physical, social sciences, math, and policy research have paid $60 to $120 per hour, with scope and depth setting the exact rate. Your role centers on research design, statistical reasoning, and written judgments that become grading rubrics for AI evaluation.
What You'll Do7
- 1Write original research problems from your field, complete with the system, data, assumptions, and the reference answer used to grade a model's attempt.
- 2Assess model-written analysis for method fit, statistical support, and whether conclusions follow from evidence.
- 3Identify outputs that seem coherent but fail peer review because of study design, sample limits, or measurement constraints.
- 4Create or refine grading criteria that turn your scientific standards into a repeatable rubric.
- 5Explain your reasoning in writing so project leads and model teams can follow the calls you made.
- 6Flag unclear task instructions or missing context instead of guessing at what a client wants.
- 7Review model interpretations across disciplines such as biology, physics, economics, or political science.
Requirements7
- 1Graduate degree in a life, physical, or social science.
- 2Hands-on research experience where you designed studies and interpreted data yourself.
- 3Ability to explain methodological and statistical choices in clear written English.
- 4Comfort working with ambiguous prompts and incomplete specifications.
- 5Willingness to raise questions when instructions lack detail.
- 6Research background close enough in time that you can defend the decisions you made.
- 7Domain depth in at least one listed area, such as chemistry, astronomy, psychology, or policy research.
Who Should Apply
Researchers who hold a graduate degree in a science field and have designed studies, run analyses, and defended their interpretations in writing fit this pool best. The work suits people who can read a model's answer, spot weak statistical support, and explain why a conclusion does not follow. Candidates without graduate-level research training or without direct data interpretation experience tend to score low, even when they know a subject well from coursework. Applications also lose fit when candidates cannot point to specific studies they designed or when their writing leaves the reasoning behind a judgment unclear. The role is less suitable for people who want a fixed schedule, a defined project from day one, or tasks with complete instructions at every step.
Salary Insight
Projects have posted at $60 to $120 an hour, and Mercor sets the exact rate per project based on scope and depth. That range places the work in the expert AI evaluation tier, where pay tracks scientific domain knowledge and the complexity of the grading task. A specific project listing names the rate, hours, and client, so final compensation depends on the assignment you match with.
Pay and demand for Machine Learning & AI roles
AggregatedTypical pay
$75/hour
This role
$60–$120/hr
Most Machine Learning & AI roles pay $55–$100 per hour. This role's pay falls inside that range.
Based on 529 similar roles that publish pay · 91 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 595
- Listed in last 30 days
- 273
- Remote
- 97%
Hiring most right now: micro1 (269) · Mercor (84) · SME Careers (61)
Most requested skills · share of roles
- python17%
- technical writing8%
- llm evaluation8%
- data annotation7%
Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$60–120/hr
Required Skills
Application Tip
Apply with a resume that names your degree, field, and at least one study you designed, then complete the about 20-minute AI interview and join every network you qualify for. In your materials, call out the methods, statistical tests, and data interpretation calls you made, since those details help reviewers match you to projects that need your exact expertise.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedData Analysis Expert - Remote AI Evaluation
Mercor maintains a standing pool for data analysis experts who want to train and evaluate frontier AI models. Research teams bring their hardest data judgment calls here, and your standards help build the rubrics that grade model answers. The pool spans business intelligence, data engineering, statistics, and machine learning, with part-time project work that starts on short notice. A matching project may arrive within a week or take several months, depending on client needs. The listing stays open even when no project is active.

Mercor
VerifiedEngineering Expert - Remote AI Evaluation
Mercor supports research teams behind well-known AI models by supplying engineering judgment that models cannot generate on their own. Your design standards and analysis habits become the rubric used to grade a model's engineering output. The company keeps a standing, part-time listing for engineers who want AI evaluation and training work, not a single opening. Projects can start with short notice, so Mercor draws from this pool when a client needs a discipline such as structural engineering, mechanical engineering, or electrical engineering. Past engineering projects posted at $60.00 to $125.00 per hour.

Mercor
VerifiedRemote Survey Research Expert for AI Evaluation
Mercor is connecting you with a leading AI research organization that needs seasoned survey professionals to evaluate its surveys for methodological quality. You will examine question wording, response scales, and sampling approaches, and articulate why certain designs yield unreliable or biased data. This remote, hourly role suits practitioners from market research, consumer insights, behavioral science, and related fields who have designed and fielded surveys in real projects. Your critical eye will help improve how AI systems think about research design and evidence quality. Candidates with hands-on experience at research agencies or internal insights teams will feel at home.

Mercor
VerifiedGeneralist AI Evaluation Expert (Remote)
Mercor runs a standing pool for people who want part-time work in AI training and model evaluation. The team connects generalists to projects that need careful human judgment, from writing prompts and reference answers to grading how well a model handles general reasoning. Contracts pay $40 to $70 per hour and open on short notice, so the listing stays active even when no project is live. You apply once, complete a short interview, and join a network that clients draw from when a matching task appears.

Mercor
VerifiedSoftware Engineer Expert (Remote) Part-Time
Mercor keeps a standing pool of software engineers who help grade and train frontier AI models. The work is part-time and remote, and projects span frontend, backend, mobile, embedded, DevOps, QA, and security engineering. Assignments range from writing original engineering problems with reference answers to reviewing model-written code for edge cases, complexity, and production failure modes. Past software engineering projects on Mercor paid between $70 and $150 per hour, with scope and depth setting the exact rate.

Mercor
VerifiedMedical Expert - Remote AI Evaluation & Review
Mercor recruits health professionals who can supply the medical judgment that top AI research teams cannot generate on their own. The medicine pool is an open call, not a single opening, so it remains live while projects start on short notice. When a project launches, you may write a case from your practice with the presentation, workup, constraints, and reference answer that sets the standard for grading a model's attempt. You may also review model output for clinical reasoning, checking whether the differential follows a sensible order, whether a guideline applies to the patient, and whether the care plan is safe to carry out. Mercor sets rates per project, and most work is remote and part-time.


