
Research Evaluation Specialist (PhD/Researcher)
Listing checked September 8, 2026 · pay as published by Micro1
Overview
This remote contractor role puts your research expertise to work in a new way. You will craft challenging questions and answers that test the limits of advanced AI models, helping them reason more effectively. Your domain knowledge is the core asset here, and no prior AI experience is needed. You will rely on primary sources, document your reasoning, and refine your work based on feedback to meet strict quality standards.
What You'll Do6
- 1Develop original, high-difficulty question-and-answer pairs that probe the boundaries of your field and challenge AI systems.
- 2Locate and verify information using authoritative references, documenting each answer with clear citations and logical reasoning.
- 3Design questions that require advanced reasoning, methodological awareness, or synthesis across multiple sources, avoiding simple or shortcut solutions.
- 4Test your questions against AI models, spot those that are too easy, and adjust them to increase difficulty while keeping accuracy intact.
- 5Write with precision and clarity, ensuring every question and answer is unambiguous and defensible.
- 6Incorporate reviewer feedback and follow all project guidelines to maintain consistent quality across your submissions.
Requirements7
- 1A completed PhD, active PhD candidacy, or equivalent research experience as a specialist, researcher, or professor.
- 2A solid record of scholarly work, deep subject matter expertise, and comfort with primary literature.
- 3Strong analytical thinking, meticulous attention to detail, and excellent written English.
- 4Proven ability to source and cross-check information from authoritative references.
- 5Self-direction and reliability in delivering expert-level output independently in a remote setup.
- 6Skill in composing original, methodologically sound, and challenging questions.
- 7Experience with AI training or evaluation is a plus but not required.
Who Should Apply
This role suits researchers and academics who enjoy the puzzle of crafting tough questions and verifying answers against primary sources. If you have a PhD or equivalent research depth and can work independently, you will fit well. It is less ideal for those who prefer structured team environments or lack the patience for meticulous citation work. Candidates often miss out when their questions are too straightforward or their sourcing is sloppy, so bring your sharpest analytical edge.
Salary Insight
Compensation is output-based, with pay per task that meets project specs. The rate is not fixed in the listing, but the range mentioned is $40-$90 per hour. Actual earnings depend on your speed and the number of tasks you complete. Minimum weekly submissions apply, and pay details are discussed after you apply.
Pay and demand for Medical Research & Public Health roles
AggregatedTypical pay
$65/hour
This role
$40–$90/hr
Most Medical Research & Public Health roles pay $36–$84 per hour. This role's pay falls inside that range.
Based on 255 similar roles that publish pay · 106 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 299
- Listed in last 30 days
- 197
- Remote
- 99%
Hiring most right now: SME Careers (108) · micro1 (71) · Mercor (39)
Most requested skills · share of roles
- llm evaluation30%
- ai training30%
- trainer feedback26%
- documentation22%
Figures from Medical Research & Public Health roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$40–90/hr
Required Skills
Application Tip
In your application, highlight a sample question you have designed that demonstrates your ability to create high-difficulty, source-backed content. Show how you would verify the answer using primary literature.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedEvaluation Specialist for AI Training and Research
Remote contractor role for Evaluation Specialists or Recent Grads who will help train next‑generation AI systems. You will craft original QA pairs, perform rigorous source triangulation, and create multi‑step questions that require synthesis. The project emphasizes high-quality, real‑world input to influence how models learn, reason, and perform. Prior AI experience isn’t required; deep domain knowledge matters and it’s welcomed.

Micro1
VerifiedAI Evaluation Specialist
As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.

Micro1
VerifiedComputer Science Expert PhD for AI Training Remote Role
A remote contract opportunity for a PhD-level computer science expert to contribute specialized knowledge to an AI training initiative. You’ll provide authoritative explanations and review complex CS data to improve model learning, reasoning, and performance. No prior AI experience is required beyond deep domain mastery. Expect to produce high-quality, clearly justified responses and collaborate with annotation leads and project managers.

AfterQuery
VerifiedResearch Scientist - Formal Methods (Remote)
AfterQuery runs a research cohort of senior domain experts who probe how frontier AI models handle hard, expert-level problems. The STEM track needs people active in formal methods and computational science, whether that means theorem proving in Lean or Coq, genomics and computational biology, or condensed-matter and quantum physics. The work leans toward evaluation and research rather than production engineering: you define what a correct or excellent answer looks like inside your own specialty. Expect roughly 4 hours a week on a remote, asynchronous contract that fits beside a research post or an industry job.

AfterQuery
VerifiedPhD Domain Expert - Remote AI Research Contract
AfterQuery seeks PhD-level researchers for Project Iter, a remote contract engagement that runs on a project-by-project basis. You bring doctoral training in a quantitative, scientific, technical, or humanities field and apply it to applied research tasks, from literature reviews to experiment design. Assignments span 2 to 3 weeks and call for about 10 hours per week, with some projects requiring 10 to 20 hours. Work is asynchronous, so you choose when to contribute while supporting model evaluation, data curation, and benchmarking. Pay ranges from $120 to $180 per hour, and the role offers direct exposure to an early-stage Y Combinator-backed company.

Mercor
VerifiedRemote Survey Research Expert for AI Evaluation
Mercor is connecting you with a leading AI research organization that needs seasoned survey professionals to evaluate its surveys for methodological quality. You will examine question wording, response scales, and sampling approaches, and articulate why certain designs yield unreliable or biased data. This remote, hourly role suits practitioners from market research, consumer insights, behavioral science, and related fields who have designed and fielded surveys in real projects. Your critical eye will help improve how AI systems think about research design and evidence quality. Candidates with hands-on experience at research agencies or internal insights teams will feel at home.


