
Physician (MD/DO) - Remote Clinical AI Evaluation
Listing checked September 17, 2026 · pay as published by Mercor
Overview
Mercor needs residency-trained physicians for non-clinical projects that shape how clinical AI systems get measured. You will use your clinical background to build grading criteria, review clinical dialogues, and perform reasoning annotation. The role involves no patient care and no live diagnosis. Mercor runs a shared expert pool, so after onboarding you may join several workstreams at once based on your specialty and availability. Assignments shift as priorities change, and you can move between streams.
What You'll Do8
- 1Turn a clinical question into discrete, checkable criteria that let a reviewer grade model responses with consistency.
- 2Review multi-turn clinical conversations for accuracy, safety, completeness, proper hedging, and correct escalation advice.
- 3Record your diagnostic process for a case, including differentials you considered and ruled out, not just the final answer.
- 4Flag hallucinated findings, dangerous omissions, unsupported certainty, and advice that is correct in theory yet unsafe for patients.
- 5Define edge cases and standards of care for your specialty so annotation stays consistent across a large clinician group.
- 6Construct difficult clinical questions that push the limits of current model reasoning.
- 7Work through tasks that range from about 45 minutes for a dialogue evaluation to over an hour for grading-criteria authoring.
- 8Meet the specific throughput target set for the workstream you join.
Requirements12
- 1MD or DO with a completed residency in any specialty.
- 2Hold an active, unrestricted medical license in the country where you practice.
- 3Bring at least 2 years of clinical experience after residency, whether you practice now or have left clinical work.
- 4Write structured clinical rationale that a reviewer outside your specialty can follow.
- 5Demonstrate fluency in written and spoken English.
- 6Commit to a minimum of 20 hours per week and can concentrate those hours when a workstream has a deadline.
- 7(Preferred) Board certification in your specialty.
- 8(Preferred) U.S. licensure plus familiarity with U.S. standards of care and clinical guidelines.
- 9(Preferred) Background in primary care, internal medicine, emergency medicine, or hospitalist work.
- 10(Preferred) Prior experience with clinical annotation, AI evaluation, medical education, or question writing.
- 11(Preferred) Experience designing grading criteria, assessing residents, or developing clinical guidelines.
- 12(Preferred) Published research or sustained technical writing (link a sample).
Who Should Apply
The ideal candidate has finished a residency in any specialty, holds an active license, and brings at least two years of post-residency clinical experience. You should enjoy turning messy clinical thinking into clear, checkable steps that others can apply. If you want direct patient care, live diagnosis, or a traditional clinical setting, this role will not fit. Reviewers filter out applicants who lack a completed residency, hold no active unrestricted license, or cannot meet the 20-hour minimum each week. Clinical writing that a non-specialist cannot follow also lowers fit scores.
Salary Insight
This role pays $150.00 per hour. That rate sits in the upper range for remote non-clinical physician work, which aligns with the requirement for a completed residency, an active license, and at least two years of post-residency experience.
Pay and demand for Medical Research & Public Health roles
AggregatedTypical pay
$88/hour
This role
$150/hr
Most Medical Research & Public Health roles pay $70–$125 per hour. This role's pay sits above that range.
Based on 61 similar roles that publish pay · 5 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 79
- Listed in last 30 days
- 50
- Remote
- 99%
Hiring most right now: micro1 (26) · Mercor (24) · AfterQuery (8)
Most requested skills · share of roles
- ai evaluation14%
- clinical reasoning14%
- technical writing11%
- medical writing9%
Figures from Medical Research & Public Health roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$150/hr
Required Skills
Application Tip
When you apply, attach a short clinical writing sample that breaks a complex case into checkable steps for a non-specialist. State your post-residency years, your license status, and your availability each week. If you have experience with grading criteria, guideline authoring, or AI evaluation, name those projects and quantify the scale.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedMedical Expert (M.D./D.O.) - Remote Part-Time
Mercor keeps this listing open for physicians who want to shape how frontier AI models handle clinical reasoning. The work fits around your practice: you might write a clinical vignette from a real case, complete with presentation, labs, and imaging, then supply the attending-level answer used as a reference. Other projects ask you to grade model output on diagnostic reasoning, checking whether a differential follows a sensible order, whether the model anchors on a first impression, and whether the proposed workup matches guidelines. You also flag recommendations that read well in a textbook but fail the patient in front of you because of comorbidities, contraindications, or access limits. Projects start on short notice, so Mercor draws from this pool when client demand matches your specialty.

Mercor
VerifiedMedical Expert - Remote AI Evaluation & Review
Mercor recruits health professionals who can supply the medical judgment that top AI research teams cannot generate on their own. The medicine pool is an open call, not a single opening, so it remains live while projects start on short notice. When a project launches, you may write a case from your practice with the presentation, workup, constraints, and reference answer that sets the standard for grading a model's attempt. You may also review model output for clinical reasoning, checking whether the differential follows a sensible order, whether a guideline applies to the patient, and whether the care plan is safe to carry out. Mercor sets rates per project, and most work is remote and part-time.

Mercor
VerifiedMultilingual Outpatient Physician - Clinical AI Evaluation
Physicians who practice in outpatient or ambulatory clinics can shape how medical AI writes and evaluates clinical notes. Mercor seeks MD and DO clinicians for a healthcare AI partner that builds documentation and decision-support tools. You will review, annotate, and score real encounter notes in EHR systems, then write structured feedback for AI evaluation. The cohort is multilingual, so you need C1 or higher speaking, listening, and writing skills in English plus one of these languages: Czech, Catalan, Danish, Dutch, Vietnamese, or Finnish. Work happens remote and part-time, with a minimum of 10 hours each week.

Mercor
VerifiedDermatologist - Clinical AI Image Evaluation (Remote)
Board-certified dermatologists join Mercor to build and review clinical AI systems without seeing patients or issuing live diagnoses. The work centers on dermatological image interpretation, structured annotation, and AI output evaluation. You enter a shared expert pool after onboarding, then move among workstreams based on subspecialty, availability, and interest. Cases can take about 10 minutes for schema-based labeling or up to an hour for full case authoring, and each stream carries its own throughput target.

SME Careers
VerifiedMedical Doctor Specialist for AI Content Review (Remote)
Work as a Medical Doctor Specialist SME, remotely on an hourly contract to review AI-generated clinical responses and, when needed, author expert medical content. You will assess clinical reasoning, verify step-by-step problem solving, and ensure responses align with the prompt. Your input helps improve AI models by validating accuracy, safety, and adherence to evidence-based guidelines. Specialists across fields including Internal Medicine, Emergency Medicine, and Cardiology are welcome to join the expert network.

Turing
VerifiedResident Medical Specialist (MD/DO)
This remote role places licensed physicians at the center of AI evaluation for clinical reasoning. You will design test scenarios, assessment frameworks, and gap analyses for medical AI systems, collaborating directly with AI researchers. The engagement is flexible and can fit around your clinical schedule, requiring up to 30 hours per week over a one-month contract, with possible extensions based on performance. Your hands-on medical judgment will shape how AI models handle real-world diagnostic and management problems.


