
Research Scientist - Formal Methods (Remote)
Listing checked September 22, 2026 · pay as published by AfterQuery
Overview
AfterQuery runs a research cohort of senior domain experts who probe how frontier AI models handle hard, expert-level problems. The STEM track needs people active in formal methods and computational science, whether that means theorem proving in Lean or Coq, genomics and computational biology, or condensed-matter and quantum physics. The work leans toward evaluation and research rather than production engineering: you define what a correct or excellent answer looks like inside your own specialty. Expect roughly 4 hours a week on a remote, asynchronous contract that fits beside a research post or an industry job.
What You'll Do6
- 1Build problem sets and realistic technical scenarios that match the difficulty of your track.
- 2Write expert reference solutions that show the intended reasoning, not only the final answer.
- 3Draft grading rubrics that separate a partly correct response from a genuinely expert one.
- 4Review AI-generated answers for correctness, depth of reasoning, and domain judgment.
- 5Record the reasoning behind each grading call so researchers and other reviewers can follow your standard.
- 6Flag responses that read as confident but rest on faulty assumptions or skipped steps.
Requirements7
- 1Hands-on work in your field, now or within the recent past.
- 2Strong technical writing: you produce explanations and structured feedback, not just answers.
- 3Ability to work alone on a remote, asynchronous schedule with little back-and-forth.
- 4Meets the experience and education bar described in the full listing.
- 5Publication record, open-source contributions, or another recognized body of work in your domain (preferred).
- 6Past experience reviewing or grading other people's technical work, such as peer review, code review, or editing (preferred).
- 7Bonus signal for formal methods and theorem proving with Lean or Coq, computational biology or genomics, or condensed-matter and quantum physics.
Who Should Apply
The best fit is a working researcher or practitioner who can point to concrete output in one of these tracks: a Lean or Coq proof library, a genomics pipeline, a physics paper, a chemistry model. You need to enjoy writing about your subject, because a large share of the job is authoring solutions, rubrics, and feedback rather than solving problems for their own sake. People who want to train or fine-tune models, or who need frequent direction from a manager, will find this role misaligned. Applications score low when the domain claim stays vague, for example naming a field without naming what you built or published in it, or when the writing sample reads thin next to the technical depth claimed on the resume.
Salary Insight
AfterQuery lists $100 to $170 per hour for this contract, with the rate set by experience. At four hours a week that is roughly $400 to $680 weekly, so the role works best as a supplement to a main position rather than a replacement. Candidates with a publication record plus review or grading experience tend to land nearer the top of the band.
Pay and demand for Medical Research & Public Health roles
AggregatedTypical pay
$75/hour
This role
$100–$170/hr
Most Medical Research & Public Health roles pay $55–$100 per hour. This role's pay sits above that range.
Based on 545 similar roles that publish pay · 91 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 618
- Listed in last 30 days
- 288
- Remote
- 98%
Hiring most right now: micro1 (278) · Mercor (92) · SME Careers (64)
Most requested skills · share of roles
- python16%
- technical writing8%
- llm evaluation8%
- ai evaluation7%
Figures from Medical Research & Public Health roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$100–170/hr
Required Skills
Application Tip
Lead with one concrete artifact from your track, a Lean or Coq proof, a published paper, or a repository, and add two or three sentences on how you would grade a flawed solution in that area. That combination of domain proof and grading judgment maps straight onto the review, rubric, and evaluation work this contract asks for.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

AfterQuery
VerifiedScientific Computing Research Expert (Remote)
AfterQuery hands working scientists a research writing job built on their own computational projects. You take a problem you already know well and shape it into an AI evaluation task that measures what a model truly understands. Every task also needs the metrics that separate a sound solution from a shallow one, while the exact passing bar stays private. Python covers anything you need to run or debug yourself. Researchers in the life, physical, and social sciences, along with mathematicians and engineers, all fit the brief.

AfterQuery
VerifiedSenior Domain Expert - AI Research Evaluation
AfterQuery builds a cohort of senior specialists across seven tracks: software and systems engineering, formal methods and computational science, ML inference and GPU kernels, enterprise operations, security, hardware design, and creative technology. The work centers on judging how frontier AI models handle hard expert-level problems, not on shipping production code. You define what counts as correct and what counts as excellent inside your domain, then turn that standard into scenarios, reference answers, and grading rubrics. The engagement runs remote and asynchronous on a contract basis, with about 4 hours of commitment each week and pay from $100.00 to $170.00 per hour.

Mercor
VerifiedScience Research Expert - AI Evaluation (Remote)
Mercor builds pools of scientific experts who help frontier AI teams judge model work in domains those teams cannot cover on their own. This listing is a standing application for part-time, remote research work, not a single opening. Accepted experts join a science pool, and Mercor invites matching people to specific projects that name the client, rate, hours, and hiring decision. Projects in life, physical, social sciences, math, and policy research have paid $60 to $120 per hour, with scope and depth setting the exact rate. Your role centers on research design, statistical reasoning, and written judgments that become grading rubrics for AI evaluation.

Micro1
VerifiedResearch Evaluation Specialist (PhD/Researcher)
This remote contractor role puts your research expertise to work in a new way. You will craft challenging questions and answers that test the limits of advanced AI models, helping them reason more effectively. Your domain knowledge is the core asset here, and no prior AI experience is needed. You will rely on primary sources, document your reasoning, and refine your work based on feedback to meet strict quality standards.

Handshake AI
VerifiedMathematics Expert - AI Evaluation Project
Handshake AI seeks mathematicians and research scientists for part-time contract work that supports AI research. You will craft expert-level math problems drawn from real proof, formalization, and computational workflows. Then you assess model answers for technical accuracy and rigor, and send structured written feedback that helps the system reason more like a working mathematician. The fellowship runs remote and asynchronous, so you choose your hours and carry the work alongside another role.

Handshake AI
VerifiedMathematics Expert (India) Part-Time Remote
Handshake AI offers this fellowship as an ongoing, part-time contract for mathematicians and research scientists who can improve model reasoning in pure and applied math. You judge AI answers for proof rigor, formalization, and computational correctness. The schedule stays open, so you can keep the project alongside a current academic or industry role. Contributions center on real mathematics workflows, from symbolic computation to theorem proving.


