
Mathematics Expert - AI Evaluation Project
Listing checked September 18, 2026 · pay as published by Handshake AI
Overview
Handshake AI seeks mathematicians and research scientists for part-time contract work that supports AI research. You will craft expert-level math problems drawn from real proof, formalization, and computational workflows. Then you assess model answers for technical accuracy and rigor, and send structured written feedback that helps the system reason more like a working mathematician. The fellowship runs remote and asynchronous, so you choose your hours and carry the work alongside another role.
What You'll Do5
- 1Design expert-level mathematics problems and scenarios that reflect genuine proof, formalization, and computation practice.
- 2Evaluate AI-generated responses for technical accuracy, mathematical rigor, and fit with how researchers solve problems.
- 3Write structured feedback that points out errors, gaps, and reasoning flaws in model outputs.
- 4Work at your own pace without a fixed schedule or minimum weekly hours.
- 5Use your sub-domain depth to pressure-test model reasoning on problems that fall outside current AI capability.
Requirements12
- 1PhD in Mathematics, pure or applied.
- 2First-author publications and a record of formalizing proofs that AI models struggle with.
- 3Depth in proof and formal methods such as Lean, Coq, or Isabelle.
- 4Depth in symbolic computation with SageMath or Mathematica, optimization, or numerical analysis.
- 5Depth in a core subfield such as algebra, topology, number theory, or combinatorics.
- 6Meeting two of the three main qualifications also qualifies you to apply.
- 7Familiarity with theorem provers and symbolic computation tools.
- 8Experience formalizing proofs or tackling problems at the edge of current AI ability.
- 9US work authorization.
- 10Strong written communication and attention to detail.
- 11Ability to explain complex mathematical concepts in clear terms.
- 12Ability to work on your own in a remote, asynchronous environment.
Who Should Apply
The strongest applicant holds a PhD in pure or applied mathematics and can point to first-author publications. They also formalize proofs in Lean, Coq, or Isabelle, or show depth in symbolic computation, optimization, numerical analysis, or a core subfield such as algebra or topology. The role suits someone who can explain a tricky proof in plain English and write precise, structured critiques of model output. The role fits less well if you lack US work authorization or have no hands-on proof formalization to show. Applicants often score low when their application lists math coursework but no first-author research, or when they cannot describe a specific problem where current AI models fail.
Salary Insight
Handshake AI lists compensation at $75.00 per hour. That figure fits the advanced profile in this listing: a PhD mathematician with first-author research and proof formalization skills. Pay for project-based AI training varies by domain and task, so compare the rate to your own consulting or research rate before you commit.
Pay and demand for Machine Learning & AI roles
AggregatedTypical pay
$75/hour
This role
$75/hr
Most Machine Learning & AI roles pay $55–$100 per hour. This role's pay falls inside that range.
Based on 513 similar roles that publish pay · 90 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 570
- Listed in last 30 days
- 252
- Remote
- 97%
Hiring most right now: micro1 (262) · Mercor (79) · SME Careers (60)
Most requested skills · share of roles
- python17%
- technical writing9%
- llm evaluation8%
- data annotation7%
Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$75/hr
Required Skills
Application Tip
Show proof formalization examples in Lean, Coq, or Isabelle and name a problem where current AI models fail, then link to a first-author publication or a public repository.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Handshake AI
VerifiedMathematics Expert (India) Part-Time Remote
Handshake AI offers this fellowship as an ongoing, part-time contract for mathematicians and research scientists who can improve model reasoning in pure and applied math. You judge AI answers for proof rigor, formalization, and computational correctness. The schedule stays open, so you can keep the project alongside a current academic or industry role. Contributions center on real mathematics workflows, from symbolic computation to theorem proving.

Handshake AI
VerifiedMath Expert AI Evaluation Fellowship Remote
Handshake AI recruits Math PhDs to sharpen how AI systems handle mathematical reasoning, proof construction, and technical problem-solving. The work centers on crafting difficult domain questions and checking AI-generated answers for accuracy, logical consistency, and mathematical rigor. Contributors join a remote, project-based fellowship with a schedule they set themselves. Most participants log 5 to 20 hours per week while a project is active, and no prior AI or technical experience is needed. Placement follows current project needs, with chances to join later projects as they open.

Turing
VerifiedMathematics Expert (Master’s/Ph.D.)
This role works on projects that assess and refine large language models using advanced mathematical reasoning. You will design challenging multi-step problems, solve them on your own, and translate proofs into the formal language of Lean. The job also involves writing reliable Python code for computational tasks, verifying numerical answers, and providing detailed feedback on model-generated solutions. A solid background in mathematics at the graduate or PhD level is essential, along with the ability to break down abstract ideas for non-experts. The engagement is a freelance contract that requires 20 to 40 hours per week, with a 4-hour overlap with Pacific Time.

Micro1
VerifiedMathematics Expert
micro1 is seeking sharp Mathematics Experts for a high-impact project that involves training the next generation of AI systems. In this fully remote contract role, you'll put your deep knowledge of advanced mathematics to work — no prior AI experience needed. Your contributions will directly influence how AI models reason and solve problems through high-quality, real-world data input.

Mercor
VerifiedMathematics Expert - Remote AI Training (Part-Time)
Mercor keeps a standing pool of mathematicians for AI training and evaluation projects rather than a single opening. Research teams at major AI labs need human judgment on math problems and proofs, and they rely on experts to build rubrics and grade model output. Work can arrive on short notice, so the pool stays open even when no project runs this week. Assignments cover tasks like writing research or competition-level problems, checking each proof step, reviewing edge cases, and spotting answers that look right but lack proof. Contracts in this field have paid $40 to $75 per hour, and Mercor sets each rate by project scope and depth.

Micro1
VerifiedMathematics Expert for AI Training Data Remote Contractor
A remote contractor role for a Mathematics Expert who will contribute to a customer project by shaping AI training data through rigorous mathematical input. You’ll author and review advanced problems, proofs, and explanations across calculus and related domains, and provide precise feedback to ensure clarity and accuracy. Your analyses and justifications should reflect research-level rigor and be suitable for diverse audiences.


