
Chemistry AI Red-Teaming Expert (Remote)
Listing checked September 15, 2026 · pay as published by Mercor
Overview
Neon seeks chemistry domain experts to red-team frontier AI models. The core test asks whether a model can separate a safe technical question from a harmful one. You will design single-turn prompts tagged benign, dual-use, or adversarial, then check how the model answers against a policy standard. The job hinges on calibration. A harmless prompt that the model rejects matters as much as a harmful prompt it approves.
What You'll Do6
- 1Write hard single-turn prompts inside your chemistry specialty and assign each one a label: benign, dual-use, or adversarial.
- 2Review model replies against the agreed policy standard and decide if each response meets that standard.
- 3Draft the reference answer that shows what a correct reply should include.
- 4Explain the technical reasoning behind your reference answer so a non-specialist can follow it.
- 5Flag cases where the model refuses a safe request or approves a dangerous one.
- 6Keep calibration central by treating false refusals and false approvals with equal weight.
Requirements6
- 1Hands-on experience in synthesis, analytical chemistry, or chemical safety.
- 2Background as a chemical defense researcher, synthetic or process chemist, analytical chemist, industrial hygienist, process safety engineer, forensic chemist, or toxicological chemist.
- 3The team prefers prior AI evaluation or red-teaming work, but does not require it.
- 4Strong technical writing. Every judgment needs a written rationale that a non-specialist can follow.
- 5Published research, expert witness work, or prior technical writing is a strong signal. Include a sample or link.
- 6You must not draw on classified or export-controlled information, material under an NDA, or content under prepublication review.
Who Should Apply
The ideal candidate has bench or field experience in synthesis, analytical chemistry, or chemical safety, plus a track record of writing clear technical rationales. Chemists who have run reactions, assessed exposure risks, or identified hazardous compounds will calibrate prompts with the right level of detail. People with no hands-on chemistry background, or those who cannot explain a technical call in plain language, tend to score low here. Applications also fall short when the writing sample is missing or when the prompts are either too easy or too far from real laboratory practice. This work is not a fit for someone looking for coding tasks or general AI training data work.
Salary Insight
The source lists a one-time payment of $75 for this task. That is a fixed fee, not an hourly or recurring rate, so the total reflects a single engagement rather than ongoing work. Treat it as compensation for the prompt set and evaluations you complete.
Pay and demand for Cybersecurity roles
AggregatedTypical pay
$65/hour
This role
$75 fixed
Most Cybersecurity roles pay $40–$85 per hour.
Based on 63 similar roles that publish pay · 6 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 81
- Listed in last 30 days
- 38
- Remote
- 96%
Hiring most right now: Neon (32) · micro1 (23) · Mercor (6)
Most requested skills · share of roles
- technical writing27%
- red teaming27%
- penetration testing23%
- reverse engineering22%
Figures from Cybersecurity roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$75 fixed
Required Skills
Application Tip
Attach a short writing sample where you explain a chemistry judgment to a non-specialist, then name your exact subdomain and any AI evaluation or red-teaming projects you have handled. That mix raises fit because the team scores calibration and written rationale ahead of general chemistry knowledge.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedAnalytical Chemist - Remote AI Red-Teaming
Chemistry and chemical safety specialists will join a red-team panel that probes how well frontier AI models handle misuse risk. The task centers on single-turn prompts written at three labels: benign, dual-use, and adversarial, and on judging each prompt against a policy standard. You also draft reference answers that show the correct response and explain the technical reasoning. Relevant method areas include GC-MS, LC-MS/MS, IR, and NMR, plus trace-level validation and unknown identification.

Mercor
VerifiedForensic Toxicological Chemist - AI Red Team
Mercor is assembling a panel of chemistry and chemical safety experts to red-team frontier AI models. Your core task is to write challenging single-turn prompts in your domain, labeled as benign, dual-use, or adversarial. You then evaluate the model's responses against a defined policy standard and judge whether each was handled correctly. Another key duty is writing the reference answer, including the technical reasoning for why a response is right or wrong. The dual-use nature of this field means you must tell a routine professional question from one that is fishing for harmful information. A model that refuses a legitimate identification question fails as much as one that answers a dangerous one.

Mercor
VerifiedChemical Defense Researcher - Remote Contract
Mercor is assembling a panel of chemistry and chemical safety specialists to red-team frontier AI models. You will write single-turn prompts in your domain, label them as benign, dual-use, or adversarial, and judge model answers against a policy standard. The work sits on the dual-use line in chemical defense, where the same toxicology that sets a countermeasure dose can also describe an exposure someone might intend. Practitioners with CWC verification, decontamination, or medical countermeasure experience can separate routine professional questions from probes for harmful detail. This remote, task-based engagement pays a one-time $75 fee.

Mercor
VerifiedEnergetic Materials Red Teaming Expert (Remote)
The panel tests frontier AI models on how they handle explosive and energetic materials requests. You write single-turn prompts split into benign, dual-use, and adversarial levels, then judge model replies against a fixed policy. Your reference answers show what a sound response looks like and explain the technical reasoning behind that call. The work matters because models both refuse routine mining or investigation questions and answer dangerous requests dressed as academic work. Energetic materials, red teaming, and AI evaluation sit at the center of the task.

Mercor
VerifiedSynthetic / Process Chemist - Remote Task-Based
A panel of process chemistry and synthetic chemistry experts is forming to red-team frontier AI models in a task-based remote engagement. The group checks whether a model can separate legitimate technical questions from dangerous ones, with prompts and answers sorted as benign, dual-use, or adversarial. You will write single-turn prompts, judge model responses against a policy standard, and draft reference answers that include the technical reasoning. The domain sits on the dual-use line: the same scale-up knowledge that turns a medicine into a product can also support a controlled compound, and an impurity control profile can reveal how a substance was made. Pay is a one-time $75 fee.

Mercor
VerifiedRadiological Safety AI Red Team Expert (Remote)
Neon seeks radiological safety practitioners to red-team frontier AI models on a task-based engagement. You will craft single-turn prompts that sit on the dual-use line, label each one as benign, dual-use, or adversarial, then judge how the model handles the request against a set policy. The work matters because a useful model must answer routine dose, shielding, or isotope questions while refusing requests that seek misuse details. Your domain knowledge decides where that boundary falls, and your writing makes the call clear to people outside radiological safety.


