
Chemical Defense Researcher - Remote Contract
Listing checked September 15, 2026 · pay as published by Mercor
Overview
Mercor is assembling a panel of chemistry and chemical safety specialists to red-team frontier AI models. You will write single-turn prompts in your domain, label them as benign, dual-use, or adversarial, and judge model answers against a policy standard. The work sits on the dual-use line in chemical defense, where the same toxicology that sets a countermeasure dose can also describe an exposure someone might intend. Practitioners with CWC verification, decontamination, or medical countermeasure experience can separate routine professional questions from probes for harmful detail. This remote, task-based engagement pays a one-time $75 fee.
What You'll Do8
- 1Write difficult single-turn prompts in your specialty, tagging each as benign, dual-use, or adversarial.
- 2Review model answers against a defined policy standard and decide if each answer was handled correctly.
- 3Draft reference answers that show what a correct response should include.
- 4Explain the technical reasoning behind each reference answer so a non-specialist can follow it.
- 5Identify borderline cases where a model should answer in full and where it must refuse.
- 6Document your rationale for every judgment in clear written form.
- 7Work through sustained periods of reading and writing about misuse scenarios in your field.
- 8Pause or step away without penalty when needed, as briefed in advance.
Requirements11
- 1Experience working chemical threats from the defensive side in an institutional setting.
- 2Background in one or more areas: countermeasures, protection, detection, or verification.
- 3Familiarity with toxicology of chemical warfare agents or medical countermeasure development.
- 4Knowledge of decontamination, protection, or detection research for chemical threats.
- 5CWC-related experience such as verification, destruction, or designated laboratory analysis.
- 6Ability to assess consequence and medical management of chemical exposure.
- 7Red-teaming experience preferred.
- 8Strong technical writing, published research, or expert witness work.
- 9Able to write rationales that non-specialists understand.
- 10Must not draw on classified or export-controlled information, or material under NDA or prepublication review.
- 11If you hold such obligations, disclose them so the team can scope the work accordingly.
Who Should Apply
The right candidate has worked chemical threats from the defensive side at an institution such as USAMRICD, DoD, CDC, or an allied national institute. You can judge whether a question about agent properties supports a treatment protocol or a threat assessment, and you can write that judgment for a non-specialist. This role is less suited to general chemists without defense, toxicology, or CWC experience, or to anyone expecting lab work instead of prompt writing and model evaluation. Applicants often score low when they submit no writing sample or link to published research, or when they cannot show red-teaming or policy-standard evaluation experience. Another common rejection reason is failing to disclose NDA, prepublication, or export-control obligations upfront.
Salary Insight
The pay is $75 for a one-time task, paid as a single engagement rather than an hourly or salary rate. That figure sits at the low end for specialized red-team prompt writing and expert model evaluation, where rates often climb with domain scarcity and writing depth. Expect the scope to match a short, defined task, so ask about expected hours before committing if you need to compare it to other work.
Location
Compensation
$75 fixed
Required Skills
Application Tip
Include a writing sample or link to published research with your application, because the role depends on rationales that a non-specialist can follow. In that sample, highlight a judgment call where you separated a routine countermeasure or PPE question from a misuse probe, and name any relevant frameworks such as CWC verification or decontamination protocols. Also state upfront whether you hold NDA, prepublication, or export-control obligations so the team can scope the task correctly.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedChemistry AI Red-Teaming Expert (Remote)
Neon seeks chemistry domain experts to red-team frontier AI models. The core test asks whether a model can separate a safe technical question from a harmful one. You will design single-turn prompts tagged benign, dual-use, or adversarial, then check how the model answers against a policy standard. The job hinges on calibration. A harmless prompt that the model rejects matters as much as a harmful prompt it approves.

Mercor
VerifiedAnalytical Chemist - Remote AI Red-Teaming
Chemistry and chemical safety specialists will join a red-team panel that probes how well frontier AI models handle misuse risk. The task centers on single-turn prompts written at three labels: benign, dual-use, and adversarial, and on judging each prompt against a policy standard. You also draft reference answers that show the correct response and explain the technical reasoning. Relevant method areas include GC-MS, LC-MS/MS, IR, and NMR, plus trace-level validation and unknown identification.

Mercor
VerifiedSynthetic / Process Chemist - Remote Task-Based
A panel of process chemistry and synthetic chemistry experts is forming to red-team frontier AI models in a task-based remote engagement. The group checks whether a model can separate legitimate technical questions from dangerous ones, with prompts and answers sorted as benign, dual-use, or adversarial. You will write single-turn prompts, judge model responses against a policy standard, and draft reference answers that include the technical reasoning. The domain sits on the dual-use line: the same scale-up knowledge that turns a medicine into a product can also support a controlled compound, and an impurity control profile can reveal how a substance was made. Pay is a one-time $75 fee.

Mercor
VerifiedForensic Toxicological Chemist - AI Red Team
Mercor is assembling a panel of chemistry and chemical safety experts to red-team frontier AI models. Your core task is to write challenging single-turn prompts in your domain, labeled as benign, dual-use, or adversarial. You then evaluate the model's responses against a defined policy standard and judge whether each was handled correctly. Another key duty is writing the reference answer, including the technical reasoning for why a response is right or wrong. The dual-use nature of this field means you must tell a routine professional question from one that is fishing for harmful information. A model that refuses a legitimate identification question fails as much as one that answers a dangerous one.

Mercor
VerifiedPublic Safety Hazardous Device Specialist (Remote)
Neon needs public safety practitioners to red-team frontier AI models that handle dangerous technical requests. You will write single-turn prompts in the hazardous device domain, label each one as benign, dual-use, or adversarial, then judge how the model responds against a policy standard. Your day-to-day work centers on render-safe procedures, post-blast investigation, and CBRNE response, where routine professional questions sit close to misuse scenarios. Every judgment you make comes with a written rationale, including a reference answer that shows what a correct response looks like and why. The engagement runs remote and task-based, with a one-time payment of $75.

Mercor
VerifiedNuclear Security Engineer (Remote, Task-Based)
A panel of nuclear materials and safeguards specialists will red-team frontier AI models for Mercor, testing how well a model judges misuse potential. You write single-turn prompts at three levels: benign, dual-use, and adversarial, then judge model replies against a policy standard and draft the reference answer with technical reasoning. The work sits on the dual-use line where a material-balance calculation can close an inspector's account or hide a gap. Expect sustained reading and writing about misuse scenarios, with briefings and freedom to pause without penalty.


