
Web Research Specialist
Overview
You will design research problems for an evaluation benchmark that tests frontier AI browsing agents. Each problem starts from a verifiable fact and works backward to create a question that a state-of-the-art model cannot solve with full web access and repeated attempts. The work demands investigative research rather than subject expertise or content writing. You will finish each assignment with a natural-language question, independently checkable clues across fact types such as dates, people, places, and records, plus a validation record that documents the obvious searches you ran.
What You'll Do6
- 1Construct natural-language research questions that start from a verifiable fact and make that fact difficult for AI agents to locate.
- 2Create clusters of independently checkable clues spanning dates, people, places, organizations, works, events, records, and quantities, each with specific constraints.
- 3Run the obvious searches yourself and document the results in a validation record that proves the question is hard.
- 4Cite exact pages, tables, and sections from primary records found in government and institutional databases, archives, and registries.
- 5Maintain a structured evidence trail for every question, since documentation makes up the majority of the work.
- 6Research unfamiliar subjects from scratch using open-web methods when no prior domain knowledge exists.
Requirements8
- 1Master's degree or more than 3 years of equivalent work experience.
- 2Demonstrated open-web research ability, including locating primary records and navigating government databases, archives, registries, and PDF documents.
- 3Precision with sourcing: cite exact pages, tables, and sections, not just homepages.
- 4Comfort researching unfamiliar subjects from scratch.
- 5Native or near-native written English.
- 6High tolerance for structured documentation, with the evidence trail as the core deliverable.
- 7Prior experience with LLM evaluation, red-teaming, or benchmark construction.
- 8Domain experience in at least one of: reference librarianship, investigative journalism, fact-checking, OSINT, due diligence, KYC, patent search, legal discovery, genealogy, or competitive quizzing.
Who Should Apply
The ideal candidate is a trained researcher who enjoys the hunt: someone from a reference librarian, investigative journalist, OSINT analyst, or fact-checker background who regularly digs through primary records and government databases. This role suits people who find satisfaction in documenting every search step, not just delivering an answer. It is less suitable for subject matter experts who want to write content or answer questions in their own field, and for writers who dislike rigid evidence trails. Common reasons candidates score low include citing vague sources such as homepages instead of exact pages or tables, and failing to show retrieval depth when the topic sits outside their specialty.
Location
Required Skills
Application Tip
When you submit the assessment, include a written validation record that shows the exact searches you ran, the databases and archives you queried, and page-level citations. This demonstrates the investigative rigor that this role scores highest on.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Evaluation Specialist – LLM Web Agent Benchmarking (Remote - US)
Your role centers on building hard research challenges for an AI browsing benchmark. You start from a fact that can be verified, then design a natural-language question that would defeat a frontier model even when it has full web access and many attempts. The output includes checkable clues across dates, people, places, organizations, works, events, records, and quantities, plus a validation record of the obvious searches you ran and what they returned. The work is investigative research, not subject-matter expertise or content writing, and the evidence trail carries most of the weight. The contract runs 8 weeks at 40 hours per week, with at least 4 hours of overlap with PST.

Mercor
VerifiedApplied Legal Benchmark Specialist
Legal experts with a JD, LLM, or SJD will craft and review multiple-choice questions for an AI research initiative. The work is remote and asynchronous, with two possible tracks: authoring original questions or verifying pre-written ones. Each question targets a core law domain such as intellectual property, securities, or antitrust. Experts rate difficulty from medium to expert and provide chain-of-thought solutions with references.

Micro1
VerifiedSocial Science Research Assistant for AI Benchmarking
Remote contractor role supporting a frontier AI benchmarking project. You’ll design and run authentic evaluation tasks drawn from social science methods, then produce realistic research artifacts and grading rubrics. Your domain knowledge guides how models should be trained to reason and perform, with emphasis on rigor and transparent documentation. Prior AI experience isn’t required; focus is on solid social science expertise and careful task construction. Key tools and areas include survey design, qualitative coding, literature reviews, and statistical outputs, all built to reflect real-world research complexity. Survey design, Qualitative coding, Statistical analysis are central to the work.

Micro1
VerifiedEvaluation Specialist for AI Training and Research
Remote contractor role for Evaluation Specialists or Recent Grads who will help train next‑generation AI systems. You will craft original QA pairs, perform rigorous source triangulation, and create multi‑step questions that require synthesis. The project emphasizes high-quality, real‑world input to influence how models learn, reason, and perform. Prior AI experience isn’t required; deep domain knowledge matters and it’s welcomed.

SME Careers
VerifiedLegal Researcher Remote Contract for AI Training Content
Operate as a remote AI training content focused Legal Researcher for model guidance. Draft prompts and gold-standard outputs such as issue outlines, research plans, memo templates, and authority summaries, using sources from westlaw or lexisnexis with bluebook style citations. Scrutinize AI outputs for accuracy, flag jurisdictional issues, and verify primary authorities to support defensible conclusions. Work independently across time zones, contributing to structured knowledge management and high quality legal writing.

Micro1
VerifiedMember of Technical Staff, Legal Research
This role sits at the intersection of legal expertise and AI research. You'll design evaluation frameworks that measure how well large language models handle complex legal reasoning, working remotely to help build the next generation of enterprise AI for the legal industry. Your work will directly shape how AI agents perform tasks like statutory interpretation, contract analysis, and workflow automation.

