Turing
TuringVerified listing
Remote

AI Evaluation Specialist – LLM Web Agent Benchmarking (Remote - US)

Remote
Posted September 1, 2026
contract

Overview

Your role centers on building hard research challenges for an AI browsing benchmark. You start from a fact that can be verified, then design a natural-language question that would defeat a frontier model even when it has full web access and many attempts. The output includes checkable clues across dates, people, places, organizations, works, events, records, and quantities, plus a validation record of the obvious searches you ran and what they returned. The work is investigative research, not subject-matter expertise or content writing, and the evidence trail carries most of the weight. The contract runs 8 weeks at 40 hours per week, with at least 4 hours of overlap with PST.

What You'll Do6

  • 1Construct natural-language research questions that start from a verifiable fact and end with a short, stable answer anyone can check.
  • 2Write individual clues that work on their own and touch several fact types, including dates, people, places, organizations, works, events, records, and quantities, with clear constraints.
  • 3Produce a validation record that shows the obvious searches you ran and the results each one returned.
  • 4Trace primary records through government and institutional databases, archives, registries, and PDF documents.
  • 5Name exact pages, tables, and sections in your sources instead of citing a homepage.
  • 6Build enough context in unfamiliar subjects to ask precise, well-scoped research questions.

Requirements9

  • 1Hold a Master's degree or bring more than 3 years of related work experience.
  • 2Show proven online research ability, including locating primary records and moving through government databases, institutional archives, registries, and PDF documents.
  • 3Keep sourcing precise: cite exact pages, tables, and sections, not just a homepage.
  • 4Work from scratch on unfamiliar subjects without hand-holding.
  • 5Write in native or near-native English.
  • 6Handle structured documentation without friction; the evidence trail is the bulk of the work.
  • 7Bring experience with LLM evaluation, red-teaming, or benchmark construction.
  • 8Demonstrate depth in at least one of these areas: reference librarianship or archival research, investigative journalism or fact-checking, OSINT or due diligence, patent or prior-art search, genealogy, or competitive quizzing and puzzle hunts.
  • 9Know JSON and structured data delivery formats, which helps even though not required.

Who Should Apply

Investigative researchers who enjoy making facts hard to find will fit this role. The work suits people with a background in reference librarianship, fact-checking, OSINT, or puzzle construction, since it demands exact citations and primary records. A pure content writer or someone expecting a subject-matter expert role should skip it. Candidates often score low when their sources end at a homepage, when their validation trail misses the obvious searches, or when they avoid the structured documentation that carries most of the assignment.

Location

Typeremote
LocationRemote
Eligible countriesUnited States
This is a remote position

Required Skills

web researchprimary source researchgovernment databasesinstitutional archivesregistriespdf documentssource citationllm evaluationred teamingbenchmark constructionfact-checkinginvestigative journalismosintdue diligencekycpatent searchprior art searchlegal discoverygenealogyrecords researchcompetitive quizzingpuzzle huntsjsonstructured dataarchival researchspecial collectionsllmuser research

Application Tip

Build a one-page sample before applying: pick a verifiable fact, write a hard-to-locate question around it, then list the obvious searches you ran and cite the exact page, table, or section where each clue lives. That mirrors the core deliverable and gives reviewers direct evidence of your sourcing precision.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

7d agoRemotecontract
Hot

Web Research Specialist

You will design research problems for an evaluation benchmark that tests frontier AI browsing agents. Each problem starts from a verifiable fact and works backward to create a question that a state-of-the-art model cannot solve with full web access and repeated attempts. The work demands investigative research rather than subject expertise or content writing. You will finish each assignment with a natural-language question, independently checkable clues across fact types such as dates, people, places, and records, plus a validation record that documents the obvious searches you ran.

Competitive salary
LLMLLM EvaluationRed Teaming+18 more
Turing

Turing

23d agoRemotecontract

Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Competitive salary
PythonJavaScriptReact+13 more
Micro1

Micro1

1mo agoRemotecontract

AI Evaluation Specialist

As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.

20–35/hr
· 50 openings
Rubric Based EvaluationStructured Observation And ReportingHigh Attention To Detail+2 more
Turing

Turing

23d agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Competitive salary
PythonJavaC+++18 more
Micro1

Micro1

11d agoRemotecontract
Hot

Senior Litigation Attorney for AI Evaluation Projects

Senior Litigation Attorney with deep US practice is sought to contribute to a frontier AI evaluation project focused on complex commercial disputes. Remote work from the United States is allowed, with contractor engagement and hourly compensation between 100 and 150. You will seed realistic multi-turn litigation scenarios, guide junior lawyers, and help shape how AI models learn from real-world inputs using public or synthetic materials. Your expertise will inform evaluation rubrics and identify edge cases to improve model reasoning and performance. Litigation strategy and adversarial reasoning will be essential as you collaborate with a small team and align with project leads.

100–150/hr
· 10 openings
Litigation StrategyAdversarial ReasoningCase Team Direction+1 more
Turing

Turing

21d agoRemotecontract

Senior Software Engineer – LLM Evaluation

In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Competitive salary
PythonJavaScriptReact+14 more