Mercor
MercorVerified listing
Remote

SWE-Bench Task Auditor | $70-$90/hr Remote

70–90/hr
Remote · Remote — United States
Posted September 1, 2026
hourly
3 openings

Overview

This role puts you inside the evaluation pipeline for a frontier AI lab's models. You will audit SWE-Bench style repository tasks, checking reference patches, test harnesses, and Docker isolation for correctness and reproducibility. Your written, rubric-based feedback shapes which tasks get used for training and evaluation. The work is remote and hourly, paying $70-$90 per hour, and it demands strong open-source credentials plus fluency in Python and at least one of Java, Go, TypeScript, or C++.

What You'll Do5

  • 1Audit repository-level benchmark tasks for correctness, reproducibility, and alignment with the intended grading criteria.
  • 2Inspect reference patches to confirm they solve the stated problem and do not contain hidden shortcuts or answer leakage.
  • 3Review test harnesses and Docker isolation setups so each task executes in a controlled environment and produces accurate pass/fail results.
  • 4Detect reward hacking attempts, task contamination, or any behavior that inflates model scores without genuine problem-solving.
  • 5Write structured, rubric-based feedback for each task, documenting issues and suggesting fixes.

Requirements8

  • 13+ years of professional software engineering experience, with recent hands-on coding in production or OSS settings.
  • 2Demonstrated open-source track record: merged PRs, committer status, or maintainer role on a real project.
  • 3Ability to audit reference patches and test runners for correctness, and to evaluate Docker isolation in CI environments.
  • 4Skill at spotting answer leakage, reward hacking, or other forms of task contamination that inflate benchmark scores.
  • 5Professional fluency in Python, plus working knowledge of at least one of Java, Go, TypeScript, or C++.
  • 6Familiarity with SWE-Bench, SWE-Bench Verified, or comparable repository-level coding benchmarks (preferred).
  • 7Maintainer history on a major Python OSS project such as Django, Flask, scikit-learn, sympy, or pytest (preferred).
  • 8Prior experience reviewing code or grading programming tasks (preferred).

Who Should Apply

The ideal candidate is a senior engineer with a public trail of merged PRs and maintainer experience who enjoys auditing code for subtle flaws. You should be comfortable reading patches from other developers and spotting the moment a test answer leaks into the training data. This role is less suitable for engineers who have never done formal code review or who have no open-source footprint. Applications often fail when the candidate lists no real merged contributions, or when they cannot articulate how reward hacking works in a repository-level benchmark. If you have triaged issues or reviewed PRs on a major Python project, that will separate you from most applicants.

Salary Insight

The rate is $70.00-$90.00 per hour. That rate sits at the senior end for contract software engineering work, reflecting the specialized nature of benchmark auditing and the need for maintainer-level judgment.

Location

Typeremote
LocationRemote — United States
Eligible countriesUnited States
This is a remote position

Required Skills

pythonjavagotypescriptc++dockerswe-benchdjangoflaskscikit-learnsympypytestgit

Application Tip

In your application, include direct links to two or three merged PRs on a public repository and state your exact role in that project. If you have worked with SWE-Bench, show one instance where you verified a task or caught a leakage issue.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

12d agoRemotecontract

SWE Bench – Data Engineer/Data Scientist

Turing, a San Francisco-based research accelerator, is hiring experienced data engineers and data scientists for benchmark-driven evaluation of advanced AI systems. The role centers on SWE Bench-style tasks: building and validating data pipelines, processing structured and unstructured datasets, and preparing features for data science workflows. You will write Python code, run local experiments, and verify outputs for correctness and reproducibility. This is a fully remote contractor assignment with required overlap with PST hours.

Competitive salary
PythonData EngineeringData Science+7 more
Micro1

Micro1

21d agoRemotecontract
Hot

Technical Writer for AI Benchmark Tasks

A remote contractor role focused on shaping AI benchmark work through precise, real-world technical input. You’ll craft expert evaluation tasks using realistic data formats like CSVs, PDFs, and spreadsheets, and assemble source materials such as product specs and code samples. You’ll define clear problem criteria and build detailed rubrics to judge AI outputs for accuracy and audience relevance. Prior experience in regulated or technical fields is valued, but direct AI domain experience isn’t required. Bolded technologies: CSV, PDF, spreadsheets, API references, regulatory documentation.

30–60/hr
· 10 openings
Precision WritingSource SynthesisAudience Calibration+1 more
Mercor

Mercor

3d agoRemotehourly

AWS Serverless & Infrastructure-as-Code Task Auditor

This contract role puts you inside the evaluation loop for training data used by a frontier AI lab. You will audit AWS serverless and infrastructure-as-code tasks for architectural correctness, cross-service integration, and IaC fidelity. Expect to work through Lambda, Step Functions, DynamoDB, EventBridge, and related services, then deliver written feedback guided by a set rubric. The position is fully remote at an hourly rate.

70–90/hr
· 3 openings
AWSLambdaAPI Gateway+14 more
Micro1

Micro1

21d agoRemotecontract
Hot

Social Science Research Assistant for AI Benchmarking

Remote contractor role supporting a frontier AI benchmarking project. You’ll design and run authentic evaluation tasks drawn from social science methods, then produce realistic research artifacts and grading rubrics. Your domain knowledge guides how models should be trained to reason and perform, with emphasis on rigor and transparent documentation. Prior AI experience isn’t required; focus is on solid social science expertise and careful task construction. Key tools and areas include survey design, qualitative coding, literature reviews, and statistical outputs, all built to reflect real-world research complexity. Survey design, Qualitative coding, Statistical analysis are central to the work.

30–50/hr
· 10 openings
Methodological RigorSource SynthesisRubric Fidelity+1 more
Turing

Turing

23d agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Competitive salary
PythonJavaC+++18 more
Mercor

Mercor

3d agoRemotehourly

Kubernetes Task Auditor

At Mercor, you will assess Kubernetes tasks that train a frontier AI lab's models. The work involves grading cluster-operation scenarios, checking manifest correctness, and judging whether failure-mode troubleshooting is realistic and complete. Each review ends with rubric-based written feedback that the AI team uses to improve model performance. This remote hourly role is open to candidates across the United States.

70–90/hr
· 3 openings
KubernetesEksGke+20 more