Mercor
MercorVerified listing
Remote

ML Challenge Task Auditor | $70-$90/hr Remote

70–90/hr
Remote · Remote — United States
Posted September 1, 2026
hourly
3 openings

Overview

Contract reviewers with at least three years of hands-on applied ML experience will audit challenge tasks used to train and evaluate models for a frontier AI lab. The work focuses on experiment design, model-selection reasoning, and evaluation methodology, with feedback delivered through a defined rubric. Each review checks for data-quality problems such as leakage, metric gaming, and weak train/test/CV hygiene. This role is not an LLM-application-building or MLOps position; it demands reproducible critique of ML claims using frameworks like PyTorch, TensorFlow, scikit-learn, and XGBoost.

What You'll Do6

  • 1Review applied ML challenge tasks for correctness, clarity, and methodological soundness before they are used in model training or evaluation.
  • 2Assess the design of experiments, including choice of baselines, hyperparameter tuning procedures, and model-selection logic.
  • 3Judge evaluation methodology for hidden biases, data leakage, and susceptibility to metric gaming.
  • 4Write clear, rubric-based feedback that explains strengths and weaknesses and suggests concrete fixes.
  • 5Reproduce or sanity-check reported results against provided code and data to verify claims.
  • 6Flag any train/test/CV hygiene issues that could invalidate benchmark outcomes.

Requirements7

  • 13+ years of hands-on applied or experimental ML work, covering experiment design, model selection, hyperparameter tuning, and evaluation methodology.
  • 2Demonstrated data-quality rigor: ability to detect leakage, spot metric gaming, and enforce train/test/CV hygiene.
  • 3Working proficiency in standard ML frameworks: PyTorch, TensorFlow, scikit-learn, and XGBoost.
  • 4Ability to critique ML claims against evidence and reproduce results from code.
  • 5Preferred: Kaggle competition or other benchmark experience.
  • 6Preferred: graduate research or publications in applied ML.
  • 7Preferred: prior experience grading tasks or performing peer review.

Who Should Apply

The ideal candidate has spent years designing and running applied ML experiments, knows how to spot data leakage and metric gaming from a mile away, and can write feedback that a researcher can act on. Reviewers thrive on adversarial evaluation: they enjoy poking holes in model-selection logic and reproducing results to see if claims hold up. This role is not a fit for builders who want to ship LLM applications or manage MLOps pipelines, because the entire job is auditing the rigor of existing ML tasks. Candidates often get rejected when they cannot provide concrete examples of catching leakage or gaming in past work, or when their experience is too narrow, such as only using AutoML without understanding the underlying methodology.

Salary Insight

The rate is $70.00 - $90.00 per hour, and the position is remote in the United States. That rate lines up with a senior applied ML practitioner who can audit experiments and write rigorous feedback, not a junior engineer.

Location

Typeremote
LocationRemote — United States
Eligible countriesUnited States
This is a remote position

Required Skills

pytorchtensorflowscikit-learnxgboostpythonkaggleexperiment designhyperparameter tuningmodel evaluationcross-validationdata leakage detectionmetric gamingreproducibility

Application Tip

In your application, include a short write-up of a time you caught a data leakage problem, a gamed metric, or a flawed evaluation design, and explain how you proved it to the original authors. That kind of evidence separates strong auditors from candidates who just list ML tools.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

SME Careers

SME Careers

1d agoRemotecontract

Python ML Quality Lead for Remote Contract Role

Remote, hourly contractor role overseeing quality for Python-driven machine learning training projects. You will review AI-generated Python code, ML pipelines, and model explanations, delivering precise feedback aligned to project rubrics. Assessments focus on code quality, data handling, reproducibility, and evaluation standards to guide contributors toward consistent results. SME Careers, a growth-focused AI data services company under SuperAnnotate, connects you with future expert opportunities within the expert network.

Up to 120/hr
PythonMachine LearningPyTorch+32 more
Turing

Turing

19d agoRemotecontract

MLE Bench – ML Engineers

This contract role centers on benchmark-driven evaluation of real-world machine learning systems. You will work directly with production-grade codebases to build, run, and modify model training, evaluation, and inference pipelines. The work blends research and engineering, using frameworks such as PyTorch, TensorFlow, or JAX and collaborating with researchers to design challenging evaluation tasks. The position is fully remote and requires a minimum 20-hour week with a 4-hour daily overlap with PST.

Competitive salary
PythonMachine LearningPyTorch+13 more
Turing

Turing

15d agoRemotecontract

MLE Bench – Data Analyst

Turing runs benchmark-driven evaluation projects for frontier AI labs, and this role focuses on data analysis for MLE Bench. You will inspect real-world machine learning outputs, define and validate metrics, and write reproducible Python and SQL analysis scripts. The 3-month contractor engagement requires at least 20 hours per week and a 4-hour daily overlap with PST.

Competitive salary
PythonSQLMachine Learning+8 more
Mercor

Mercor

3d agoRemotehourly

Kubernetes Task Auditor

At Mercor, you will assess Kubernetes tasks that train a frontier AI lab's models. The work involves grading cluster-operation scenarios, checking manifest correctness, and judging whether failure-mode troubleshooting is realistic and complete. Each review ends with rubric-based written feedback that the AI team uses to improve model performance. This remote hourly role is open to candidates across the United States.

70–90/hr
· 3 openings
KubernetesEksGke+20 more
Mercor

Mercor

3d agoRemotehourly

SWE-Bench Task Auditor

This role puts you inside the evaluation pipeline for a frontier AI lab's models. You will audit SWE-Bench style repository tasks, checking reference patches, test harnesses, and Docker isolation for correctness and reproducibility. Your written, rubric-based feedback shapes which tasks get used for training and evaluation. The work is remote and hourly, paying $70-$90 per hour, and it demands strong open-source credentials plus fluency in Python and at least one of Java, Go, TypeScript, or C++.

70–90/hr
· 3 openings
PythonJavaGo+10 more
Turing

Turing

23d agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Competitive salary
PythonJavaC+++18 more