
MLE Bench – Data Analyst
Overview
Turing runs benchmark-driven evaluation projects for frontier AI labs, and this role focuses on data analysis for MLE Bench. You will inspect real-world machine learning outputs, define and validate metrics, and write reproducible Python and SQL analysis scripts. The 3-month contractor engagement requires at least 20 hours per week and a 4-hour daily overlap with PST.
What You'll Do7
- 1Analyze structured and unstructured datasets generated by ML training, inference, and evaluation pipelines to uncover patterns and anomalies.
- 2Define and compute metrics that measure model performance, then validate those metrics for accuracy and relevance.
- 3Investigate data distributions, model outputs, failure modes, and edge cases tied to benchmark tasks.
- 4Write and run Python and SQL code to produce reports and support the evaluation workflow.
- 5Check datasets and experimental results for quality, consistency, and correctness.
- 6Document your analytical methods and create reproducible workflows that others can follow.
- 7Collaborate with ML engineers and researchers to design realistic evaluation scenarios for MLE Bench.
Requirements8
- 1At least 3 years of experience as a data analyst or analytics-focused engineer.
- 2Strong Python skills for data analysis, including writing clean and well-documented code.
- 3Solid experience with SQL and working with relational datasets.
- 4Hands-on experience analyzing ML outputs and evaluation metrics.
- 5A firm grasp of statistics and analytical reasoning.
- 6Ability to work with large, complex datasets and draw reliable conclusions.
- 7Excellent written and spoken English for clear communication.
- 8Comfort with documenting analysis and creating reproducible processes.
Who Should Apply
The ideal candidate has several years of hands-on data analysis experience, deep familiarity with Python and SQL, and a genuine interest in how machine learning models are evaluated. This role is less suitable for analysts who prefer static dashboards over probing model behavior or who lack comfort with statistical reasoning. Candidates often get rejected when they struggle with the live coding interview, especially on SQL queries or Python data manipulation. Another common miss is failing to show a clear work schedule that accommodates the required 4-hour overlap with PST.
Location
Required Skills
Application Tip
Before you apply, run through a timed Python and SQL coding exercise that mimics a data analysis task, such as computing a metric from a sample dataset. Be ready to explain how you validate metrics and handle messy data, and mention any past work with ML evaluation in your application.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedMLE Bench – ML Engineers
This contract role centers on benchmark-driven evaluation of real-world machine learning systems. You will work directly with production-grade codebases to build, run, and modify model training, evaluation, and inference pipelines. The work blends research and engineering, using frameworks such as PyTorch, TensorFlow, or JAX and collaborating with researchers to design challenging evaluation tasks. The position is fully remote and requires a minimum 20-hour week with a 4-hour daily overlap with PST.

Turing
VerifiedPython Machine Learning Engineer
Turing pairs frontier AI labs with data, training pipelines, and specialized researchers, and helps enterprises move AI from prototype to production systems that deliver measurable business results. This remote contract role focuses on machine learning solution delivery using Python, with ownership across data pipelines, model design, deployment, and monitoring. The engineer sets technical direction, mentors peers, and keeps ML initiatives aligned with business priorities. Hands-on experience in Kaggle competitions or ML benchmarks is a strong signal for this position.

Turing
VerifiedSWE Bench – Data Engineer/Data Scientist
Turing, a San Francisco-based research accelerator, is hiring experienced data engineers and data scientists for benchmark-driven evaluation of advanced AI systems. The role centers on SWE Bench-style tasks: building and validating data pipelines, processing structured and unstructured datasets, and preparing features for data science workflows. You will write Python code, run local experiments, and verify outputs for correctness and reproducibility. This is a fully remote contractor assignment with required overlap with PST hours.

Turing
VerifiedData Scientist/Analyst
This contract role focuses on improving AI model performance through hands-on Python development and rigorous data analysis. You will build and maintain code for model training, run evaluations, and rank model responses across diverse domains. The work includes creating high-quality datasets for supervised fine-tuning and collaborating with researchers on RLHF efforts. The position is fully remote and requires a minimum of 20 hours per week with 4 hours of overlap with Pacific Time.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

SME Careers
VerifiedPython ML Quality Lead for Remote Contract Role
Remote, hourly contractor role overseeing quality for Python-driven machine learning training projects. You will review AI-generated Python code, ML pipelines, and model explanations, delivering precise feedback aligned to project rubrics. Assessments focus on code quality, data handling, reproducibility, and evaluation standards to guide contributors toward consistent results. SME Careers, a growth-focused AI data services company under SuperAnnotate, connects you with future expert opportunities within the expert network.

