
MLE Bench – ML Engineers
Overview
This contract role centers on benchmark-driven evaluation of real-world machine learning systems. You will work directly with production-grade codebases to build, run, and modify model training, evaluation, and inference pipelines. The work blends research and engineering, using frameworks such as PyTorch, TensorFlow, or JAX and collaborating with researchers to design challenging evaluation tasks. The position is fully remote and requires a minimum 20-hour week with a 4-hour daily overlap with PST.
What You'll Do8
- 1Support MLE Bench-style evaluation tasks by working directly with real-world ML codebases.
- 2Develop, execute, and adjust model training, evaluation, and inference pipelines.
- 3Assemble datasets, engineer features, and define metrics for benchmarking and validation.
- 4Identify and fix bugs, refactor code, and enhance production-like ML systems for correctness and speed.
- 5Analyze model behavior, failure modes, and edge cases in benchmark scenarios.
- 6Produce clean, reproducible, and well-documented Python code for ML workflows.
- 7Take part in code reviews to ensure high engineering standards.
- 8Work alongside researchers and engineers to craft challenging ML engineering tasks for evaluating AI systems.
Requirements9
- 1At least 3 years of experience as a Machine Learning Engineer or ML-focused Software Engineer.
- 2Advanced Python skills for machine learning and data workflows.
- 3Hands-on experience building and tuning model training, evaluation, and inference pipelines.
- 4Solid grasp of ML fundamentals, including supervised and unsupervised learning, evaluation metrics, and optimization.
- 5Experience with ML frameworks like PyTorch, TensorFlow, or JAX.
- 6Ability to read, navigate, and modify complex real-world ML codebases.
- 7Track record of writing readable, reusable, and maintainable production-quality code.
- 8Strong problem-solving and debugging abilities.
- 9Excellent spoken and written English communication skills.
Who Should Apply
The candidate who thrives here brings at least 3 years of ML engineering experience and can move fluidly between research ideas and production code. This role is less suitable for engineers who prefer model experimentation only and avoid dealing with messy, real-world codebases. A common reason candidates get rejected is struggling with the 60-minute live coding challenge, especially when asked to modify an existing ML pipeline under time pressure. Others score low on fit when they cannot articulate evaluation metric choices or fail to demonstrate strong debugging skills.
Location
Required Skills
Application Tip
Practice refactoring an existing ML training pipeline in Python within a 60-minute window. Be ready to explain your evaluation metric choices and debug code out loud, as the live coding interview will test both speed and ML fundamentals.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedMLE Bench – Data Analyst
Turing runs benchmark-driven evaluation projects for frontier AI labs, and this role focuses on data analysis for MLE Bench. You will inspect real-world machine learning outputs, define and validate metrics, and write reproducible Python and SQL analysis scripts. The 3-month contractor engagement requires at least 20 hours per week and a 4-hour daily overlap with PST.

Turing
VerifiedPython Machine Learning Engineer
Turing pairs frontier AI labs with data, training pipelines, and specialized researchers, and helps enterprises move AI from prototype to production systems that deliver measurable business results. This remote contract role focuses on machine learning solution delivery using Python, with ownership across data pipelines, model design, deployment, and monitoring. The engineer sets technical direction, mentors peers, and keeps ML initiatives aligned with business priorities. Hands-on experience in Kaggle competitions or ML benchmarks is a strong signal for this position.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Micro1
VerifiedMachine Learning Engineer
A remote contract role for building and refining machine learning models using Python. You’ll work on real-world data to help train next‑generation AI systems, shaping how models learn, reason, and perform. The project emphasizes data handling, model evaluation, and clear documentation of experiments and results. Python, MongoDB, and core ML workflows are central to this work.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Mercor
VerifiedMachine Learning Engineer Talent Network
This is an open, rolling application for a network of machine learning engineers seeking remote contract work with top-tier AI research labs. As an expert in this network, you'll contribute to training, evaluating, and improving AI models by completing real-world tasks and delivering domain-specific feedback. Projects typically require 15-30 hours per week and offer flexible schedules with pay ranging from $70 to $250 per hour.

