
SWE Bench – Data Engineer/Data Scientist
Overview
Turing, a San Francisco-based research accelerator, is hiring experienced data engineers and data scientists for benchmark-driven evaluation of advanced AI systems. The role centers on SWE Bench-style tasks: building and validating data pipelines, processing structured and unstructured datasets, and preparing features for data science workflows. You will write Python code, run local experiments, and verify outputs for correctness and reproducibility. This is a fully remote contractor assignment with required overlap with PST hours.
What You'll Do8
- 1Process structured and unstructured datasets to support SWE Bench-style evaluation tasks.
- 2Design, build, and validate data pipelines for benchmarking and evaluation workflows.
- 3Perform data cleaning, analysis, feature preparation, and validation for data science use cases.
- 4Write and modify Python code to process data and run local experiments.
- 5Assess data quality, transformations, and outputs for correctness and reproducibility.
- 6Create clean, documented, reusable data workflows for benchmarking.
- 7Review peers' code to maintain code quality and maintainability standards.
- 8Collaborate with researchers and engineers to design real-world data engineering and data science tasks for AI systems.
Requirements8
- 1At least 3 years of experience as a Data Engineer, Data Scientist, or data-focused Software Engineer.
- 2Strong command of Python for data engineering and data science workflows.
- 3Proven experience with data processing, analysis, and model-related workflows.
- 4Solid grasp of machine learning and data science fundamentals.
- 5Hands-on experience with both structured and unstructured data.
- 6Ability to understand, navigate, and modify complex, real-world codebases.
- 7Track record of writing readable, reusable, maintainable, and well-documented code.
- 8Strong problem-solving skills on algorithmic or data-intensive problems, plus clear spoken and written English.
Who Should Apply
The ideal candidate has 3+ years of hands-on data engineering or data science experience and can write production-quality Python under time pressure. They should enjoy untangling messy data and making pipelines reproducible for benchmark tasks. Someone who prefers dashboarding or high-level analytics over hands-on pipeline work will be a weaker fit. A common reason candidates fail is stumbling on the 60-minute live coding challenge, especially when asked to process data quickly. Others get screened out when they cannot explain how they validate data quality or reproduce results in past projects.
Location
Required Skills
Application Tip
Practice a 60-minute timed coding exercise that involves building a small data processing pipeline in Python. Be ready to talk about how you validated data quality and ensured reproducibility in a recent project, since that is a core theme of the evaluation.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedSWE-Bench Task Auditor
This role puts you inside the evaluation pipeline for a frontier AI lab's models. You will audit SWE-Bench style repository tasks, checking reference patches, test harnesses, and Docker isolation for correctness and reproducibility. Your written, rubric-based feedback shapes which tasks get used for training and evaluation. The work is remote and hourly, paying $70-$90 per hour, and it demands strong open-source credentials plus fluency in Python and at least one of Java, Go, TypeScript, or C++.

Turing
VerifiedMLE Bench – Data Analyst
Turing runs benchmark-driven evaluation projects for frontier AI labs, and this role focuses on data analysis for MLE Bench. You will inspect real-world machine learning outputs, define and validate metrics, and write reproducible Python and SQL analysis scripts. The 3-month contractor engagement requires at least 20 hours per week and a 4-hour daily overlap with PST.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Micro1
VerifiedTechnical Writer for AI Benchmark Tasks
A remote contractor role focused on shaping AI benchmark work through precise, real-world technical input. You’ll craft expert evaluation tasks using realistic data formats like CSVs, PDFs, and spreadsheets, and assemble source materials such as product specs and code samples. You’ll define clear problem criteria and build detailed rubrics to judge AI outputs for accuracy and audience relevance. Prior experience in regulated or technical fields is valued, but direct AI domain experience isn’t required. Bolded technologies: CSV, PDF, spreadsheets, API references, regulatory documentation.

Micro1
VerifiedData Engineer
A full-time remote role for a data engineer who can build and maintain scalable ETL pipelines and handle both structured and unstructured data. You will transform raw information into reliable datasets that support research, analytics, and AI model work, with exposure to machine learning tools viewed as valuable. You’ll collaborate with researchers, data scientists, and engineers to prep data for AI initiatives, and you’ll work with familiar tools and environments to ensure solid data quality.

Micro1
VerifiedData Engineer
This role puts your data engineering skills to work in an exciting way: you'll help build and refine the datasets that teach next-generation AI models. No prior AI experience is needed—your deep knowledge of databases, pipelines, and data quality is what counts. You'll work remotely on a contract basis, earning a competitive hourly rate while making a tangible impact on how machines learn.

