Turing
TuringVerified listing
Remote

SWE Bench – Data Engineer/Data Scientist

Remote
Posted August 23, 2026
contract

Overview

Turing, a San Francisco-based research accelerator, is hiring experienced data engineers and data scientists for benchmark-driven evaluation of advanced AI systems. The role centers on SWE Bench-style tasks: building and validating data pipelines, processing structured and unstructured datasets, and preparing features for data science workflows. You will write Python code, run local experiments, and verify outputs for correctness and reproducibility. This is a fully remote contractor assignment with required overlap with PST hours.

What You'll Do8

  • 1Process structured and unstructured datasets to support SWE Bench-style evaluation tasks.
  • 2Design, build, and validate data pipelines for benchmarking and evaluation workflows.
  • 3Perform data cleaning, analysis, feature preparation, and validation for data science use cases.
  • 4Write and modify Python code to process data and run local experiments.
  • 5Assess data quality, transformations, and outputs for correctness and reproducibility.
  • 6Create clean, documented, reusable data workflows for benchmarking.
  • 7Review peers' code to maintain code quality and maintainability standards.
  • 8Collaborate with researchers and engineers to design real-world data engineering and data science tasks for AI systems.

Requirements8

  • 1At least 3 years of experience as a Data Engineer, Data Scientist, or data-focused Software Engineer.
  • 2Strong command of Python for data engineering and data science workflows.
  • 3Proven experience with data processing, analysis, and model-related workflows.
  • 4Solid grasp of machine learning and data science fundamentals.
  • 5Hands-on experience with both structured and unstructured data.
  • 6Ability to understand, navigate, and modify complex, real-world codebases.
  • 7Track record of writing readable, reusable, maintainable, and well-documented code.
  • 8Strong problem-solving skills on algorithmic or data-intensive problems, plus clear spoken and written English.

Who Should Apply

The ideal candidate has 3+ years of hands-on data engineering or data science experience and can write production-quality Python under time pressure. They should enjoy untangling messy data and making pipelines reproducible for benchmark tasks. Someone who prefers dashboarding or high-level analytics over hands-on pipeline work will be a weaker fit. A common reason candidates fail is stumbling on the 60-minute live coding challenge, especially when asked to process data quickly. Others get screened out when they cannot explain how they validate data quality or reproduce results in past projects.

Location

Typeremote
LocationRemote
Eligible countriesIndia, Pakistan, Nigeria, Kenya, Egypt +5 more
This is a remote position

Required Skills

pythondata engineeringdata sciencedata pipelinesmachine learningdata analysisfeature engineeringstructured dataunstructured databenchmarking

Application Tip

Practice a 60-minute timed coding exercise that involves building a small data processing pipeline in Python. Be ready to talk about how you validated data quality and ensured reproducibility in a recent project, since that is a core theme of the evaluation.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

3d agoRemotehourly

SWE-Bench Task Auditor

This role puts you inside the evaluation pipeline for a frontier AI lab's models. You will audit SWE-Bench style repository tasks, checking reference patches, test harnesses, and Docker isolation for correctness and reproducibility. Your written, rubric-based feedback shapes which tasks get used for training and evaluation. The work is remote and hourly, paying $70-$90 per hour, and it demands strong open-source credentials plus fluency in Python and at least one of Java, Go, TypeScript, or C++.

70–90/hr
· 3 openings
PythonJavaGo+10 more
Turing

Turing

15d agoRemotecontract

MLE Bench – Data Analyst

Turing runs benchmark-driven evaluation projects for frontier AI labs, and this role focuses on data analysis for MLE Bench. You will inspect real-world machine learning outputs, define and validate metrics, and write reproducible Python and SQL analysis scripts. The 3-month contractor engagement requires at least 20 hours per week and a 4-hour daily overlap with PST.

Competitive salary
PythonSQLMachine Learning+8 more
Turing

Turing

23d agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Competitive salary
PythonJavaC+++18 more
Micro1

Micro1

21d agoRemotecontract
Hot

Technical Writer for AI Benchmark Tasks

A remote contractor role focused on shaping AI benchmark work through precise, real-world technical input. You’ll craft expert evaluation tasks using realistic data formats like CSVs, PDFs, and spreadsheets, and assemble source materials such as product specs and code samples. You’ll define clear problem criteria and build detailed rubrics to judge AI outputs for accuracy and audience relevance. Prior experience in regulated or technical fields is valued, but direct AI domain experience isn’t required. Bolded technologies: CSV, PDF, spreadsheets, API references, regulatory documentation.

30–60/hr
· 10 openings
Precision WritingSource SynthesisAudience Calibration+1 more
Micro1

Micro1

30d agoRemotefull-time

Data Engineer

A full-time remote role for a data engineer who can build and maintain scalable ETL pipelines and handle both structured and unstructured data. You will transform raw information into reliable datasets that support research, analytics, and AI model work, with exposure to machine learning tools viewed as valuable. You’ll collaborate with researchers, data scientists, and engineers to prep data for AI initiatives, and you’ll work with familiar tools and environments to ensure solid data quality.

4–5/hr
PythonSQLAi/ml+3 more
Micro1

Micro1

1mo agoRemotefull-time

Data Engineer

This role puts your data engineering skills to work in an exciting way: you'll help build and refine the datasets that teach next-generation AI models. No prior AI experience is needed—your deep knowledge of databases, pipelines, and data quality is what counts. You'll work remotely on a contract basis, earning a competitive hourly rate while making a tangible impact on how machines learn.

30–130/hr
MysqlPythonEtl+1 more