Turing
TuringVerified listing
Remote

Senior Python Developer

Posted August 20, 2026
contract

Overview

This contract role supports a foundational LLM company by producing high-quality data used to fine-tune and benchmark their models. You will write Python solutions to code-based prompts, evaluate responses from two model versions, and create detailed rationale for ranking decisions. Day-to-day work involves SFT and RLHF data generation, designing evaluation strategies, and reviewing code with a small team. The position does not involve building or fine-tuning models directly, but your outputs directly inform their improvement.

What You'll Do10

  • 1Design, develop, and maintain Python code that supports AI model training.
  • 2Run evaluations to benchmark model performance and analyze results for continuous improvement.
  • 3Rank AI model responses across diverse domains and score them against predefined criteria.
  • 4Write clear explanations and rationales for each evaluation decision.
  • 5Lead Supervised Fine-Tuning (SFT) efforts by creating and maintaining high-quality, task-specific datasets.
  • 6Collaborate with researchers and annotators on Reinforcement Learning with Human Feedback (RLHF) and reward model refinement.
  • 7Create and refine model responses that emphasize clarity, relevance, and technical accuracy.
  • 8Perform thorough peer reviews of code and documentation, offering constructive feedback.
  • 9Work with cross-functional teams to improve model performance and contribute to product enhancements.
  • 10Research and integrate new tools and methods to improve AI training processes.

Requirements7

  • 13+ years of hands-on Python development experience.
  • 2Solid understanding of code quality, formatting, and software development best practices.
  • 3Experience with Python testing frameworks, including unit, integration, and property-based testing.
  • 4Working knowledge of multi-threading and asynchronous programming in Python.
  • 5Ability to apply architectural patterns and refactor code without introducing regressions.
  • 6Strong debugging skills, especially around memory and concurrency problems.
  • 7Fluent conversational and written English communication skills.

Who Should Apply

This role suits a Python developer with at least 3 years of experience who enjoys writing clear, testable code and has a strong grasp of testing frameworks, async, and debugging. The work is contract-based and focused on data generation and model evaluation, so candidates who want to build or fine-tune LLMs directly will find it a poor fit. A common rejection reason is insufficient depth in Python's testing ecosystem, especially property-based testing. Another is weak written communication, since you must explain evaluation rationales to the client and teammates.

Location

TypeNot specified
LocationRemote
This is a remote position

Required Skills

pythonllmsftrlhfmodel evaluationunit testingintegration testingproperty-based testingmultithreadingasynchronous programmingcode review

Application Tip

Highlight your experience with property-based testing and async Python in your application, and be ready to explain a specific example of how you debugged a concurrency issue.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

19d agoRemotecontract

Python + Full-Stack (JS) Developer

You will build and maintain code that trains and refines AI models in a remote contractor setup. Your day-to-day includes running model evaluations, creating datasets for Supervised Fine-Tuning (SFT), and helping researchers implement RLHF. You will also review peers' code and write clear rationales for model scoring. This role demands production-grade Python and JavaScript/TypeScript skills, plus mandatory Docker knowledge.

Competitive salary
PythonJavaScriptTypeScript+12 more
Turing

Turing

12d agoRemotecontract

LLM Java Developer

This role centers on building and maintaining Java back-end components that support LLM training and refinement. Daily work includes running model evaluations, scoring AI responses, and writing clear rationales for those judgments. You will also contribute to Supervised Fine-Tuning (SFT) datasets and collaborate on Reinforcement Learning with Human Feedback (RLHF). The position is a fully remote contractor assignment with a required 4-hour overlap with PST and a commitment of 20, 30, or 40 hours per week.

Competitive salary
JavaLLMAI Model Training+11 more
Turing

Turing

24d agoRemotecontract

Data Scientist/Analyst

This contract role focuses on improving AI model performance through hands-on Python development and rigorous data analysis. You will build and maintain code for model training, run evaluations, and rank model responses across diverse domains. The work includes creating high-quality datasets for supervised fine-tuning and collaborating with researchers on RLHF efforts. The position is fully remote and requires a minimum of 20 hours per week with 4 hours of overlap with Pacific Time.

Competitive salary
PythonData AnalysisData Science+9 more
Turing

Turing

21d agoRemotecontract

Senior Software Engineer – LLM Evaluation

In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Competitive salary
PythonJavaScriptReact+14 more
Turing

Turing

23d agoRemotecontract

Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Competitive salary
PythonJavaScriptReact+13 more
Turing

Turing

14d agoRemotecontract

Senior Software Engineer – Python (LLM Evaluation & Repository Validation)

GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.

Competitive salary
PythonGitDocker+6 more