
Senior Software Engineer – Python (LLM Evaluation & Repository Validation)
Overview
GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.
What You'll Do6
- 1Triage and analyze incoming issues across high-traffic open-source GitHub projects, flagging those suitable for evaluation tasks.
- 2Configure repository environments from scratch, including Docker containerization and dependency setup.
- 3Inspect unit test suites to judge their coverage and ability to catch regressions.
- 4Run and modify real-world codebases locally to test how LLMs perform on bug-fixing tasks.
- 5Partner with researchers to select repositories and issues that present a real challenge for current LLMs.
- 6Optionally lead a small team of junior engineers working on dataset expansion.
Requirements8
- 1At least 3 years of professional software engineering experience.
- 2Strong command of Python for development and scripting tasks.
- 3Working proficiency with Git, Docker, and standard CI/pipeline tooling.
- 4Comfort navigating large, unfamiliar codebases and tracing logic across modules.
- 5Ability to set up, run, and modify real-world projects on a local machine.
- 6Prior open-source contribution or code review experience is preferred.
- 7Bonus: previous work on LLM research or model evaluation projects.
- 8Bonus: experience building developer tools or automation agents.
Who Should Apply
The ideal candidate is a Python engineer with 3+ years of experience who is comfortable digging into unfamiliar GitHub repositories and assessing test quality. You should be happy spending your day triaging issues, setting up Docker containers, and running others' code locally, because this role is about evaluation, not greenfield development. The role is less suitable for developers who only want to write new features or who shy away from infrastructure setup. Common reasons candidates score low include weak Git and Docker fundamentals, or an inability to demonstrate a practical workflow for running and modifying open-source code. In interviews, be ready to explain how you judge whether a test suite truly covers the intended behavior.
Location
Required Skills
Application Tip
Prepare a concrete example of a GitHub issue you triaged or fixed in an open-source project. Walk through how you set up the environment, which tests you ran, and how you verified the fix, since the interview will likely probe your hands-on workflow.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation & Repository Validation
This contract role centers on building LLM evaluation datasets from public repository histories. You will triage GitHub issues, set up Docker environments, and run codebases on your own machine to score how well language models fix real bugs. The work targets popular repositories with 500+ stars and spans languages such as Python, JavaScript, and Go. The project uses a human-in-the-loop approach to expand task coverage across difficulty levels and programming languages.

Turing
VerifiedSenior Software Engineer – C#(LLM Evaluation & Repository Validation)
This project builds LLM evaluation and training datasets that help models solve realistic software engineering problems. The work involves creating verifiable software engineering tasks from public repository histories using a synthetic, human-in-the-loop approach. Engineers analyze GitHub issues, set up Docker-based environments, and evaluate test coverage to judge model performance in bug-fixing scenarios. The role fits a tech lead-level engineer comfortable with C#, Git, and running complex codebases locally.

Turing
VerifiedSenior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
The project builds LLM evaluation datasets and verifiable software engineering tasks drawn from public repository histories. Ruby engineers at tech-lead level will triage GitHub issues, set up repositories with Docker, and run codebases locally to score model performance. The day-to-day includes test coverage review, environment automation, and collaboration with researchers to pick challenging bug-fixing problems. The role also offers the chance to lead a small team of junior engineers while working on advanced AI projects.

Turing
VerifiedSenior Software Engineer – Go (LLM Evaluation & Repository Validation)
This role centers on building LLM evaluation datasets that teach models to solve realistic software engineering problems. You will triage GitHub issues from popular open-source repositories, set up Docker containers and code environments, and run modified codebases in local environments to measure how well models handle bug-fixing tasks. The work is hands-on and includes evaluating unit test coverage, collaborating with researchers to identify challenging tasks, and the option to lead junior engineers. Projects draw on public repository histories with a human-in-the-loop, synthetic approach to expand task coverage across Go and other programming languages.

Turing
VerifiedSenior Software Engineer – C++ (LLM Evaluation & Repository Validation)
Turing is assembling a team to build LLM evaluation and training datasets that teach models to solve real software engineering tasks. This project constructs verifiable SWE problems by mining public repository histories and using a synthetic, human-in-the-loop approach. A senior C++ engineer in this role will analyze trending GitHub issues, triage bugs, and evaluate unit test quality across open-source libraries. The work includes setting up repositories with Docker, running and modifying codebases locally, and partnering with researchers to design challenges that stretch LLM capabilities. This is a fully remote contractor assignment with required PST overlap.

Turing
VerifiedSenior Software Engineer – Rust (LLM Evaluation & Repository Validation)
This role sits inside a project that produces LLM evaluation and training datasets for real-world software engineering problems. The team builds verifiable SWE tasks from public repository histories using a synthetic approach with human-in-the-loop review, then expands coverage across programming languages and difficulty levels. You will analyze and triage GitHub issues in trending open-source libraries, configure repos with Docker, run codebases locally, and assess how well LLMs handle bug-fixing scenarios. The work is hands-on and spans environment automation, test coverage evaluation, and collaboration with researchers to select challenging repositories and issues. You also have the option to lead a small team of junior engineers.

