
Senior Software Engineer – Rust (LLM Evaluation & Repository Validation)
Overview
This role sits inside a project that produces LLM evaluation and training datasets for real-world software engineering problems. The team builds verifiable SWE tasks from public repository histories using a synthetic approach with human-in-the-loop review, then expands coverage across programming languages and difficulty levels. You will analyze and triage GitHub issues in trending open-source libraries, configure repos with Docker, run codebases locally, and assess how well LLMs handle bug-fixing scenarios. The work is hands-on and spans environment automation, test coverage evaluation, and collaboration with researchers to select challenging repositories and issues. You also have the option to lead a small team of junior engineers.
What You'll Do6
- 1Triage and categorize GitHub issues from popular open-source libraries to find tasks suitable for LLM evaluation.
- 2Configure code repositories and build local development environments, including Docker containers and dependency setup.
- 3Assess unit test coverage and quality across candidate repositories to identify meaningful evaluation targets.
- 4Run and modify real-world codebases locally to test how well LLMs perform on bug-fixing tasks.
- 5Work alongside researchers to select repositories and issues that present a genuine challenge for current LLMs.
- 6Lead a group of junior engineers on project tasks when the workload calls for it.
Requirements9
- 1At least 3 years of professional software engineering experience.
- 2Strong command of Rust for real-world development and debugging tasks.
- 3Working knowledge of Git, Docker, and standard software pipeline setup.
- 4Comfort reading and navigating large, unfamiliar codebases.
- 5Ability to run, modify, and test existing codebases in a local environment.
- 6Experience contributing to or evaluating open-source projects is a plus.
- 7Prior involvement in LLM research or evaluation projects is nice to have.
- 8Familiarity with building or testing developer tools or automation agents is a plus.
- 9A track record that supports operating at a tech lead level, from scoping tasks to guiding less experienced engineers.
Who Should Apply
The ideal candidate is a senior software engineer with deep Rust experience and a habit of digging into open-source repos, triaging issues, and assessing test coverage. This role fits someone who enjoys hands-on repo setup with Git and Docker and can reason about why an LLM fails on a specific bug-fixing task. The role is less suitable for engineers who prefer feature development over evaluation work or who are uncomfortable running unfamiliar codebases locally. Candidates often score low when they lack evidence of working with large public codebases, or when they cannot clearly explain how they would design a reproducible environment for a given repository.
Location
Required Skills
Application Tip
Create a short portfolio entry around a real GitHub repository: write a Dockerfile, describe how you would triage an open issue, and explain how you would assess test coverage. Mention any LLM evaluation or open-source contribution work directly in your resume since it maps to the exact tasks in this role.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation & Repository Validation
This contract role centers on building LLM evaluation datasets from public repository histories. You will triage GitHub issues, set up Docker environments, and run codebases on your own machine to score how well language models fix real bugs. The work targets popular repositories with 500+ stars and spans languages such as Python, JavaScript, and Go. The project uses a human-in-the-loop approach to expand task coverage across difficulty levels and programming languages.

Turing
VerifiedSenior Software Engineer – C#(LLM Evaluation & Repository Validation)
This project builds LLM evaluation and training datasets that help models solve realistic software engineering problems. The work involves creating verifiable software engineering tasks from public repository histories using a synthetic, human-in-the-loop approach. Engineers analyze GitHub issues, set up Docker-based environments, and evaluate test coverage to judge model performance in bug-fixing scenarios. The role fits a tech lead-level engineer comfortable with C#, Git, and running complex codebases locally.

Turing
VerifiedSenior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
The project builds LLM evaluation datasets and verifiable software engineering tasks drawn from public repository histories. Ruby engineers at tech-lead level will triage GitHub issues, set up repositories with Docker, and run codebases locally to score model performance. The day-to-day includes test coverage review, environment automation, and collaboration with researchers to pick challenging bug-fixing problems. The role also offers the chance to lead a small team of junior engineers while working on advanced AI projects.

Turing
VerifiedSenior Software Engineer – Go (LLM Evaluation & Repository Validation)
This role centers on building LLM evaluation datasets that teach models to solve realistic software engineering problems. You will triage GitHub issues from popular open-source repositories, set up Docker containers and code environments, and run modified codebases in local environments to measure how well models handle bug-fixing tasks. The work is hands-on and includes evaluating unit test coverage, collaborating with researchers to identify challenging tasks, and the option to lead junior engineers. Projects draw on public repository histories with a human-in-the-loop, synthetic approach to expand task coverage across Go and other programming languages.

Turing
VerifiedSenior Software Engineer – C++ (LLM Evaluation & Repository Validation)
Turing is assembling a team to build LLM evaluation and training datasets that teach models to solve real software engineering tasks. This project constructs verifiable SWE problems by mining public repository histories and using a synthetic, human-in-the-loop approach. A senior C++ engineer in this role will analyze trending GitHub issues, triage bugs, and evaluate unit test quality across open-source libraries. The work includes setting up repositories with Docker, running and modifying codebases locally, and partnering with researchers to design challenges that stretch LLM capabilities. This is a fully remote contractor assignment with required PST overlap.

Turing
VerifiedSenior Software Engineer – Python (LLM Evaluation & Repository Validation)
GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.

