
Senior Software Engineer – LLM Evaluation & Repository Validation
Overview
This contract role centers on building LLM evaluation datasets from public repository histories. You will triage GitHub issues, set up Docker environments, and run codebases on your own machine to score how well language models fix real bugs. The work targets popular repositories with 500+ stars and spans languages such as Python, JavaScript, and Go. The project uses a human-in-the-loop approach to expand task coverage across difficulty levels and programming languages.
What You'll Do6
- 1Triage GitHub issues from trending open-source libraries and select those worth turning into LLM evaluation tasks.
- 2Set up repository environments by writing Dockerfiles, installing dependencies, and configuring build steps.
- 3Assess unit test suites for coverage and quality, then flag gaps that weaken evaluation reliability.
- 4Check out, modify, and run real codebases to test how LLMs perform on bug-fixing tasks.
- 5Collaborate with researchers to identify repositories and issues that challenge LLMs.
- 6Lead a small team of junior engineers working on related project milestones.
Requirements6
- 1Strong command of at least one language from this set: Python, JavaScript, Java, Go, Rust, C/C++, C#, or Ruby.
- 2Solid working knowledge of Git, Docker, and standard build pipeline configuration.
- 3Ability to read and move through complex, unfamiliar codebases without hand-holding.
- 4Comfort with checking out, changing, and executing real-world projects on a local machine.
- 5Prior open-source contribution or evaluation experience is a plus.
- 6Exposure to LLM research, LLM evaluation, or developer tooling and automation agents is preferred.
Who Should Apply
This role suits a senior or tech lead engineer who enjoys digging into popular open-source repositories and turning messy GitHub issues into structured evaluation tasks. The role works less well for developers who prefer building greenfield products over maintaining and debugging existing code. Candidates often get screened out when they lack Docker or Git fluency, or when they cannot show experience with high-star repositories (500+ stars). A history of contributing to open-source projects or evaluating LLM outputs gives you a clear edge.
Location
Required Skills
Application Tip
Show specific examples of setting up Docker environments for a popular open-source repo (500+ stars) and triaging or resolving issues in that codebase. Mention the exact programming languages you used and any LLM evaluation or automation projects you have worked on.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – Python (LLM Evaluation & Repository Validation)
GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.

Turing
VerifiedSenior Software Engineer – C#(LLM Evaluation & Repository Validation)
This project builds LLM evaluation and training datasets that help models solve realistic software engineering problems. The work involves creating verifiable software engineering tasks from public repository histories using a synthetic, human-in-the-loop approach. Engineers analyze GitHub issues, set up Docker-based environments, and evaluate test coverage to judge model performance in bug-fixing scenarios. The role fits a tech lead-level engineer comfortable with C#, Git, and running complex codebases locally.

Turing
VerifiedSenior Software Engineer – Go (LLM Evaluation & Repository Validation)
This role centers on building LLM evaluation datasets that teach models to solve realistic software engineering problems. You will triage GitHub issues from popular open-source repositories, set up Docker containers and code environments, and run modified codebases in local environments to measure how well models handle bug-fixing tasks. The work is hands-on and includes evaluating unit test coverage, collaborating with researchers to identify challenging tasks, and the option to lead junior engineers. Projects draw on public repository histories with a human-in-the-loop, synthetic approach to expand task coverage across Go and other programming languages.

Turing
VerifiedSenior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
The project builds LLM evaluation datasets and verifiable software engineering tasks drawn from public repository histories. Ruby engineers at tech-lead level will triage GitHub issues, set up repositories with Docker, and run codebases locally to score model performance. The day-to-day includes test coverage review, environment automation, and collaboration with researchers to pick challenging bug-fixing problems. The role also offers the chance to lead a small team of junior engineers while working on advanced AI projects.

Turing
VerifiedSenior Software Engineer – Rust (LLM Evaluation & Repository Validation)
This role sits inside a project that produces LLM evaluation and training datasets for real-world software engineering problems. The team builds verifiable SWE tasks from public repository histories using a synthetic approach with human-in-the-loop review, then expands coverage across programming languages and difficulty levels. You will analyze and triage GitHub issues in trending open-source libraries, configure repos with Docker, run codebases locally, and assess how well LLMs handle bug-fixing scenarios. The work is hands-on and spans environment automation, test coverage evaluation, and collaboration with researchers to select challenging repositories and issues. You also have the option to lead a small team of junior engineers.

Turing
VerifiedSenior Software Engineer – C++ (LLM Evaluation & Repository Validation)
Turing is assembling a team to build LLM evaluation and training datasets that teach models to solve real software engineering tasks. This project constructs verifiable SWE problems by mining public repository histories and using a synthetic, human-in-the-loop approach. A senior C++ engineer in this role will analyze trending GitHub issues, triage bugs, and evaluate unit test quality across open-source libraries. The work includes setting up repositories with Docker, running and modifying codebases locally, and partnering with researchers to design challenges that stretch LLM capabilities. This is a fully remote contractor assignment with required PST overlap.

