
Senior Software Engineer – Go (LLM Evaluation & Repository Validation)
Overview
This role centers on building LLM evaluation datasets that teach models to solve realistic software engineering problems. You will triage GitHub issues from popular open-source repositories, set up Docker containers and code environments, and run modified codebases in local environments to measure how well models handle bug-fixing tasks. The work is hands-on and includes evaluating unit test coverage, collaborating with researchers to identify challenging tasks, and the option to lead junior engineers. Projects draw on public repository histories with a human-in-the-loop, synthetic approach to expand task coverage across Go and other programming languages.
What You'll Do6
- 1Triage incoming GitHub issues from trending open-source projects to identify suitable candidates for LLM evaluation.
- 2Set up and configure repository environments, including Docker containers and dependency installation.
- 3Assess the coverage and quality of unit tests across selected codebases.
- 4Modify and run codebases in local environments to benchmark how well LLMs resolve bug-fixing tasks.
- 5Partner with researchers to select repositories and issues that present meaningful challenges for models.
- 6Guide and supervise a small team of junior engineers on project delivery.
Requirements7
- 1At least 3 years of professional software engineering experience.
- 2Strong command of Go and the ability to write, debug, and evaluate code in it.
- 3Working proficiency with Git and Docker, plus the ability to set up basic software pipelines.
- 4Skill at reading, navigating, and modifying complex codebases.
- 5Comfort running and testing real-world projects in local environments.
- 6Prior open-source contribution or code evaluation experience is a plus.
- 7Any past work on LLM research, evaluation, or developer tools is a plus.
Who Should Apply
The ideal candidate is a senior engineer with strong Go skills and a track record of working with high-quality open-source repositories. You should be comfortable setting up Docker environments, triaging GitHub issues, and judging test coverage from concrete evidence. If you have contributed to open-source projects or participated in LLM evaluation work before, that fits well. This role is less suitable for engineers who prefer hands-off research or who cannot commit to at least 20 hours per week with 4 hours of PST overlap. Candidates often score low when they lack demonstrated experience running and modifying unfamiliar codebases, or when they struggle to explain how they assess the quality of a test suite.
Location
Required Skills
Application Tip
Highlight concrete examples of Go projects you have set up with Docker and tested in local environments. Name the GitHub repositories you contributed to or evaluated, and quantify your experience with issue triaging and test coverage assessment. If you have any background in LLM evaluation or developer tools, put it near the top of your resume.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation & Repository Validation
This contract role centers on building LLM evaluation datasets from public repository histories. You will triage GitHub issues, set up Docker environments, and run codebases on your own machine to score how well language models fix real bugs. The work targets popular repositories with 500+ stars and spans languages such as Python, JavaScript, and Go. The project uses a human-in-the-loop approach to expand task coverage across difficulty levels and programming languages.

Turing
VerifiedSenior Software Engineer – C#(LLM Evaluation & Repository Validation)
This project builds LLM evaluation and training datasets that help models solve realistic software engineering problems. The work involves creating verifiable software engineering tasks from public repository histories using a synthetic, human-in-the-loop approach. Engineers analyze GitHub issues, set up Docker-based environments, and evaluate test coverage to judge model performance in bug-fixing scenarios. The role fits a tech lead-level engineer comfortable with C#, Git, and running complex codebases locally.

Turing
VerifiedSenior Software Engineer – Python (LLM Evaluation & Repository Validation)
GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.

Turing
VerifiedSenior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
The project builds LLM evaluation datasets and verifiable software engineering tasks drawn from public repository histories. Ruby engineers at tech-lead level will triage GitHub issues, set up repositories with Docker, and run codebases locally to score model performance. The day-to-day includes test coverage review, environment automation, and collaboration with researchers to pick challenging bug-fixing problems. The role also offers the chance to lead a small team of junior engineers while working on advanced AI projects.

Turing
VerifiedSenior Software Engineer – Rust (LLM Evaluation & Repository Validation)
This role sits inside a project that produces LLM evaluation and training datasets for real-world software engineering problems. The team builds verifiable SWE tasks from public repository histories using a synthetic approach with human-in-the-loop review, then expands coverage across programming languages and difficulty levels. You will analyze and triage GitHub issues in trending open-source libraries, configure repos with Docker, run codebases locally, and assess how well LLMs handle bug-fixing scenarios. The work is hands-on and spans environment automation, test coverage evaluation, and collaboration with researchers to select challenging repositories and issues. You also have the option to lead a small team of junior engineers.

Turing
VerifiedSenior Software Engineer – C++ (LLM Evaluation & Repository Validation)
Turing is assembling a team to build LLM evaluation and training datasets that teach models to solve real software engineering tasks. This project constructs verifiable SWE problems by mining public repository histories and using a synthetic, human-in-the-loop approach. A senior C++ engineer in this role will analyze trending GitHub issues, triage bugs, and evaluate unit test quality across open-source libraries. The work includes setting up repositories with Docker, running and modifying codebases locally, and partnering with researchers to design challenges that stretch LLM capabilities. This is a fully remote contractor assignment with required PST overlap.

