
Senior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
Overview
The project builds LLM evaluation datasets and verifiable software engineering tasks drawn from public repository histories. Ruby engineers at tech-lead level will triage GitHub issues, set up repositories with Docker, and run codebases locally to score model performance. The day-to-day includes test coverage review, environment automation, and collaboration with researchers to pick challenging bug-fixing problems. The role also offers the chance to lead a small team of junior engineers while working on advanced AI projects.
What You'll Do6
- 1Examine and categorize incoming issues in popular open-source repositories to identify tasks suitable for LLM evaluation.
- 2Prepare and configure codebases, including containerizing projects with Docker and establishing local environments.
- 3Assess the depth and reliability of unit tests within each repository.
- 4Run and alter local codebases to measure how well LLMs fix real bugs.
- 5Work with researchers to pinpoint repositories and issues that pose a meaningful challenge for LLMs.
- 6Oversee and coordinate a group of junior engineers on collaborative project tasks.
Requirements8
- 1At least 3 years of overall software engineering experience.
- 2Strong working knowledge of Ruby.
- 3Practical skill with Git, Docker, and standard software pipeline configuration.
- 4Ability to understand and navigate complex codebases.
- 5Comfort running, modifying, and testing real-world projects locally.
- 6Experience contributing to or evaluating open-source projects.
- 7Previous involvement in LLM research or evaluation projects.
- 8Experience building or testing developer tools or automation agents.
Who Should Apply
An ideal candidate is a senior Ruby engineer with 3+ years of hands-on experience and a habit of exploring high-quality open-source repositories. This person is comfortable setting up Dockerized environments, running unfamiliar codebases locally, and assessing test coverage. The role suits engineers who enjoy repo-level detective work and LLM benchmarking over feature building. Developers who prefer product-facing tasks or rarely touch open-source code will find this role a poor fit. Common reasons candidates get rejected include weak Ruby depth beyond syntax, no practical Git/Docker pipeline skills, and an inability to commit to the required 4-hour PST overlap.
Location
Required Skills
Application Tip
Before applying, pick one Ruby open-source repository you know well, Dockerize it locally, and note the test coverage you found. Mention this in your application along with your GitHub profile, because the role centers on triaging issues and assessing test quality in real repos.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation & Repository Validation
This contract role centers on building LLM evaluation datasets from public repository histories. You will triage GitHub issues, set up Docker environments, and run codebases on your own machine to score how well language models fix real bugs. The work targets popular repositories with 500+ stars and spans languages such as Python, JavaScript, and Go. The project uses a human-in-the-loop approach to expand task coverage across difficulty levels and programming languages.

Turing
VerifiedSenior Software Engineer – C#(LLM Evaluation & Repository Validation)
This project builds LLM evaluation and training datasets that help models solve realistic software engineering problems. The work involves creating verifiable software engineering tasks from public repository histories using a synthetic, human-in-the-loop approach. Engineers analyze GitHub issues, set up Docker-based environments, and evaluate test coverage to judge model performance in bug-fixing scenarios. The role fits a tech lead-level engineer comfortable with C#, Git, and running complex codebases locally.

Turing
VerifiedSenior Software Engineer – Python (LLM Evaluation & Repository Validation)
GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.

Turing
VerifiedSenior Software Engineer – Rust (LLM Evaluation & Repository Validation)
This role sits inside a project that produces LLM evaluation and training datasets for real-world software engineering problems. The team builds verifiable SWE tasks from public repository histories using a synthetic approach with human-in-the-loop review, then expands coverage across programming languages and difficulty levels. You will analyze and triage GitHub issues in trending open-source libraries, configure repos with Docker, run codebases locally, and assess how well LLMs handle bug-fixing scenarios. The work is hands-on and spans environment automation, test coverage evaluation, and collaboration with researchers to select challenging repositories and issues. You also have the option to lead a small team of junior engineers.

Turing
VerifiedSenior Software Engineer – Go (LLM Evaluation & Repository Validation)
This role centers on building LLM evaluation datasets that teach models to solve realistic software engineering problems. You will triage GitHub issues from popular open-source repositories, set up Docker containers and code environments, and run modified codebases in local environments to measure how well models handle bug-fixing tasks. The work is hands-on and includes evaluating unit test coverage, collaborating with researchers to identify challenging tasks, and the option to lead junior engineers. Projects draw on public repository histories with a human-in-the-loop, synthetic approach to expand task coverage across Go and other programming languages.

Turing
VerifiedSenior Software Engineer – C++ (LLM Evaluation & Repository Validation)
Turing is assembling a team to build LLM evaluation and training datasets that teach models to solve real software engineering tasks. This project constructs verifiable SWE problems by mining public repository histories and using a synthetic, human-in-the-loop approach. A senior C++ engineer in this role will analyze trending GitHub issues, triage bugs, and evaluate unit test quality across open-source libraries. The work includes setting up repositories with Docker, running and modifying codebases locally, and partnering with researchers to design challenges that stretch LLM capabilities. This is a fully remote contractor assignment with required PST overlap.

