
Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Overview
Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.
What You'll Do6
- 1Curate code examples and develop precise solutions in Python, JavaScript/ReactJS, C/C++, Java, Rust, and Go for AI training datasets.
- 2Review and correct AI-generated code to improve its efficiency, scalability, and reliability.
- 3Collaborate with researchers and engineers to benchmark AI-driven coding solutions against industry performance standards.
- 4Build software agents that inspect code quality and detect recurring error patterns.
- 5Analyze the software engineering lifecycle from prototyping and API design to production deployment and monitoring, then design experiments to test model capabilities at each stage.
- 6Design automated verification mechanisms that confirm whether a proposed solution correctly solves a given software engineering task.
Requirements4
- 1At least 3 years of professional software engineering experience.
- 2Strong record of building full-stack applications and deploying production-grade software with modern tools.
- 3Deep understanding of software architecture, design, development, debugging, and code quality review.
- 4Excellent written and verbal communication skills for producing clear, structured evaluation rationales.
Who Should Apply
The ideal candidate has at least 3 years of software engineering experience and has built and shipped full-stack applications in production. You should be comfortable reviewing code line by line, writing clear evaluation rationales, and switching between Python, JavaScript/ReactJS, C/C++, Rust, and other languages. This role is less suitable for engineers who need full-time employment or prefer building features over assessing AI output. Rejections often happen when candidates present vague debugging examples or write explanations that lack structure and evidence.
Location
Required Skills
Application Tip
Practice a three-minute response about a time you debugged or reviewed AI-generated code, and structure it as problem, diagnosis, fix, and measurable result. Your evaluation rationale will be scored on clarity, so avoid rambling and quantify the outcome with concrete numbers.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Turing
VerifiedSenior Software Engineer – C++ (LLM Evaluation & Repository Validation)
Turing is assembling a team to build LLM evaluation and training datasets that teach models to solve real software engineering tasks. This project constructs verifiable SWE problems by mining public repository histories and using a synthetic, human-in-the-loop approach. A senior C++ engineer in this role will analyze trending GitHub issues, triage bugs, and evaluate unit test quality across open-source libraries. The work includes setting up repositories with Docker, running and modifying codebases locally, and partnering with researchers to design challenges that stretch LLM capabilities. This is a fully remote contractor assignment with required PST overlap.

Turing
VerifiedSenior Software Engineer – Python (LLM Evaluation & Repository Validation)
GitHub repository histories become training ground for LLM evaluation at this contract role. You will create verifiable software engineering tasks through a synthetic, human-in-the-loop pipeline that broadens dataset coverage across programming languages and difficulty levels. Day-to-day work involves triaging issues in trending open-source libraries, configuring Docker environments, and measuring unit test quality. You will run and modify local codebases to see how well LLMs handle real bug-fixing scenarios, then share findings with researchers. There is also room to take on a team lead role with junior engineers.

Turing
VerifiedSenior Software Engineer – Ruby (LLM Evaluation & Repository Validation)
The project builds LLM evaluation datasets and verifiable software engineering tasks drawn from public repository histories. Ruby engineers at tech-lead level will triage GitHub issues, set up repositories with Docker, and run codebases locally to score model performance. The day-to-day includes test coverage review, environment automation, and collaboration with researchers to pick challenging bug-fixing problems. The role also offers the chance to lead a small team of junior engineers while working on advanced AI projects.

Mercor
VerifiedSenior Software Engineer, Full Stack (Python, Java, Rust, C#, C++)
This role places senior full-stack engineers inside a top-tier AI lab's generative AI team, where you'll build real applications and internal tools on top of frontier large language models before they're publicly released. You'll work directly with the lab's engineering manager, moving quickly and shipping working code from day one. It's a fully remote, 40-hour-per-week W-2 contract through Cincinnatus LLC, with a competitive hourly rate.

