
Senior Software Engineer - AI Coding Agent Evaluation
Listing checked September 22, 2026 · pay as published by Turing
Overview
Turing seeks experienced software engineers to assess how well AI coding agents handle real repositories. You will treat the agent like a colleague whose pull requests you review: checking whether it read the task, picked a sound design, and produced working code. The work spans Python, TypeScript/JavaScript, and Go codebases, plus the evaluation rubrics and data pipelines that sharpen model behavior. Researchers rely on your written judgment to turn one-off observations into repeatable quality signals. This is a freelance contract, roughly three months long, with 40 hours per week preferred.
What You'll Do10
- 1Audit AI-generated code and solutions from live software repositories for technical accuracy.
- 2Inspect agent actions, tool calls, and diffs to judge correctness and maintainability.
- 3Spot recurring failure modes, flawed design choices, and gaps in model reasoning.
- 4Rank competing model outputs and explain why one implementation outperforms another.
- 5Draft and refine scoring rubrics that define quality for coding benchmarks.
- 6Generate preference and evaluation data that feeds back into model training.
- 7Maintain pipelines and tooling for collecting, generating, and reviewing evaluation data.
- 8Distill findings into concise memos and status updates for research partners.
- 9Partner with researchers and engineers to scale qualitative review into repeatable processes.
- 10Report actionable results to AI researchers and engineering stakeholders.
Requirements8
- 15+ years of hands-on software engineering across production systems.
- 2Deep fluency in Python, TypeScript/JavaScript, Go, or a comparable production language.
- 3Track record working inside large, mature codebases rather than greenfield projects.
- 4Keen code review instincts and the technical judgment to defend your assessment.
- 5Ability to articulate why an implementation is correct, broken, or ripe for improvement.
- 6Clear written communication that stands on its own without follow-up questions.
- 7Regular use of modern LLMs or AI coding assistants in your workflow.
- 8Nice to have: exposure to LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training.
Who Should Apply
Engineers with at least five years of production coding experience who enjoy dissecting why a solution works, not just whether it compiles. The role rewards people comfortable in big codebases and fluent in Python, TypeScript/JavaScript, or Go, and who can write a crisp justification for their verdict. Candidates weak on written explanations tend to score low here, because the deliverable is often a clear rationale rather than a patch. Applying without real LLM or AI coding tool exposure also hurts fit, even though it is not strictly required. If you prefer building features all day over reviewing and critiquing code, this contract will feel like a poor match.
Salary Insight
Compensation is described as market rate, with candidates asked to provide a specific hourly rate expectation. No figures are stated.
Pay and demand for Machine Learning & AI roles
AggregatedTypical pay
$72/hour
This role
Pay not disclosed
Most Machine Learning & AI roles pay $50–$95 per hour.
Based on 653 similar roles that publish pay · 168 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 730
- Listed in last 30 days
- 378
- Remote
- 97%
Hiring most right now: micro1 (292) · SME Careers (139) · Mercor (91)
Most requested skills · share of roles
- llm evaluation16%
- python15%
- ai training15%
- trainer feedback11%
Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
Undisclosed by Turing
Comparable roles pay
653 roles$50 – $95/ hr
Not this role’s pay. Turing has not published a range.
Confirmed during Turing profile screening.
Required Skills
Application Tip
State a specific hourly rate with your application and name a project where you reviewed code in a large repository, preferably one involving Python or Go. If you have used LLM evaluation, RLHF, or rubric design, add one sentence on the outcome and quantify it.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Micro1
VerifiedCoding Research Technical Staff - Remote
This full-time remote role centers on advancing how frontier coding agents are evaluated and improved. You will work where AI research, software engineering, and model evaluation meet, creating benchmarks, scoring methods, and data systems that shape how next-generation coding models are measured. The position involves close collaboration with researchers, engineers, and applied AI teams. Python or C++ skills will be central, along with familiarity with LLMs and coding agents.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Micro1
VerifiedMember of Technical Staff, Coding Research
This role sits at the intersection of AI research and software engineering, focusing on how next-generation coding agents are evaluated and improved. You’ll design the benchmarks, methodologies, and data systems that directly shape the measurement of model performance across complex software engineering tasks. It’s a high-impact position for someone who thrives on building rigorous evaluation frameworks and translating model behavior into actionable improvements.

Micro1
VerifiedSenior Software Engineer for AI Code Evaluation and Model Review
micro1 is hiring senior software engineers to work with a leading AI lab. Your main tasks involve evaluating AI-generated code for correctness and design, writing challenging coding problems with reference solutions and test cases, and reviewing model reasoning on architecture, debugging, and system design. You will provide detailed technical feedback to help improve model performance on real-world software tasks. This is a remote contractor role with a pay range of $245 to $280 per hour.


