Turing
TuringVerified listing
Remote

Senior Software Engineer - AI Coding Agent Evaluation

Remote · North America, LATAM, India
Posted September 22, 2026
contract
Not Sure? Upload your resume to see every role you match

Listing checked September 22, 2026 · pay as published by Turing

Overview

Turing seeks experienced software engineers to assess how well AI coding agents handle real repositories. You will treat the agent like a colleague whose pull requests you review: checking whether it read the task, picked a sound design, and produced working code. The work spans Python, TypeScript/JavaScript, and Go codebases, plus the evaluation rubrics and data pipelines that sharpen model behavior. Researchers rely on your written judgment to turn one-off observations into repeatable quality signals. This is a freelance contract, roughly three months long, with 40 hours per week preferred.

What You'll Do10

  • 1Audit AI-generated code and solutions from live software repositories for technical accuracy.
  • 2Inspect agent actions, tool calls, and diffs to judge correctness and maintainability.
  • 3Spot recurring failure modes, flawed design choices, and gaps in model reasoning.
  • 4Rank competing model outputs and explain why one implementation outperforms another.
  • 5Draft and refine scoring rubrics that define quality for coding benchmarks.
  • 6Generate preference and evaluation data that feeds back into model training.
  • 7Maintain pipelines and tooling for collecting, generating, and reviewing evaluation data.
  • 8Distill findings into concise memos and status updates for research partners.
  • 9Partner with researchers and engineers to scale qualitative review into repeatable processes.
  • 10Report actionable results to AI researchers and engineering stakeholders.

Requirements8

  • 15+ years of hands-on software engineering across production systems.
  • 2Deep fluency in Python, TypeScript/JavaScript, Go, or a comparable production language.
  • 3Track record working inside large, mature codebases rather than greenfield projects.
  • 4Keen code review instincts and the technical judgment to defend your assessment.
  • 5Ability to articulate why an implementation is correct, broken, or ripe for improvement.
  • 6Clear written communication that stands on its own without follow-up questions.
  • 7Regular use of modern LLMs or AI coding assistants in your workflow.
  • 8Nice to have: exposure to LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training.

Who Should Apply

Engineers with at least five years of production coding experience who enjoy dissecting why a solution works, not just whether it compiles. The role rewards people comfortable in big codebases and fluent in Python, TypeScript/JavaScript, or Go, and who can write a crisp justification for their verdict. Candidates weak on written explanations tend to score low here, because the deliverable is often a clear rationale rather than a patch. Applying without real LLM or AI coding tool exposure also hurts fit, even though it is not strictly required. If you prefer building features all day over reviewing and critiquing code, this contract will feel like a poor match.

Salary Insight

Compensation is described as market rate, with candidates asked to provide a specific hourly rate expectation. No figures are stated.

Pay and demand for Machine Learning & AI roles

Aggregated

Typical pay

$72/hour

This role

Pay not disclosed

Most Machine Learning & AI roles pay $50–$95 per hour.

Based on 653 similar roles that publish pay · 168 publish only a top rate; those count at the rate they gave

Typical rangeMedian pay

Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.

Live similar roles
730
Listed in last 30 days
378
Remote
97%

Hiring most right now: micro1 (292) · SME Careers (139) · Mercor (91)

Most requested skills · share of roles

  • llm evaluation
    16%
  • python
    15%
  • ai training
    15%
  • trainer feedback
    11%

Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.

Compare your resume against these roles

Location

Typeremote
LocationNorth America, LATAM, India
This is a remote position

Compensation

Undisclosed by Turing

Comparable roles pay

653 roles

$50$95/ hr

Not this role’s pay. Turing has not published a range.

Confirmed during Turing profile screening.

Check your fit score

Required Skills

pythontypescriptjavascriptgojavagithubci/cdopen sourcellmcode reviewcode analysisllm evaluationcoding agentsrlhfpreference datarubric designpost-trainingcode reviews

Application Tip

State a specific hourly rate with your application and name a project where you reviewed code in a large repository, preferably one involving Python or Go. If you have used LLM evaluation, RLHF, or rubric design, add one sentence on the outcome and quantify it.

Share:

See NearSkill jobs more often in your search

Application & verification flow

  1. 1Instant rubric match

    Your resume is scanned against this role’s requirements to check qualification fit.

  2. 2Screened before the employer sees it

    Only profiles that clear screening are passed on.

  3. 3Outcome by email

    We notify you at the address on your resume once the screening is reviewed.

Test your fit score before applying

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

1mo agoRemotecontract

Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Pay not published
PythonJavaScriptReact+13 more
Turing

Turing

1mo agoRemotecontract

Senior Software Engineer – LLM Evaluation

In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Pay not published
PythonJavaScriptReact+14 more
Micro1

Micro1

7d agoRemotefull-time
Hot

Coding Research Technical Staff - Remote

This full-time remote role centers on advancing how frontier coding agents are evaluated and improved. You will work where AI research, software engineering, and model evaluation meet, creating benchmarks, scoring methods, and data systems that shape how next-generation coding models are measured. The position involves close collaboration with researchers, engineers, and applied AI teams. Python or C++ skills will be central, along with familiarity with LLMs and coding agents.

8–9/hr
LlmsCoding EvaluationAI Evaluation+1 more
Turing

Turing

1mo agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Pay not published
PythonJavaC+++18 more
Micro1

Micro1

2mo agoRemotecontract
Hot

Member of Technical Staff, Coding Research

This role sits at the intersection of AI research and software engineering, focusing on how next-generation coding agents are evaluated and improved. You’ll design the benchmarks, methodologies, and data systems that directly shape the measurement of model performance across complex software engineering tasks. It’s a high-impact position for someone who thrives on building rigorous evaluation frameworks and translating model behavior into actionable improvements.

8–9/hr
LlmsCoding EvaluationAI Evaluation+1 more
Micro1

Micro1

26d agoRemotecontract
Hot

Senior Software Engineer for AI Code Evaluation and Model Review

micro1 is hiring senior software engineers to work with a leading AI lab. Your main tasks involve evaluating AI-generated code for correctness and design, writing challenging coding problems with reference solutions and test cases, and reviewing model reasoning on architecture, debugging, and system design. You will provide detailed technical feedback to help improve model performance on real-world software tasks. This is a remote contractor role with a pay range of $245 to $280 per hour.

245–280/hr
· 100 openings
System DesignCode ReviewDebugging+3 more