
Coding Research Technical Staff - Remote
Listing checked September 15, 2026 · pay as published by Micro1
Overview
This full-time remote role centers on advancing how frontier coding agents are evaluated and improved. You will work where AI research, software engineering, and model evaluation meet, creating benchmarks, scoring methods, and data systems that shape how next-generation coding models are measured. The position involves close collaboration with researchers, engineers, and applied AI teams. Python or C++ skills will be central, along with familiarity with LLMs and coding agents.
What You'll Do8
- 1Create and maintain evaluation frameworks for coding agents, covering benchmark specs, scoring rubrics, and quality standards.
- 2Drive research projects from start to finish that measure and enhance coding model performance on varied software engineering tasks.
- 3Produce high-quality datasets, golden examples, and evaluation protocols for reliable assessment of frontier coding systems.
- 4Examine model behavior and failure modes to spot systematic weaknesses and turn those findings into concrete improvements for training and evaluation.
- 5Build tooling and infrastructure for large-scale experimentation, data generation, review workflows, and evaluation pipelines.
- 6Set best practices for coding-agent assessment that ensure methodological rigor, reproducibility, and measurement quality.
- 7Work with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
- 8Contribute to technical reports, benchmark studies, and client-facing research that communicate model performance and insights.
Requirements8
- 1Strong software engineering background with expertise in Python, C++, or similar languages.
- 2At least 3 years in software engineering, machine learning, AI research, evaluation, or related technical fields.
- 3Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
- 4Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
- 5Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
- 6Strong analytical skills to investigate model behavior and derive insights from complex technical systems.
- 7Excellent written and verbal communication, including the ability to explain technical findings to diverse audiences.
- 8Comfortable working in fast-moving research settings with significant ambiguity and shifting priorities.
Who Should Apply
This role suits engineers and researchers who have spent at least three years building or evaluating AI systems, especially those with hands-on experience in Python or C++ and a track record of designing benchmarks or evaluation protocols. Candidates who enjoy digging into model failure modes and turning those insights into better training data or scoring methods will find the work rewarding. The position is less ideal for someone who prefers stable, well-defined tasks or who has not worked with LLMs, coding agents, or evaluation methodologies. A common reason applicants get rejected is lacking concrete examples of building evaluation frameworks or tooling for large-scale experiments. Another is failing to show how their work directly improved model performance or measurement quality.
Salary Insight
The source lists a rate of $8 to $9 per hour. No other compensation details are provided in the job description.
Pay and demand for Machine Learning & AI roles
AggregatedTypical pay
$75/hour
This role
$8–$9/hr
Most Machine Learning & AI roles pay $55–$100 per hour. This role's pay sits below that range.
Based on 559 similar roles that publish pay · 93 publish only a top rate; those count at the rate they gave
Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.
- Live similar roles
- 631
- Listed in last 30 days
- 292
- Remote
- 96%
Hiring most right now: micro1 (283) · Mercor (86) · SME Careers (63)
Most requested skills · share of roles
- python17%
- technical writing8%
- llm evaluation7%
- ai evaluation6%
Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.
Compare your resume against these rolesLocation
Compensation
$8–9/hr
Required Skills
Application Tip
Highlight specific projects where you designed evaluation frameworks, built tooling for large-scale experiments, or analyzed model failure modes. Quantify the impact, such as the number of benchmarks created or the percentage improvement in model performance, and mention exact technologies like Python, C++, LLMs, or coding agents.
See NearSkill jobs more often in your search
Application & verification flow
1Instant rubric match
Your resume is scanned against this role’s requirements to check qualification fit.
2Screened before the employer sees it
Only profiles that clear screening are passed on.
3Outcome by email
We notify you at the address on your resume once the screening is reviewed.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedMember of Technical Staff, Coding Research
This role sits at the intersection of AI research and software engineering, focusing on how next-generation coding agents are evaluated and improved. You’ll design the benchmarks, methodologies, and data systems that directly shape the measurement of model performance across complex software engineering tasks. It’s a high-impact position for someone who thrives on building rigorous evaluation frameworks and translating model behavior into actionable improvements.

Turing
VerifiedSenior Software Engineer - AI Coding Agent Evaluation
Turing seeks experienced software engineers to assess how well AI coding agents handle real repositories. You will treat the agent like a colleague whose pull requests you review: checking whether it read the task, picked a sound design, and produced working code. The work spans Python, TypeScript/JavaScript, and Go codebases, plus the evaluation rubrics and data pipelines that sharpen model behavior. Researchers rely on your written judgment to turn one-off observations into repeatable quality signals. This is a freelance contract, roughly three months long, with 40 hours per week preferred.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Turing
VerifiedSenior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Micro1
VerifiedSenior Software Engineer for AI Code Evaluation and Model Review
micro1 is hiring senior software engineers to work with a leading AI lab. Your main tasks involve evaluating AI-generated code for correctness and design, writing challenging coding problems with reference solutions and test cases, and reviewing model reasoning on architecture, debugging, and system design. You will provide detailed technical feedback to help improve model performance on real-world software tasks. This is a remote contractor role with a pay range of $245 to $280 per hour.


