Micro1
Micro1Verified listing
RemoteHot

Coding Research Technical Staff - Remote

8–9/hr
Remote
Posted September 15, 2026
full-time
Not Sure? Upload your resume to see every role you match

Listing checked September 15, 2026 · pay as published by Micro1

Overview

This full-time remote role centers on advancing how frontier coding agents are evaluated and improved. You will work where AI research, software engineering, and model evaluation meet, creating benchmarks, scoring methods, and data systems that shape how next-generation coding models are measured. The position involves close collaboration with researchers, engineers, and applied AI teams. Python or C++ skills will be central, along with familiarity with LLMs and coding agents.

What You'll Do8

  • 1Create and maintain evaluation frameworks for coding agents, covering benchmark specs, scoring rubrics, and quality standards.
  • 2Drive research projects from start to finish that measure and enhance coding model performance on varied software engineering tasks.
  • 3Produce high-quality datasets, golden examples, and evaluation protocols for reliable assessment of frontier coding systems.
  • 4Examine model behavior and failure modes to spot systematic weaknesses and turn those findings into concrete improvements for training and evaluation.
  • 5Build tooling and infrastructure for large-scale experimentation, data generation, review workflows, and evaluation pipelines.
  • 6Set best practices for coding-agent assessment that ensure methodological rigor, reproducibility, and measurement quality.
  • 7Work with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
  • 8Contribute to technical reports, benchmark studies, and client-facing research that communicate model performance and insights.

Requirements8

  • 1Strong software engineering background with expertise in Python, C++, or similar languages.
  • 2At least 3 years in software engineering, machine learning, AI research, evaluation, or related technical fields.
  • 3Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • 4Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
  • 5Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
  • 6Strong analytical skills to investigate model behavior and derive insights from complex technical systems.
  • 7Excellent written and verbal communication, including the ability to explain technical findings to diverse audiences.
  • 8Comfortable working in fast-moving research settings with significant ambiguity and shifting priorities.

Who Should Apply

This role suits engineers and researchers who have spent at least three years building or evaluating AI systems, especially those with hands-on experience in Python or C++ and a track record of designing benchmarks or evaluation protocols. Candidates who enjoy digging into model failure modes and turning those insights into better training data or scoring methods will find the work rewarding. The position is less ideal for someone who prefers stable, well-defined tasks or who has not worked with LLMs, coding agents, or evaluation methodologies. A common reason applicants get rejected is lacking concrete examples of building evaluation frameworks or tooling for large-scale experiments. Another is failing to show how their work directly improved model performance or measurement quality.

Salary Insight

The source lists a rate of $8 to $9 per hour. No other compensation details are provided in the job description.

Pay and demand for Machine Learning & AI roles

Aggregated

Typical pay

$75/hour

This role

$8–$9/hr

Most Machine Learning & AI roles pay $55–$100 per hour. This role's pay sits below that range.

Based on 559 similar roles that publish pay · 93 publish only a top rate; those count at the rate they gave

Typical rangeMedian payThis role

Rates shown per hour. Yearly and monthly pay converted; one-time fees and non-USD pay are not included.

Live similar roles
631
Listed in last 30 days
292
Remote
96%

Hiring most right now: micro1 (283) · Mercor (86) · SME Careers (63)

Most requested skills · share of roles

  • python
    17%
  • technical writing
    8%
  • llm evaluation
    7%
  • ai evaluation
    6%

Figures from Machine Learning & AI roles live on NearSkill when this page loaded. A role can close before you apply, so check the listing itself.

Compare your resume against these roles

Location

TypeRemote
LocationRemote
This is a remote position

Compensation

$8–9/hr

Required Skills

LLMsCoding EvaluationAI EvaluationML Systems

Application Tip

Highlight specific projects where you designed evaluation frameworks, built tooling for large-scale experiments, or analyzed model failure modes. Quantify the impact, such as the number of benchmarks created or the percentage improvement in model performance, and mention exact technologies like Python, C++, LLMs, or coding agents.

Share:

See NearSkill jobs more often in your search

Application & verification flow

  1. 1Instant rubric match

    Your resume is scanned against this role’s requirements to check qualification fit.

  2. 2Screened before the employer sees it

    Only profiles that clear screening are passed on.

  3. 3Outcome by email

    We notify you at the address on your resume once the screening is reviewed.

Test your fit score before applying

Similar open positions

Explore active roles that match your skills and interests.

Micro1

Micro1

2mo agoRemotecontract
Hot

Member of Technical Staff, Coding Research

This role sits at the intersection of AI research and software engineering, focusing on how next-generation coding agents are evaluated and improved. You’ll design the benchmarks, methodologies, and data systems that directly shape the measurement of model performance across complex software engineering tasks. It’s a high-impact position for someone who thrives on building rigorous evaluation frameworks and translating model behavior into actionable improvements.

8–9/hr
LlmsCoding EvaluationAI Evaluation+1 more
Turing

Turing

11h agoRemotecontract

Senior Software Engineer - AI Coding Agent Evaluation

Turing seeks experienced software engineers to assess how well AI coding agents handle real repositories. You will treat the agent like a colleague whose pull requests you review: checking whether it read the task, picked a sound design, and produced working code. The work spans Python, TypeScript/JavaScript, and Go codebases, plus the evaluation rubrics and data pipelines that sharpen model behavior. Researchers rely on your written judgment to turn one-off observations into repeatable quality signals. This is a freelance contract, roughly three months long, with 40 hours per week preferred.

Pay not published
PythonTypeScriptJavaScript+15 more
Turing

Turing

1mo agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Pay not published
PythonJavaC+++18 more
Turing

Turing

1mo agoRemotecontract

Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Pay not published
PythonJavaScriptReact+13 more
Turing

Turing

1mo agoRemotecontract

Senior Software Engineer – LLM Evaluation

In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Pay not published
PythonJavaScriptReact+14 more
Micro1

Micro1

26d agoRemotecontract
Hot

Senior Software Engineer for AI Code Evaluation and Model Review

micro1 is hiring senior software engineers to work with a leading AI lab. Your main tasks involve evaluating AI-generated code for correctness and design, writing challenging coding problems with reference solutions and test cases, and reviewing model reasoning on architecture, debugging, and system design. You will provide detailed technical feedback to help improve model performance on real-world software tasks. This is a remote contractor role with a pay range of $245 to $280 per hour.

245–280/hr
· 100 openings
System DesignCode ReviewDebugging+3 more