Micro1
Micro1Verified listing
RemoteHotAI & Machine LearningSoftware Engineering

Member of Technical Staff, Coding Research | $8-$9/hr Remote

8–9/hr
Remote
Posted July 24, 2026
contract

Overview

This role sits at the intersection of AI research and software engineering, focusing on how next-generation coding agents are evaluated and improved. You’ll design the benchmarks, methodologies, and data systems that directly shape the measurement of model performance across complex software engineering tasks. It’s a high-impact position for someone who thrives on building rigorous evaluation frameworks and translating model behavior into actionable improvements.

What You'll Do8

  • 1Own the design of evaluation frameworks for coding agents, including benchmark specs, scoring rubrics, and quality standards.
  • 2Lead end-to-end research initiatives that measure and boost coding model performance on a wide range of engineering challenges.
  • 3Create high-quality datasets, golden examples, and evaluation protocols to ensure reliable assessment of frontier AI systems.
  • 4Analyze model behavior and failure modes, pinpointing systematic weaknesses and turning findings into training and evaluation improvements.
  • 5Build tooling and infrastructure to support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
  • 6Establish best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
  • 7Collaborate with researchers and applied AI teams to design experiments and evaluate emerging model capabilities.
  • 8Contribute to technical reports, benchmark studies, and client-facing research that communicate model performance insights.

Requirements8

  • 1Strong software engineering background with proficiency in Python, C++, or similar languages.
  • 2At least 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical field.
  • 3Proven experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • 4Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
  • 5Ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
  • 6Strong analytical skills to investigate model behavior and derive insights from complex technical systems.
  • 7Excellent written and verbal communication skills, with the ability to articulate technical findings to diverse audiences.
  • 8Comfort operating in fast-moving research environments with significant ambiguity and evolving priorities.

Who Should Apply

You’re the ideal candidate if you love working at the intersection of AI research and software engineering, are passionate about measuring and improving coding model performance, and enjoy building robust evaluation systems from scratch. You thrive in ambiguous, fast-paced environments where you can independently drive projects from concept to execution. A strong track record in benchmark design, data systems, or agentic workflows will set you apart.

Salary Insight

The base salary for this full-time remote position ranges from $160,000 to $220,000 USD, plus equity compensation and performance-based bonuses. micro1 also offers a comprehensive benefits package including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with company match, and additional remote-first perks.

Location

TypeRemote
LocationRemote
This is a remote position

Required Skills

LLMsCoding EvaluationAI EvaluationML Systems

Application Tip

In your application, detail a specific benchmark or evaluation framework you’ve built for coding models—explain the metrics you chose, why, and how your work led to measurable improvements in model performance or reliability. Concrete examples of tooling or pipeline automation you’ve created will also help you stand out.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Micro1

Micro1

1mo agoRemotecontract
Hot

Member of Technical Staff, Legal Research

This role sits at the intersection of legal expertise and AI research. You'll design evaluation frameworks that measure how well large language models handle complex legal reasoning, working remotely to help build the next generation of enterprise AI for the legal industry. Your work will directly shape how AI agents perform tasks like statutory interpretation, contract analysis, and workflow automation.

7–8/hr
Legal ResearchEnterprise AIBenchmarking+1 more
Turing

Turing

21d agoRemotecontract

Senior Software Engineer – LLM Evaluation

In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Competitive salary
PythonJavaScriptReact+14 more
Turing

Turing

23d agoRemotecontract

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Competitive salary
PythonJavaC+++18 more
Micro1

Micro1

30d agoRemotefull-time
Hot

Member of Technical Staff, Frontier AI

We’re looking for a Member of Technical Staff to own the bridge between research, data, and production AI systems. In this hands-on role, you’ll drive model and system improvements through rigorous evaluation, failure analysis, and fast iteration. You’ll collaborate with researchers and domain experts to turn experimental insights into measurable, real-world performance gains.

100–130/hr
Research Signal JudgmentML Oriented Data DesignOps To Research Translation+1 more
Micro1

Micro1

1mo agoRemotecontract

AI Evaluation Specialist

As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.

20–35/hr
· 50 openings
Rubric Based EvaluationStructured Observation And ReportingHigh Attention To Detail+2 more
Micro1

Micro1

8d agoRemotecontract
Hot

Senior Software Engineer for AI Code Evaluation and Model Review

Senior Software Engineer contracted for remote work with micro1 in collaboration with a top AI lab. You’ll assess AI-generated code for correctness and quality, craft demanding coding problems with reference solutions and test cases, and review model reasoning on architecture, debugging, and system design. Your feedback will guide model performance improvements on real-world software tasks. Expect to evaluate code, write tests, and communicate technical insights clearly.

245–280/hr
· 5 openings
System DesignCode ReviewDebugging+3 more