
AI Developer Trace Task Auditor | $70-$90/hr Remote
Overview
Every coding trace you receive comes from AI-assisted developer sessions. You will assess each trace for code correctness, workflow soundness, and reasoning, then write clear feedback that follows a defined rubric. A frontier AI lab uses these evaluations to train and refine its models, so your judgments shape model behavior. The role is remote within the United States and pays $70 to $90 per hour.
What You'll Do6
- 1Review complete coding traces created by AI-assisted developer tools.
- 2Judge each session for code correctness, workflow soundness, and reasoning quality.
- 3Write structured feedback that maps to a specific grading rubric.
- 4Identify bugs, logic errors, and deviations from engineering best practices in multi-step coding workflows.
- 5Verify whether the AI tool followed the developer's specification and intent across the session.
- 6Document your findings to support model training and evaluation.
Requirements8
- 13+ years of professional software development experience.
- 2Daily use of AI-assisted coding tools such as Cursor, GitHub Copilot, Claude Code, or equivalent.
- 3Experience with agentic or spec-driven development workflows.
- 4Strong code-reading and debugging skills in full-stack or backend systems.
- 5Ability to evaluate long coding trajectories for correctness and engineering best practice.
- 6Familiarity with Kiro or Amazon CodeCatalyst (preferred).
- 7Prior experience grading or evaluating AI-generated code (preferred).
- 8Contributions to developer tooling (preferred).
Who Should Apply
Developers who use AI coding tools daily and enjoy reviewing code from others will fit this role well. You need at least 3 years of professional development experience plus a sharp eye for logic errors across full-stack or backend systems. This job is less suitable for developers who prefer writing new code over analyzing existing traces. A common rejection reason is a lack of hands-on agentic workflow experience. Another is providing feedback that is too subjective to fit a rubric.
Salary Insight
Pay is $70 to $90 per hour. This rate matches senior-level contract work and accounts for the specialized focus on AI-assisted coding evaluation and rubric-based feedback.
Location
Required Skills
Application Tip
In your application, list the AI coding tools you use daily and describe a specific session where you caught a critical bug in AI-generated code. If you have graded or evaluated code before, mention the rubric you used.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedAgentic AI Expert
micro1 is looking for experienced technical professionals to help train next-generation AI systems by working hands-on with autonomous AI coding agents. In this contract role, you'll apply your software engineering expertise to guide AI models, evaluate their outputs, and document best practices. No prior AI experience is required—your deep domain knowledge is what matters most.

Micro1
VerifiedAI Evaluation Specialist
As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Micro1
VerifiedAi Trainer for Domain Experts Remote Contract
A remote contractor role focused on shaping AI performance through real-world input. You’ll review and create examples for model outputs, collaborate with a distributed team, and provide clear written rationales to support learning objectives. Your domain knowledge in a specialized field drives the quality of prompts, tasks, and data annotations that guide AI reasoning. This project relies on precise feedback, adaptable methods, and thorough documentation of progress.

SME Careers
VerifiedC Engineer for AI Code Review and Reference Apps
Remote contract role for a seasoned C engineer focused on AI data workflows. You will review AI-generated C code, evaluate low-level system designs, and craft high-quality reference implementations plus step-by-step reasoning to illuminate complex problems. Expect to assess accuracy, memory management, and concurrency while ensuring alignment with prompts. This fully remote position offers hourly compensation through SME Careers within a growing AI data services network.

Micro1
VerifiedAI Engineer
micro1 hires software engineers to build evaluation environments that push AI systems into realistic engineering workflows. As a remote contractor, you set your own schedule at roughly 15 hours per week and focus on one thing: creating reinforcement learning environments that force an AI model to use Model Context Protocol (MCP) tools while solving bugs, implementing features, or refactoring code. Each environment needs deterministic verification and a golden reference solution so results are measurable and repeatable. Prior AI experience isn't required; your coding background and judgment are what the project needs.

