
Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)
Overview
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.
What You'll Do8
- 1Review AI-generated code for correctness, efficiency, maintainability, and requirement adherence.
- 2Analyze software engineering tasks and judge whether proposed solutions meet expected outcomes.
- 3Debug code, reproduce reported issues, and verify fixes across different programming environments.
- 4Assess model-generated explanations and implementation approaches for technical accuracy.
- 5Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics.
- 6Identify edge cases, failure modes, and recurring weaknesses in AI coding systems.
- 7Document findings and give structured feedback to keep evaluation quality consistent.
- 8Work with project teams to define quality standards and evaluation methodologies.
Requirements10
- 1Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical field.
- 2At least 3 years of professional software engineering experience.
- 3Strong proficiency in one or more of these languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
- 4Solid understanding of data structures, algorithms, software design principles, and debugging methodologies.
- 5Proven experience performing code reviews and evaluating code quality in production or large-scale codebases.
- 6Ability to analyze complex technical problems and judge solution correctness with minimal supervision.
- 7Familiarity with Git and standard modern development workflows.
- 8Clear written communication and careful attention to detail.
- 9Bonus: experience with AI/ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects.
- 10Highly preferred: prior experience evaluating AI-generated code, creating benchmarks, or assessing software quality.
Who Should Apply
The candidate who fits this role is a software engineer with at least 3 years of professional experience, strong code review skills, and real comfort in Python and Docker. They enjoy debugging, analyzing why a solution fails, and turning those findings into structured feedback for AI model improvement. This role is less suitable for engineers who want a long-term staff position with benefits, or who prefer building features over evaluating and benchmarking code. Common rejection reasons include failing the online automated coding challenge in Python and Docker, and not being available for the required 4-hour daily PST overlap or the one-month contract commitment.
Location
Required Skills
Application Tip
Before applying, practice Python coding challenges that run inside Docker containers and study the RHLF-style evaluation format used for AI response ranking. In your application, highlight specific examples of code review and bug-finding work with measurable outcomes, such as error rates reduced or defects caught before release.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Micro1
VerifiedSenior Software Engineer for AI Code Evaluation and Model Review
Senior Software Engineer contracted for remote work with micro1 in collaboration with a top AI lab. You’ll assess AI-generated code for correctness and quality, craft demanding coding problems with reference solutions and test cases, and review model reasoning on architecture, debugging, and system design. Your feedback will guide model performance improvements on real-world software tasks. Expect to evaluate code, write tests, and communicate technical insights clearly.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

SME Careers
VerifiedC Engineer for AI Code Review and Reference Apps
Remote contract role for a seasoned C engineer focused on AI data workflows. You will review AI-generated C code, evaluate low-level system designs, and craft high-quality reference implementations plus step-by-step reasoning to illuminate complex problems. Expect to assess accuracy, memory management, and concurrency while ensuring alignment with prompts. This fully remote position offers hourly compensation through SME Careers within a growing AI data services network.

Micro1
VerifiedAI Evaluation Specialist
As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.

Micro1
VerifiedMember of Technical Staff, Coding Research
This role sits at the intersection of AI research and software engineering, focusing on how next-generation coding agents are evaluated and improved. You’ll design the benchmarks, methodologies, and data systems that directly shape the measurement of model performance across complex software engineering tasks. It’s a high-impact position for someone who thrives on building rigorous evaluation frameworks and translating model behavior into actionable improvements.

