Overview
This role centers on building and maintaining Java back-end components that support LLM training and refinement. Daily work includes running model evaluations, scoring AI responses, and writing clear rationales for those judgments. You will also contribute to Supervised Fine-Tuning (SFT) datasets and collaborate on Reinforcement Learning with Human Feedback (RLHF). The position is a fully remote contractor assignment with a required 4-hour overlap with PST and a commitment of 20, 30, or 40 hours per week.
What You'll Do11
- 1Design and write maintainable code for training and optimizing AI models.
- 2Run evaluation suites to benchmark model performance and analyze the results.
- 3Rank AI responses across different domains against defined quality criteria.
- 4Produce detailed explanations that justify each evaluation and rating.
- 5Lead dataset creation and maintenance for Supervised Fine-Tuning (SFT).
- 6Work with researchers and annotators to execute RLHF and improve reward models.
- 7Design new evaluation methods that strengthen model alignment with user needs and ethical guidelines.
- 8Craft and refine ideal responses to improve clarity, relevance, and technical accuracy.
- 9Perform peer reviews of code and documentation and give constructive feedback.
- 10Partner with cross-functional teams to improve model performance and product features.
- 11Evaluate and integrate new tools and techniques into the AI training pipeline.
Requirements8
- 1Bachelor's or Master's degree in Engineering, Computer Science, or equivalent practical experience.
- 2Strong proficiency in Java, including its syntax and standard conventions.
- 3Experience building web applications with modular, scalable architectures and a focus on code readability, security, and stability through testing.
- 4Ability to write clear, well-organized, correctly annotated code that is easy to classify and review.
- 5Familiarity with training LLM models using high-quality back-end components and modern coding best practices.
- 6Experience participating in code reviews and maintaining high code quality standards.
- 7Nice to have: prior experience in software quality assurance and test planning.
- 8Excellent spoken and written English communication skills.
Who Should Apply
The ideal candidate has strong Java experience and can show real projects where they built scalable back-end systems. They are comfortable writing detailed evaluation rationales and have some exposure to LLM training or fine-tuning workflows. This role is not a good fit for someone who wants a permanent, benefits-eligible position or who cannot maintain a fixed 4-hour overlap with PST. Rejections often come from candidates who describe Java skills without tangible examples, or who produce vague, unstructured evaluation explanations during the interview.
Location
Required Skills
Application Tip
Come to the technical interview prepared to walk through a Java back-end project you built, highlighting how you handled modularity, testing, and security. Also practice writing a concise evaluation rationale for a sample AI response, because the interview will probe how you justify quality judgments.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedSenior Python Developer
This contract role supports a foundational LLM company by producing high-quality data used to fine-tune and benchmark their models. You will write Python solutions to code-based prompts, evaluate responses from two model versions, and create detailed rationale for ranking decisions. Day-to-day work involves SFT and RLHF data generation, designing evaluation strategies, and reviewing code with a small team. The position does not involve building or fine-tuning models directly, but your outputs directly inform their improvement.

Turing
VerifiedLLM C/ C++ Developer
A company building next-generation dialog agents for education, entertainment, and question-answering needs a C++ engineer to review and validate AI-generated code. In this contract role, you will debug C/C++ code produced by AI systems, help define new features with cross-functional teams, and contribute to training LLM back-end components. You will also improve public GitHub repositories and mentor other developers through code review. The position is remote and contract-based, with flexible hours and a two-step interview process that includes a 60-minute technical session.

Turing
VerifiedSenior Software Engineer – LLM Evaluation (US/Canada/WEU based)
Turing, a San Francisco-based research accelerator, is hiring a contract software engineer to evaluate AI-generated code and build datasets for large language models. You will curate code examples and write precise corrections in Python, JavaScript/ReactJS, C/C++, Rust, plus Java and Go. The role includes scoring model outputs for efficiency, scalability, and reliability, designing verification mechanisms, and partnering with researchers to strengthen enterprise coding solutions. This remote engagement runs one month, offers 10 to 40 flexible hours per week, and accepts candidates in the US, Canada, and Western Europe.

Turing
VerifiedLLM Go Developer
An established company building the next generation of dialog agents for education, entertainment, and question-answering needs a Go engineer to review and validate AI-generated code. You will collaborate with cross-functional teams to define, design, and deliver new features, and apply your Go expertise to resolve difficult coding issues that surface during AI validation. The role includes managing development cycles, setting goals and deadlines, and giving teammates constructive code feedback. As a contractor, you get flexibility in contract duration and committed hours, and the process includes two internal interviews: a 60-minute technical round and a 15-30 minute cultural and offer discussion.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Turing
VerifiedJavaScript / TypeScript Full-Stack Developer
This role centers on building and refining the code that trains and evaluates AI models for leading LLM companies. You'll write production-grade JavaScript/TypeScript across the stack, using back-end frameworks like Node.js or Nest.js and front-end libraries such as React, Vue, or Angular. Daily work includes running model evaluations, ranking responses, and writing rationales that feed Supervised Fine-Tuning and RLHF processes. You'll also review peers' code, design evaluation strategies, and collaborate with researchers to align model behavior with user needs and ethical guidelines.

