
LLM Trainer - Agent Function call
Overview
Work on a project with a foundational LLM company on a 6-week contractor assignment that produces high-quality proprietary data. You will design multi-turn conversations between a user and a smart assistant, playing both sides and simulating function-calling tools like calendar, email, maps, and drive. These dialogues help fine-tune models and benchmark their performance against competitors. The role demands strong technical reasoning, API fluency, and consistent adherence to an internal formatting playbook.
What You'll Do8
- 1Write multi-turn dialogues that mirror realistic user and assistant exchanges, using apps such as calendar, email, maps, and drive.
- 2Play both the user and the assistant, adding assistant tool calls only when a correction is needed.
- 3Decide when and how the assistant invokes available tools, keeping the flow logical and function usage correct.
- 4Create exchanges that show natural language, sensible behavior, and context awareness across several turns.
- 5Build examples that show the assistant completing feasible tasks, identifying infeasible ones, and handling general chat without tools.
- 6Follow the internal playbook to keep every conversation aligned with formatting and quality standards.
- 7Revise conversation samples based on feedback to boost realism, clarity, and training value.
- 8Work with peers and reviewers to keep deliverables consistent and up to standard.
Requirements8
- 1Strong technical reasoning and the ability to model realistic assistant behavior through tool-based APIs.
- 2Skill at breaking down complex tasks and producing realistic dialogues that match user expectations and assistant limits.
- 3Comfort with any programming language or tech stack; a solid grasp of APIs, JSON, and logical thinking matters more than specific tools.
- 4Excellent written English, with clear tone and coherent instructions.
- 5Creativity and attention to detail when building realistic scenarios and responses.
- 6Familiarity with LLMs, virtual assistants, or function-calling frameworks is a plus.
- 7Consistent adherence to detailed guidelines and formatting standards.
- 8At least 3 years of professional experience in a technical or analytical field.
Who Should Apply
The ideal candidate combines strong technical reasoning with a writer's touch, crafting realistic multi-turn conversations that show an assistant using tools the right way. This role suits people who enjoy the detail work of simulating user and assistant behavior in written form, with a solid grasp of APIs and JSON. People who prefer hands-on coding over writing structured dialogue examples will find this role less suitable, as will those who struggle to follow strict formatting guidelines. Common rejection reasons include vague or sloppy written communication, missing the 3+ years of technical experience, and submitting samples that don't demonstrate logical tool selection across multiple turns.
Location
Required Skills
Application Tip
During the application, submit a short sample dialogue that demonstrates correct function-call syntax, realistic user turns, and a clear decision point where the assistant recognizes an infeasible request. Mention your experience with JSON and any prior work with LLM evaluation or data labeling.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedLLM Java Developer
This role centers on building and maintaining Java back-end components that support LLM training and refinement. Daily work includes running model evaluations, scoring AI responses, and writing clear rationales for those judgments. You will also contribute to Supervised Fine-Tuning (SFT) datasets and collaborate on Reinforcement Learning with Human Feedback (RLHF). The position is a fully remote contractor assignment with a required 4-hour overlap with PST and a commitment of 20, 30, or 40 hours per week.

Turing
VerifiedLLM C/ C++ Developer
A company building next-generation dialog agents for education, entertainment, and question-answering needs a C++ engineer to review and validate AI-generated code. In this contract role, you will debug C/C++ code produced by AI systems, help define new features with cross-functional teams, and contribute to training LLM back-end components. You will also improve public GitHub repositories and mentor other developers through code review. The position is remote and contract-based, with flexible hours and a two-step interview process that includes a 60-minute technical session.

Turing
VerifiedLLM Go Developer
An established company building the next generation of dialog agents for education, entertainment, and question-answering needs a Go engineer to review and validate AI-generated code. You will collaborate with cross-functional teams to define, design, and deliver new features, and apply your Go expertise to resolve difficult coding issues that surface during AI validation. The role includes managing development cycles, setting goals and deadlines, and giving teammates constructive code feedback. As a contractor, you get flexibility in contract duration and committed hours, and the process includes two internal interviews: a 60-minute technical round and a 15-30 minute cultural and offer discussion.

Turing
VerifiedSenior Software Engineer – LLM Evaluation
In this contract role, you will build and refine training datasets that help large language models improve their coding skills. You will write, correct, and evaluate code in Python, JavaScript (ReactJS), C/C++, Java, Rust, and Go, collaborating with researchers and cross-functional teams. Your daily work includes assessing AI-generated code for efficiency, scalability, and reliability, plus building verification agents that catch error patterns. The one-month engagement runs 10-40 hours per week with partial PST overlap.

Turing
VerifiedAI Evaluation Specialist – LLM Web Agent Benchmarking (Remote - US)
Your role centers on building hard research challenges for an AI browsing benchmark. You start from a fact that can be verified, then design a natural-language question that would defeat a frontier model even when it has full web access and many attempts. The output includes checkable clues across dates, people, places, organizations, works, events, records, and quantities, plus a validation record of the obvious searches you ran and what they returned. The work is investigative research, not subject-matter expertise or content writing, and the evidence trail carries most of the weight. The contract runs 8 weeks at 40 hours per week, with at least 4 hours of overlap with PST.

Turing
VerifiedSenior LLM Engineer
A remote role based in India centers on designing and building Generative AI and LLM systems with Python and Langchain. The engineer will create RAG pipelines, prompt techniques, and agent-based workflows that run in production. The position expects 7-12 years of experience and close work with engineering teams, business SMEs, and data teams to shape the LLM roadmap. Strong SQL and cloud familiarity across AWS, Azure, or GCP support the day-to-day work.

