
Domain Expert- Music
Overview
This contractor role focuses on refining large language models through expert-level music knowledge. You will create advanced prompts that span music theory, genres, composers, artists, production, and instruments. Work involves assessing AI outputs for factual accuracy, reasoning quality, and completeness, while documenting shortcomings with evidence. You will also build benchmark datasets and adversarial test cases to expose edge cases. The position requires a 40-hour weekly commitment with at least 4 hours of PST overlap.
What You'll Do7
- 1Design complex prompts that draw on music theory, composers, artists, genres, instruments, and production techniques.
- 2Judge AI-generated answers for factual correctness, reasoning depth, completeness, and subtlety in musical topics.
- 3Spot hallucinations, logical errors, stale references, and unusual edge cases in model responses.
- 4Create benchmark datasets and adversarial test cases to probe model weaknesses.
- 5Deliver feedback grounded in authoritative sources and verifiable references.
- 6Work with AI researchers to refine model performance based on your annotations.
- 7Keep annotation standards consistent and document your evaluation process.
Requirements7
- 1Hold a master's degree or higher in any discipline; a music-focused degree such as musicology, ethnomusicology, music theory, music production, or audio engineering is preferred.
- 2Demonstrate deep musical expertise across theory, composition, genres, composers, artists, instruments, music history, notation, production, and modern trends.
- 3Bring at least 3 years of professional work in music education, performance, composition, production, journalism, research, content creation, or audio engineering.
- 4Write clear, accurate English and show solid research and analytical skills.
- 5Evaluate music information with precision, detect factual inconsistencies, and reason across a broad range of musical subjects.
- 6Preferred: hands-on experience with LLMs, generative AI, prompt engineering, or AI evaluation.
- 7Preferred: published research, teaching experience, or industry recognition.
Who Should Apply
The ideal candidate for this contract role holds an advanced degree in music or a related field and has years of hands-on professional experience with music content, whether through performance, production, research, or education. They bring a meticulous eye for factual inaccuracies and can defend their evaluations with strong references. This role is less suitable for early-career professionals without a master's degree or for those without deep knowledge of diverse musical genres and production techniques. Candidates often get rejected when their feedback lacks concrete evidence or when their writing fails to meet the high standard of clarity required. Another common pitfall is not being able to commit to 40 hours per week with the required PST overlap.
Location
Required Skills
Application Tip
Include a short sample evaluation of an AI-generated music response. Point out a factual or reasoning error and back it with a reliable source. Also mention your experience with prompt design or AI evaluation to show you can craft high-quality test cases.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedDomain Expert - Art
The role centers on evaluating and improving Large Language Models (LLMs) using deep knowledge of art history, visual arts, architecture, and design. You will design prompts that range from simple to advanced and assess model answers for factual accuracy, reasoning, completeness, and nuance. Each day involves comparing multiple AI outputs, ranking them, and explaining correct or incorrect responses with reliable references. You will also create benchmark datasets and adversarial test cases to expose hallucinations and logical gaps. The position is an 8-week contractor role requiring 40 hours per week with at least 4 hours of PST overlap.

Turing
VerifiedMusic and Audio experts
Turing, based in San Francisco, California, supplies frontier AI labs and global enterprises with high-quality training data and specialized researchers. This project recruits music professionals to evaluate AI-generated audio by naming the instrument family and the specific instrument heard in short clips. The work is fully remote, with flexible hours inside a two-day window, and strong performers may receive additional project work. Candidates begin with a 15-minute qualifying assessment before onboarding, and formal music education plus experience in audio engineering or production is expected.

Micro1
VerifiedMusic Annotation Expert for Field Recordings and AI Review
Remote contract role focused on listening and labeling field recordings for a music and audio tech project. You’ll document exact audible details such as instruments, tempo, key, and background sounds using a dedicated web platform. Later stages involve assessing AI model outputs against what you heard, with weekly onboarding and calibration.

Turing
VerifiedDomain Expert- Sports
This role focuses on improving Large Language Models (LLMs) by applying deep sports knowledge and analytical rigor. The expert will craft challenging prompts across global sports, leagues, athletes, tournaments, rules, statistics, and analytics. Each response gets checked for factual accuracy, reasoning quality, completeness, and nuance, with evidence-backed feedback. The work also involves building benchmark datasets and adversarial test cases to expose model gaps. This is a remote contractor assignment lasting 8 weeks with a 40-hour weekly commitment.

Turing
VerifiedDomain Expert - TV Show & Movies
Turing is hiring a TV Shows & Movies Domain Expert for an 8-week contractor assignment focused on evaluating and improving large language models. The day-to-day work centers on creating prompts of varying difficulty across entertainment topics, then comparing and ranking AI responses. The expert will judge responses for factual accuracy, reasoning, completeness, and nuance, and will flag hallucinations, outdated information, and edge cases. This remote role is open to US-based candidates who can commit 40 hours per week with 4 hours of daily overlap with PST.

Micro1
VerifiedAi Domain Expert Focused on Domain Knowledge and Evaluation
Remote, part-time contractor role focused on guiding AI systems through real-world input. You’ll review AI outputs for accuracy, craft prompts that test reasoning, and provide precise written feedback to improve model performance. Your domain knowledge matters most, even if you’re not required to have AI prior experience. You’ll join a global, distributed team and help shape how models learn and reason. Bolded areas reflect core competencies like data annotation, prompt engineering, and ethical awareness in AI.

