Turing
TuringVerified listing
Remote

Domain Expert - TV Show & Movies

Remote
Posted August 26, 2026
contract

Overview

Turing is hiring a TV Shows & Movies Domain Expert for an 8-week contractor assignment focused on evaluating and improving large language models. The day-to-day work centers on creating prompts of varying difficulty across entertainment topics, then comparing and ranking AI responses. The expert will judge responses for factual accuracy, reasoning, completeness, and nuance, and will flag hallucinations, outdated information, and edge cases. This remote role is open to US-based candidates who can commit 40 hours per week with 4 hours of daily overlap with PST.

What You'll Do9

  • 1Write advanced prompts that cover film, television, streaming, actors, directors, genres, awards, franchises, and entertainment history.
  • 2Assess AI-generated responses for factual accuracy, reasoning quality, completeness, and nuance.
  • 3Spot hallucinations, logical inconsistencies, outdated facts, and edge cases in model outputs.
  • 4Build benchmark datasets and adversarial test cases to uncover model weaknesses.
  • 5Explain why a model response is correct or incorrect using authoritative references.
  • 6Compare multiple AI responses, rank them, and justify each ranking.
  • 7Work with AI researchers to turn evaluation findings into model improvements.
  • 8Maintain high annotation quality and document your evaluation decisions.
  • 9Identify ambiguous questions and propose clearer prompt alternatives.

Requirements7

  • 1Hold a master's degree or higher in any field; entertainment-related disciplines like Film Studies, Television Studies, Media Studies, Journalism, Communications, English, or Screen Studies are preferred.
  • 2Show strong knowledge of movies and television, including actors, directors, creators, genres, franchises, awards, streaming platforms, and industry developments.
  • 3Bring 3+ years of professional experience in entertainment journalism, film or TV criticism, entertainment media, content creation, research, analysis, or reporting.
  • 4Write and research in excellent English with strong analytical reasoning.
  • 5Check movie and television information with consistent attention to detail.
  • 6Prefer candidates with experience using LLMs, generative AI, prompt engineering, or AI evaluation.
  • 7Published research, industry recognition, or teaching experience is a plus.

Who Should Apply

This role belongs to an entertainment expert who enjoys close factual verification and can defend judgments with references. Someone with a graduate degree and 3+ years in criticism, journalism, or media research will fit the structured evaluation workflow. The role is less suitable for writers who want open-ended creative assignments rather than repeated scoring and annotation tasks. Candidates often fall short when they cannot cite authoritative references for their verdicts or when their entertainment knowledge is broad but not deep enough for nuanced prompt evaluation.

Location

Typeremote
LocationRemote
Eligible countriesUS
This is a remote position

Required Skills

llmgenerative aiprompt engineeringai evaluationfilm studiestelevision studiesmedia studiesjournalismcommunicationsenglish literaturescreen studiesentertainment historystreaming platformsbenchmarkingadversarial testingdata annotationresearchfilm criticismentertainment journalismcommunication

Application Tip

Show concrete prompt engineering or LLM evaluation experience in your application. Include a sample film or TV evaluation that cites authoritative sources and demonstrates how you would rank model responses. Mention any past work with benchmark datasets or annotation volumes if you have them.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

9d agoRemotecontract

Domain Expert- Politics

Turing seeks a Politics Domain Expert to evaluate and improve Large Language Models with a focus on political science, governance, elections, public policy, and international relations. The role involves designing challenging prompts that test factual accuracy, reasoning, completeness, and nuance in AI-generated responses. You will build benchmark datasets and adversarial test cases to expose hallucinations, logical inconsistencies, outdated information, and edge cases. Work happens remotely on a contractor basis for 8 weeks at 40 hours per week with a 4-hour PST overlap.

Competitive salary
Political ScienceGovernanceElections+12 more
Turing

Turing

9d agoRemotecontract

Domain Expert - Art

The role centers on evaluating and improving Large Language Models (LLMs) using deep knowledge of art history, visual arts, architecture, and design. You will design prompts that range from simple to advanced and assess model answers for factual accuracy, reasoning, completeness, and nuance. Each day involves comparing multiple AI outputs, ranking them, and explaining correct or incorrect responses with reliable references. You will also create benchmark datasets and adversarial test cases to expose hallucinations and logical gaps. The position is an 8-week contractor role requiring 40 hours per week with at least 4 hours of PST overlap.

Competitive salary
Art HistoryFine ArtsVisual Arts+14 more
Turing

Turing

9d agoRemotecontract

Domain Expert- Sports

This role focuses on improving Large Language Models (LLMs) by applying deep sports knowledge and analytical rigor. The expert will craft challenging prompts across global sports, leagues, athletes, tournaments, rules, statistics, and analytics. Each response gets checked for factual accuracy, reasoning quality, completeness, and nuance, with evidence-backed feedback. The work also involves building benchmark datasets and adversarial test cases to expose model gaps. This is a remote contractor assignment lasting 8 weeks with a 40-hour weekly commitment.

Competitive salary
Large Language ModelsGenerative AIPrompt Engineering+9 more
Micro1

Micro1

1mo agoRemotecontract
Hot

Data Science Domain Expert for AI Evaluation and Prompting

Remote contractor role focusing on AI data science with a domain expert lens. You’ll assess AI outputs, refine prompts, and annotate data to support high-quality model training. The work centers on applying deep domain knowledge to review research-style documents, technical reports, and experiment notes, shaping how models learn and reason. Strong writing, precise attention to detail, and independent research are essential as you contribute to rubric-based evaluations and content quality.

100–200/hr
· 15 openings
Critical ThinkingAnalytical ReasoningAttention To Detail+22 more
Micro1

Micro1

1mo agoRemotecontract
Hot

Ai Domain Expert Focused on Domain Knowledge and Evaluation

Remote, part-time contractor role focused on guiding AI systems through real-world input. You’ll review AI outputs for accuracy, craft prompts that test reasoning, and provide precise written feedback to improve model performance. Your domain knowledge matters most, even if you’re not required to have AI prior experience. You’ll join a global, distributed team and help shape how models learn and reason. Bolded areas reflect core competencies like data annotation, prompt engineering, and ethical awareness in AI.

140–200/hr
· 15 openings
Data AnnotationPrompt EngineeringCritical Thinking+10 more
Turing

Turing

9d agoRemotecontract

Domain Expert- Music

This contractor role focuses on refining large language models through expert-level music knowledge. You will create advanced prompts that span music theory, genres, composers, artists, production, and instruments. Work involves assessing AI outputs for factual accuracy, reasoning quality, and completeness, while documenting shortcomings with evidence. You will also build benchmark datasets and adversarial test cases to expose edge cases. The position requires a 40-hour weekly commitment with at least 4 hours of PST overlap.

Competitive salary
Music TheoryMusicologyEthnomusicology+13 more