
Board Game Reasoning Expert (AI Training & Evaluation)
Overview
Turing seeks Board Game Reasoning Experts to build and assess tasks that sharpen the reasoning of frontier AI models. The work centers on board games, game mechanics, and complex rule systems, requiring analysis of strategic decision trees and scenario-based puzzles. Experts create, review, and score game-based datasets, checking AI outputs for logical consistency and rule adherence. The role is a fully remote contractor assignment lasting about 2 months, with a commitment of at least 4 hours per day and 20 hours per week, including 4 hours of overlap with PST.
What You'll Do7
- 1Create and review game-based reasoning tasks that test AI decision-making and rule comprehension.
- 2Analyze board game scenarios, including strategic decision trees and rule-based systems, to identify edge cases and logical gaps.
- 3Evaluate AI-generated responses for correctness, consistency, and reasoning quality across strategy-heavy prompts.
- 4Write prompts, rubrics, and evaluation guidelines for board game and strategy-focused tasks.
- 5Flag logical errors, rule violations, and flawed reasoning in AI outputs with clear justifications.
- 6Support the creation of benchmarks and quality assurance processes for AI evaluation projects.
- 7Work with project teams to refine dataset quality and evaluation methods.
Requirements10
- 1Bachelor's degree in Computer Science, Mathematics, Cognitive Science, Game Design, Philosophy, Economics, or a related analytical field.
- 2At least 2 years of professional or semi-professional experience in board game design, playtesting, tabletop game communities, or strategy-focused environments.
- 3Solid command of logic, probabilistic reasoning, game mechanics, and complex rule systems.
- 4Background in game theory, behavioral economics, decision science, or formal logic.
- 5Familiarity with modern board games, trading card games (TCGs), tabletop RPGs, strategy games, or competitive game systems.
- 6Strong analytical and problem-solving skills, plus clear written communication.
- 7Prior experience in data annotation, AI training, prompt engineering, QA, game design, puzzle design, or rules-based system analysis.
- 8Experience building evaluation rubrics, benchmark datasets, or quality assurance frameworks.
- 9Working knowledge of Python, SQL, or data analysis tools.
- 10Hands-on experience with RLHF, model evaluation, synthetic data generation, or LLM benchmarking.
Who Should Apply
Candidates with a strong background in board games, game design, or tabletop communities who can reason about probabilistic scenarios and rule systems will fit naturally. The role suits people who enjoy careful analysis, writing evaluation criteria, and spotting subtle logical errors in AI outputs. The role is less suitable for those who want to design original games or lead creative work, since the focus is on dataset creation and evaluation. Applications tend to score low when they lack concrete examples of rule-system mastery or prior annotation or benchmark experience. Reviewers also reject candidates who cannot commit to the required PST overlap and weekly hours, so confirm availability before applying.
Location
Required Skills
Application Tip
Show off your rule-system depth by naming specific board games or TCGs you have mastered and describe how you would break down their decision trees. Quantify any playtesting or dataset work, such as the number of tasks reviewed or rubrics created, and call out direct experience with LLM evaluation or RLHF in your application.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

SME Careers
VerifiedPython Data Scientist for AI Reasoning Review
Work remotely as a contract Data Scientist who reviews AI-generated analytical reasoning, code, and model outputs, and crafts precise reference solutions for data problems. You will assess prompts for accuracy and clarity, then write step-by-step explanations that demonstrate correct methods. Rate and compare AI responses on correctness and reasoning quality to guide improvements in machine learning workflows. This remote, hourly role supports projects in data science and analytics for a fast-growing AI data services company.

Turing
VerifiedSoftware Engineer – AI Code Evaluation & Benchmarking (US candidates only)
This contract role places an experienced software engineer inside an evaluation workflow for frontier AI models. You will review AI-generated code for correctness, efficiency, and maintainability, validate solutions against real engineering tasks, and debug failures across different environments. You will also help build and refine evaluation datasets, benchmarks, and grading rubrics. The assignment runs for one month, requires at least 4 hours per day and 20 hours per week with a 4-hour overlap with PST, and is open only to candidates in the US. Work is fully remote.

Micro1
VerifiedEvaluation Specialist for AI Training and Research
Remote contractor role for Evaluation Specialists or Recent Grads who will help train next‑generation AI systems. You will craft original QA pairs, perform rigorous source triangulation, and create multi‑step questions that require synthesis. The project emphasizes high-quality, real‑world input to influence how models learn, reason, and perform. Prior AI experience isn’t required; deep domain knowledge matters and it’s welcomed.

Turing
VerifiedAI Evaluation Specialist – LLM Web Agent Benchmarking (Remote - US)
Your role centers on building hard research challenges for an AI browsing benchmark. You start from a fact that can be verified, then design a natural-language question that would defeat a frontier model even when it has full web access and many attempts. The output includes checkable clues across dates, people, places, organizations, works, events, records, and quantities, plus a validation record of the obvious searches you ran and what they returned. The work is investigative research, not subject-matter expertise or content writing, and the evidence trail carries most of the weight. The contract runs 8 weeks at 40 hours per week, with at least 4 hours of overlap with PST.

Micro1
VerifiedData Science Domain Expert for AI Evaluation and Prompting
Remote contractor role focusing on AI data science with a domain expert lens. You’ll assess AI outputs, refine prompts, and annotate data to support high-quality model training. The work centers on applying deep domain knowledge to review research-style documents, technical reports, and experiment notes, shaping how models learn and reason. Strong writing, precise attention to detail, and independent research are essential as you contribute to rubric-based evaluations and content quality.

Micro1
VerifiedAI Evaluation Specialist
As an AI Evaluation Specialist, you'll help train next-generation AI systems by designing and executing hands-on evaluation tasks. Your insights will directly shape how models learn, reason, and perform on practical computer-based workflows. This is a fully remote contract role where meticulous observation and clear documentation are key.

