Turing
TuringVerified listing
Remote

AI Quality Analyst - English

Remote
Posted August 11, 2026
contract

Overview

This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

What You'll Do8

  • 1Design and run multi-turn conversation prompts that span 1 to 5 turns and require the model to draw on your personal context.
  • 2Assess each model response against the intent you set in the starting prompt to determine whether personalization was applied.
  • 3Flag grounding failures when the model makes claims about you that lack supporting evidence or depend on flawed inferences.
  • 4Evaluate integration quality to ensure personal details are woven into responses without robotic overnarrating.
  • 5Compare two model responses side-by-side (SxS) and stack-rank them by helpfulness, ease of use, and enjoyment.
  • 6Write concise, defensible rationales that reference exact turn numbers to justify your rankings.
  • 7Inspect Debug Info to confirm the model used chat summaries and data sources as intended.
  • 8Delete evaluation conversations after each session to prevent them from affecting future chat history.

Requirements12

  • 1Native or near-native English reading and writing ability, since English is the project's focus language.
  • 2Willingness to connect a primary personal Google account (not a test account) and enable personal data sources for real-world evaluation.
  • 3Openness to work at least 4 hours per day, up to 40 hours per week, with 4 hours overlapping PST, plus full-time availability in your local time zone.
  • 4Sharp analytical thinking to judge nuanced, ambiguous AI responses and spot personalization errors, poor inferences, and forced connections.
  • 5Experience designing creative, multi-turn prompts that use personal context to challenge model capabilities.
  • 6Keen attention to detail for reviewing side-by-side (SxS) responses and detecting subtle differences in naturalness and overnarrating.
  • 7Strong written communication skills to produce structured rationales with explicit turn numbers.
  • 8Ability to provide constructive feedback and detailed annotations to support model improvement.
  • 9Self-motivation and comfort working on your own in a remote environment.
  • 10Reliable desktop or laptop and a stable internet connection.
  • 11BS/BA degree or equivalent experience in a related analytical field such as policy, law, ethics, linguistics, journalism, or computer science.
  • 12Previous experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.

Who Should Apply

The candidate who thrives here is comfortable letting Gemini read their personal Gmail and Search history, enjoys crafting creative multi-turn prompts, and can defend a side-by-side ranking with precise turn-level evidence. This role is not a fit for anyone who wants to keep their personal Google data separate from work, or who prefers repetitive labeling over open-ended evaluation and written rationales. Rejections often happen when a candidate misses the 24-hour assessment window or submits rationales that are too vague to show they caught grounding and integration issues.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

geminigmailgoogle searchyoutubeprompt engineeringai quality evaluationdata annotationcontent moderationdebuggingpersonalizationnatural language processingside-by-side evaluationdomain-specific languages

Application Tip

When you complete the assessment, cite explicit turn numbers in your rationales and quote the exact signal that broke grounding or overnarrating. Showing that you can reference conversation details with precision is the strongest way to meet the written communication bar for this role.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

10h agoRemotecontract
Hot

AI Quality Analyst - English

You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Competitive salary
GeminiPersonalizationPrompt Engineering+15 more
Turing

Turing

21d agoRemotecontract

AI Quality Analyst (Personalization) - Turkish

In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

15–15/hr
GeminiGmailSearch+14 more
Turing

Turing

16d agoRemotecontract

AI Quality Analyst (Personalization) - Polish

Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

20–20/hr
GeminiGmailSearch+12 more
Turing

Turing

22d agoRemotecontract

AI Quality Analyst (Personalization) - Russian

You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

15–15/hr
RussianGeminiGmail+12 more
Turing

Turing

12d agoRemotecontract
Hot

AI Quality Analyst (Personalization) -Japanese

Gemini's new personalization feature depends on human judgment to know if it works. You will create and test multi-turn prompts in Japanese using your own Google data from Gmail, Search, and YouTube. Each task involves comparing two model outputs and deciding which one handles grounding, integration, and helpfulness better. The role runs for 3 months as a contractor position, with a minimum daily commitment of 4 hours and a required overlap with Pacific Time.

15–15/hr
GeminiJapanesePrompt Engineering+11 more
Turing

Turing

11d agoRemotecontract

AI Quality Analyst (Personalization) - Vietnamese

This role puts you inside the evaluation loop for a new personalization feature in Gemini. You will design multi-turn prompts (1-5 turns) that draw on your own Google account data, including past chats, Gmail, Search, and YouTube history. Your core job is to judge whether the model's personalized responses are relevant, grounded in real evidence, and natural, using dimensions like Grounding and Integration. Work is conducted in Vietnamese, comparing two model outputs side-by-side and writing rationales that cite exact conversation turns. The engagement lasts 3 months and pays $15 per hour, with a daily commitment that includes PST overlap.

15–15/hr
VietnameseGeminiGmail+8 more