Turing
TuringVerified listing
RemoteHot

AI Quality Analyst - English

Remote
Posted September 4, 2026
contract

Overview

You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

What You'll Do8

  • 1Build and run multi-turn chat prompts that span 1 to 5 turns and pull from your personal data and experiences.
  • 2Review each model response against your original intent and flag any instance where the personalization did not hit the target.
  • 3Check for Grounding failures by verifying that statements about you trace back to real evidence instead of guesses or hallucinations.
  • 4Assess Integration quality to catch replies that bolt in personal details in a stiff or overwrought way.
  • 5Compare two versions of a response side by side and decide which one ranks higher for clarity, usefulness, and ease.
  • 6Write concise rationales for each ranking and cite the exact turns where the model succeeded or fell short.
  • 7Open the model's Debug Info to confirm it referenced the right chat summaries and data sources.
  • 8Delete each evaluation conversation after review so old tests cannot contaminate your future chat history.

Requirements12

  • 1Read and write English at a professional level; English is the project's focus language.
  • 2Use your personal Google account as the primary test account, not a separate testing profile.
  • 3Work a full-time schedule in your local time zone and support global 24-hour coverage needs.
  • 4Break down nuanced, ambiguous AI responses and decide whether the personalization fits the user's intent.
  • 5Design creative, multi-turn prompts using your own life context to stress-test Gemini.
  • 6Identify forced personalization, faulty inferences, and other signs that the model overreached.
  • 7Spot small differences between two responses in tone, naturalness, and tendency to overnarrate.
  • 8Explain your ranking decisions in clear written rationales that reference specific conversation turns.
  • 9Share constructive feedback and detailed annotations with the team.
  • 10Work alone in a remote setup with a dependable laptop or desktop and stable internet.
  • 11Hold a bachelor's degree or equivalent background in an analytical field like policy, law, linguistics, journalism, or computer science.
  • 12Prior experience in data annotation, AI quality evaluation, or content moderation is a meaningful plus.

Who Should Apply

The right candidate enjoys dissecting why an AI model chose specific personal details and can defend those judgments in writing. You should feel comfortable turning your real-world habits into conversation prompts and then ranking model replies with turn-level evidence. The role is less suitable for people who do not want their daily Google activity connected or who prefer tasks with one clear answer. Candidates often fail when their rationales lean on vague impressions instead of naming exact turns from the conversation. Bland starter prompts that do not exercise personalization also undercut an application, since the assessment rewards original scenarios.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

geminipersonalizationprompt engineeringmulti-turn promptsgroundingintegrationhelpfulnessside-by-side evaluationsxs evaluationmodel evaluationdata annotationai quality evaluationcontent moderationdebug infogmailgoogle searchyoutubedomain-specific languages

Application Tip

Write practice side-by-side evaluations before the assessment. Pick a personal conversation scenario, generate two possible responses, then rank them with a rationale that names specific turns and calls out any grounding or integration flaws.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

24d agoRemotecontract

AI Quality Analyst - English

This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Competitive salary
GeminiGmailSearch+10 more
Turing

Turing

16d agoRemotecontract

AI Quality Analyst (Personalization) - Polish

Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

20–20/hr
GeminiGmailSearch+12 more
Turing

Turing

23d agoRemotecontract

AI Quality Analyst (Gemini) - Chinese

You will evaluate a new personalization feature in Gemini that uses your past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The role blends creative prompt design with analytical review. You will write multi-turn prompts based on personal context and then score two model responses side-by-side on Grounding, Integration, and Helpfulness. The project centers on Chinese-language content, so strong reading and writing ability in Chinese is required.

15–15/hr
ChineseGeminiPrompt Engineering+8 more
Turing

Turing

20d agoRemotecontract

AI Quality Analyst (Personalization) - Spanish

Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

15–15/hr
SpanishGeminiGmail+12 more
Turing

Turing

21d agoRemotecontract

AI Quality Analyst (Personalization) - Turkish

In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

15–15/hr
GeminiGmailSearch+14 more
Turing

Turing

22d agoRemotecontract

AI Quality Analyst (Personalization) - Russian

You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

15–15/hr
RussianGeminiGmail+12 more