Turing
TuringVerified listing
Remote

AI Quality Analyst (Personalization) - Hindi

15–15/hr
Remote
Posted August 25, 2026
contract

Overview

Turing runs a global evaluation team for Gemini's personalization feature, and this contract role focuses on Hindi responses. You will create multi-turn prompts from your own Google history, including Gmail, Search, YouTube, and past conversations, then rate how well the model uses that context. Scoring centers on Grounding, Integration, and Helpfulness, with side-by-side response comparisons and written justifications. The role requires at least 4 hours per day with PST overlap, up to 40 hours per week, for a 3-month contract.

What You'll Do8

  • 1Create and run multi-turn chat prompts (1 to 5 turns) that force the model to draw on your personal Google data and experiences.
  • 2Rate each response against your original intent and decide whether the personalization was applied as intended.
  • 3Check for Grounding errors by verifying that claims about you are supported by real evidence rather than guesses or hallucinations.
  • 4Judge Integration quality by checking if personal data appears in the answer without forced or over-explanatory language.
  • 5Compare two model responses side-by-side and rank them based on overall helpfulness, usability, and user enjoyment.
  • 6Write concise, structured rationales for each comparison, citing specific turn numbers and exact spots where the model succeeded or failed.
  • 7Inspect the model's Debug Info to confirm it used the right chat summaries and data sources.
  • 8Delete every evaluation conversation after rating to keep personal history clean and prevent contamination of future sessions.

Requirements11

  • 1Read and write Hindi with enough proficiency to evaluate nuanced AI responses in that language.
  • 2Use your primary personal Google account and enable personal data sources like Gmail, Search, and YouTube for evaluation.
  • 3Maintain full-time availability in your home time zone with at least 4 hours of overlap with PST.
  • 4Show strong analytical thinking when judging ambiguous or contradictory AI responses.
  • 5Create original, multi-turn prompts based on personal context to stress-test the personalization feature.
  • 6Spot personalization failures, such as incorrect inferences or forced connections between user data and the answer.
  • 7Compare side-by-side model responses and catch small differences in naturalness and overnarrating.
  • 8Write clear rationales that reference specific turn numbers and evidence from the conversation.
  • 9Work on your own and stay organized in a remote setting.
  • 10Hold a BS/BA degree or equivalent experience in a field like Linguistics, Journalism, Computer Science, or a related analytical discipline.
  • 11Bring prior experience in data annotation, AI quality evaluation, or content moderation (preferred).

Who Should Apply

The ideal candidate is a fluent Hindi speaker who enjoys dissecting how AI uses personal data to shape answers. This person should be comfortable turning their own Gmail, Search, YouTube, and chat history into test prompts and then judging the results with written evidence. The role suits someone who can commit 4 hours daily with PST overlap for a 3-month contract and who treats data privacy with care. People who dislike sharing personal Google data or who prefer fixed, repetitive tasks will find this role less suitable. Candidates often get rejected when they fail the written assessment within 24 hours or when their rationales lack specific turn references.

Salary Insight

The project pays $15 per hour and runs for 3 months as a contractor engagement.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

hindigeminigoogle searchgmailyoutubeprompt engineeringmulti-turn conversationsai quality evaluationside-by-side evaluationgroundingintegrationhelpfulnessdata annotationpersonalizationdomain-specific languages

Application Tip

Before you apply, write a sample SxS evaluation for two AI responses and reference the exact turn numbers where each response succeeded or failed. The screening assessment will test whether you can justify rankings with concrete evidence.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

24d agoRemotecontract

AI Quality Analyst - English

This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Competitive salary
GeminiGmailSearch+10 more
Turing

Turing

12d agoRemotecontract

AI Quality Analyst (Personalization) - Vietnamese

This role puts you inside the evaluation loop for a new personalization feature in Gemini. You will design multi-turn prompts (1-5 turns) that draw on your own Google account data, including past chats, Gmail, Search, and YouTube history. Your core job is to judge whether the model's personalized responses are relevant, grounded in real evidence, and natural, using dimensions like Grounding and Integration. Work is conducted in Vietnamese, comparing two model outputs side-by-side and writing rationales that cite exact conversation turns. The engagement lasts 3 months and pays $15 per hour, with a daily commitment that includes PST overlap.

15–15/hr
VietnameseGeminiGmail+8 more
Turing

Turing

21d agoRemotecontract

AI Quality Analyst (Personalization) - Turkish

In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

15–15/hr
GeminiGmailSearch+14 more
Turing

Turing

12d agoRemotecontract
Hot

AI Quality Analyst (Personalization) -Japanese

Gemini's new personalization feature depends on human judgment to know if it works. You will create and test multi-turn prompts in Japanese using your own Google data from Gmail, Search, and YouTube. Each task involves comparing two model outputs and deciding which one handles grounding, integration, and helpfulness better. The role runs for 3 months as a contractor position, with a minimum daily commitment of 4 hours and a required overlap with Pacific Time.

15–15/hr
GeminiJapanesePrompt Engineering+11 more
Turing

Turing

23d agoRemotecontract

AI Quality Analyst (Personalization) - Indonesian

This role sits on a global team that checks how well Gemini tailors responses using a user's own history, including past chats, Gmail, Google Search, and YouTube activity. As an AI Quality Analyst, you craft personal, multi-turn prompts in Indonesian and then judge whether the model grounded its replies in your stated context. You also rank two model responses side by side to decide which is more helpful and natural. The work is remote, but you must keep a set schedule that overlaps with PST hours.

15–15/hr
IndonesianGeminiGmail+13 more
Turing

Turing

20d agoRemotecontract

AI Quality Analyst (Personalization) - Spanish

Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

15–15/hr
SpanishGeminiGmail+12 more