Turing
TuringVerified listing
Remote

AI Quality Analyst (Personalization) - Russian

15–15/hr
Remote
Posted August 13, 2026
contract

Overview

You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

What You'll Do8

  • 1Design and run multi-turn conversational prompts spanning 1-5 turns that require the model to draw on your personal context and experiences.
  • 2Evaluate each response against the intent of your starting prompt and decide if the personalization was applied in a way that matches that intent.
  • 3Check responses for grounding failures, such as unsupported claims, flawed inferences, or hallucinated details about you.
  • 4Assess integration quality to catch robotic overnarrating and verify that personal data appears in a natural way.
  • 5Compare two model responses side-by-side (SxS) and stack-rank them by helpfulness, ease of use, and enjoyment.
  • 6Write clear, defensible rationales that reference exact turn numbers and cite where issues or strengths appeared.
  • 7Extract and verify Debug Info to confirm the model used chat summaries and connected data sources.
  • 8Delete evaluation conversations after use to keep your chat history clean and avoid contaminating future assessments.

Requirements10

  • 1Read and write Russian with high competence, since Russian is the focus language for all evaluations.
  • 2Use your primary personal Google account and enable Gmail, Google Search, YouTube, and Gemini history for genuine testing.
  • 3Commit to at least 4 hours per day (up to 40 hours per week) with a 4-hour overlap with PST.
  • 4Show strong analytical thinking when judging nuanced, ambiguous, and flawed AI outputs.
  • 5Prove creative prompt engineering ability by building multi-turn prompts from personal context.
  • 6Understand personalization concepts like incorrect personalization, poor inferences, and forced connections.
  • 7Spot subtle differences in naturalness and overnarrating when reviewing Side-by-Side responses.
  • 8Write clear, concise rationales with explicit turn-number references for every ranking.
  • 9Hold a BS/BA degree or equivalent experience in a relevant field such as linguistics, journalism, computer science, policy, law, or ethics.
  • 10Bring prior experience in data annotation, AI quality evaluation, content moderation, or a similar role.

Who Should Apply

The right candidate is a detail-oriented Russian speaker with real curiosity about how AI personalizes information and a comfort connecting personal Google data for evaluation work. This role suits someone who can independently draft varied prompts and then write rigorous, turn-specific rationales for their rankings. It is less suitable for people who are hesitant to connect personal accounts, cannot keep a consistent PST overlap schedule, or prefer fixed, repetitive tasks. Common rejection reasons include missing the 24-hour assessment window or writing vague rationales that do not name specific turn numbers and fail to distinguish grounding issues from integration issues.

Salary Insight

Contractor rate is $15 per hour for a 3-month engagement.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

russiangeminigmailgoogle searchyoutubepersonalizationprompt engineeringmulti-turn promptsgroundingintegrationside-by-side evaluationdebug infodata annotationcontent moderationdomain-specific languages

Application Tip

When you take the assessment, write each rationale as if you are walking a reviewer through the conversation: name the turn number, state what the model claimed, and explain whether it was grounded in your input or a forced inference. This mirrors the day-to-day deliverable and proves you can handle ambiguity.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

16d agoRemotecontract

AI Quality Analyst (Personalization) - Polish

Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

20–20/hr
GeminiGmailSearch+12 more
Turing

Turing

24d agoRemotecontract

AI Quality Analyst - English

This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Competitive salary
GeminiGmailSearch+10 more
Turing

Turing

21d agoRemotecontract

AI Quality Analyst (Personalization) - Turkish

In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

15–15/hr
GeminiGmailSearch+14 more
Turing

Turing

12h agoRemotecontract
Hot

AI Quality Analyst - English

You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Competitive salary
GeminiPersonalizationPrompt Engineering+15 more
Turing

Turing

15d agoRemotecontract

AI Quality Analyst (Personalization) - Korean

This role evaluates a new personalization feature for Gemini that draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own personal context and then judge how well the model grounds responses in real evidence rather than inferences. The work includes side-by-side (SxS) ranking of two model outputs, writing structured rationales that cite specific turn numbers, and checking Debug Info to verify chat summaries and data sources. Korean is the focus language, so reading and writing fluency in Korean is required. The position is remote, full-time with at least 4 hours of daily overlap with PST, and runs as a 3-month contractor engagement.

15–15/hr
GeminiKoreanGmail+11 more
Turing

Turing

20d agoRemotecontract

AI Quality Analyst (Personalization) - Spanish

Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

15–15/hr
SpanishGeminiGmail+12 more