Turing
TuringVerified listing
Remote

AI Quality Analyst (Personalization) - Portuguese

15–15/hr
Remote
Posted August 11, 2026
contract

Overview

You will evaluate a new personalization feature for Gemini by testing how it draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own life and judge whether the model's responses are grounded, integrated, and genuinely helpful. Your daily routine involves comparing two model answers side-by-side and writing rationales that reference specific turns. The role demands Portuguese fluency, a willingness to use your personal Google account, and full-time availability aligned to global coverage.

What You'll Do8

  • 1Create multi-turn conversation prompts ranging from 1 to 5 turns that force Gemini to use your personal data and past interactions.
  • 2Score each model answer against your original intent and decide whether the personalization worked as intended.
  • 3Review responses for Grounding by confirming that claims about you are supported by evidence rather than guesses or hallucinations.
  • 4Assess Integration quality to see whether personal details appear in the response without sounding forced or overnarrated.
  • 5Compare two model responses side-by-side and rank them on helpfulness, ease of use, and overall enjoyment.
  • 6Write concise rationales for your rankings that cite specific turn numbers and point to exact moments in the conversation.
  • 7Pull and inspect Debug Info from the model to verify that chat summaries and data sources were used as expected.
  • 8Delete evaluation conversations after each task so they do not contaminate your future chat history.

Requirements11

  • 1Native or near-native reading and writing ability in Portuguese; the project runs in Portuguese as the focus language.
  • 2Willingness to use your primary personal Google account, not a test account, and enable personal data sources for realistic evaluation.
  • 3Full-time availability in your local time zone, with the ability to supply at least 4 hours of overlap with PST per day.
  • 4Strong analytical skill for judging vague or ambiguous AI outputs, with a focus on personalization quality.
  • 5Experience designing creative, multi-turn prompts grounded in personal context.
  • 6Understanding of personalization concepts and the ability to flag incorrect personalization, poor inferences, or forced connections.
  • 7Sharp attention to detail when comparing side-by-side model responses for naturalness and overnarrating.
  • 8Excellent written communication, including the ability to write clear rationales that reference exact turn numbers.
  • 9Ability to work independently in a remote role and maintain strict data hygiene.
  • 10Bachelor's degree or equivalent experience in fields such as linguistics, journalism, computer science, policy, law, or ethics.
  • 11Prior experience in data annotation, AI quality evaluation, or content moderation gives you an edge.

Who Should Apply

The ideal candidate for this contract position is a Portuguese-speaking evaluator who enjoys dissecting how AI uses personal context and can write precise, evidence-based rationales. Someone who is comfortable connecting their own Gmail, Search, and YouTube history to a live product evaluation will fit well here. This role is less suitable for people who are not willing to use their primary personal Google account or who cannot commit to at least 4 hours of overlap with PST. Candidates often get rejected when their assessment is incomplete, when their rationales lack turn-specific evidence, or when they miss the 24-hour deadline for the evaluation. Strong prior experience in data annotation or AI quality evaluation tends to separate finalists from the rest.

Salary Insight

The offered rate for this project is $15 per hour, with a contractor engagement lasting 3 months.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

geminigmailgoogle searchyoutubeportugueseprompt engineeringside-by-side evaluationdata annotationpersonalizationgroundingintegrationmulti-turn conversationdebug infodomain-specific languages

Application Tip

When you complete the assessment, write rationales with explicit turn numbers and describe where the model grounded its response in your stated personal context. Submit the assessment within 24 hours and be prepared to show that you understand the Side-by-Side (SxS) ranking process and the Debug Info verification step.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

12d agoRemotecontract

AI Quality Analyst - Portuguese (Portugal)

Turing hires an AI Quality Analyst to evaluate a new Gemini personalization feature that pulls from your past conversations, Gmail, Google Search, and YouTube activity. You will write multi-turn prompts in Portuguese from your own experiences and judge how well the model responds. Each review weighs personalization quality across Grounding, Integration, and Helpfulness. The contractor role runs for 3 months, pays $15 per hour, and asks for 4 hours of daily overlap with PST.

15–15/hr
PortugueseGeminiGmail+13 more
Turing

Turing

20d agoRemotecontract

AI Quality Analyst (Personalization) - Spanish

Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

15–15/hr
SpanishGeminiGmail+12 more
Turing

Turing

16d agoRemotecontract

AI Quality Analyst (Personalization) - Polish

Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

20–20/hr
GeminiGmailSearch+12 more
Turing

Turing

13h agoRemotecontract
Hot

AI Quality Analyst - English

You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Competitive salary
GeminiPersonalizationPrompt Engineering+15 more
Turing

Turing

17d agoRemotecontract

AI Quality Analyst (Personalization) - German

An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

15–15/hr
GeminiGermanGmail+10 more
Turing

Turing

23d agoRemotecontract

AI Quality Analyst (Personalization) - Arabic

You will evaluate a new personalization feature in Gemini. You will design prompts that draw on your own life and judge how well the model uses data from past conversations, Gmail, Google Search, and YouTube to tailor responses. Each review weighs Grounding, Integration, and helpfulness. The project focuses on Arabic, so strong reading and writing skills in that language are required. Work is remote and requires a 4-hour overlap with Pacific Time each day.

15–15/hr
ArabicGeminiGmail+10 more