
AI Quality Analyst (Personalization) - Portuguese
Overview
You will evaluate a new personalization feature for Gemini by testing how it draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own life and judge whether the model's responses are grounded, integrated, and genuinely helpful. Your daily routine involves comparing two model answers side-by-side and writing rationales that reference specific turns. The role demands Portuguese fluency, a willingness to use your personal Google account, and full-time availability aligned to global coverage.
What You'll Do8
- 1Create multi-turn conversation prompts ranging from 1 to 5 turns that force Gemini to use your personal data and past interactions.
- 2Score each model answer against your original intent and decide whether the personalization worked as intended.
- 3Review responses for Grounding by confirming that claims about you are supported by evidence rather than guesses or hallucinations.
- 4Assess Integration quality to see whether personal details appear in the response without sounding forced or overnarrated.
- 5Compare two model responses side-by-side and rank them on helpfulness, ease of use, and overall enjoyment.
- 6Write concise rationales for your rankings that cite specific turn numbers and point to exact moments in the conversation.
- 7Pull and inspect Debug Info from the model to verify that chat summaries and data sources were used as expected.
- 8Delete evaluation conversations after each task so they do not contaminate your future chat history.
Requirements11
- 1Native or near-native reading and writing ability in Portuguese; the project runs in Portuguese as the focus language.
- 2Willingness to use your primary personal Google account, not a test account, and enable personal data sources for realistic evaluation.
- 3Full-time availability in your local time zone, with the ability to supply at least 4 hours of overlap with PST per day.
- 4Strong analytical skill for judging vague or ambiguous AI outputs, with a focus on personalization quality.
- 5Experience designing creative, multi-turn prompts grounded in personal context.
- 6Understanding of personalization concepts and the ability to flag incorrect personalization, poor inferences, or forced connections.
- 7Sharp attention to detail when comparing side-by-side model responses for naturalness and overnarrating.
- 8Excellent written communication, including the ability to write clear rationales that reference exact turn numbers.
- 9Ability to work independently in a remote role and maintain strict data hygiene.
- 10Bachelor's degree or equivalent experience in fields such as linguistics, journalism, computer science, policy, law, or ethics.
- 11Prior experience in data annotation, AI quality evaluation, or content moderation gives you an edge.
Who Should Apply
The ideal candidate for this contract position is a Portuguese-speaking evaluator who enjoys dissecting how AI uses personal context and can write precise, evidence-based rationales. Someone who is comfortable connecting their own Gmail, Search, and YouTube history to a live product evaluation will fit well here. This role is less suitable for people who are not willing to use their primary personal Google account or who cannot commit to at least 4 hours of overlap with PST. Candidates often get rejected when their assessment is incomplete, when their rationales lack turn-specific evidence, or when they miss the 24-hour deadline for the evaluation. Strong prior experience in data annotation or AI quality evaluation tends to separate finalists from the rest.
Salary Insight
The offered rate for this project is $15 per hour, with a contractor engagement lasting 3 months.
Location
Required Skills
Application Tip
When you complete the assessment, write rationales with explicit turn numbers and describe where the model grounded its response in your stated personal context. Submit the assessment within 24 hours and be prepared to show that you understand the Side-by-Side (SxS) ranking process and the Debug Info verification step.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst - Portuguese (Portugal)
Turing hires an AI Quality Analyst to evaluate a new Gemini personalization feature that pulls from your past conversations, Gmail, Google Search, and YouTube activity. You will write multi-turn prompts in Portuguese from your own experiences and judge how well the model responds. Each review weighs personalization quality across Grounding, Integration, and Helpfulness. The contractor role runs for 3 months, pays $15 per hour, and asks for 4 hours of daily overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Spanish
Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

Turing
VerifiedAI Quality Analyst (Personalization) - Polish
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

Turing
VerifiedAI Quality Analyst - English
You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Turing
VerifiedAI Quality Analyst (Personalization) - German
An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

Turing
VerifiedAI Quality Analyst (Personalization) - Arabic
You will evaluate a new personalization feature in Gemini. You will design prompts that draw on your own life and judge how well the model uses data from past conversations, Gmail, Google Search, and YouTube to tailor responses. Each review weighs Grounding, Integration, and helpfulness. The project focuses on Arabic, so strong reading and writing skills in that language are required. Work is remote and requires a 4-hour overlap with Pacific Time each day.

