
AI Quality Analyst (Personalization) - Spanish
Overview
Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.
What You'll Do8
- 1Create multi-turn conversation prompts (usually 1-5 turns) that force the model to use your personal context and life experiences.
- 2Judge whether each model response matches your original intent and whether personalization was applied appropriately.
- 3Check responses for Grounding errors by verifying that claims about you are backed by evidence rather than guesses or hallucinations.
- 4Review Integration quality to confirm personal data appears naturally in the reply and does not sound robotic or overnarrated.
- 5Compare two model responses side-by-side and stack-rank them based on helpfulness, ease of use, and enjoyment.
- 6Write concise, defensible rationales for your rankings, citing specific turn numbers where issues or strengths appear.
- 7Pull and verify Debug Info from the model to confirm chat summaries and personal data sources were used correctly.
- 8Delete evaluation conversations after each task to keep personal data from contaminating your future chat history.
Requirements10
- 1Native-level reading and writing in Spanish, since this project evaluates Spanish-language responses.
- 2Willingness to use your personal Google account and turn on personal data sources for authentic assessments.
- 3Ability to commit at least 4 hours per day, up to 40 hours per week, with 4 hours of overlap with PST.
- 4Strong analytical thinking for judging nuanced, ambiguous AI outputs and personalization quality.
- 5Experience creating creative, multi-turn prompts based on your own personal context.
- 6Understanding of personalization concepts, including spotting incorrect personalization, poor inferences, and forced connections.
- 7Sharp attention to detail for comparing side-by-side model responses and noticing subtle differences in naturalness and overnarrating.
- 8Excellent written communication, with the ability to write structured rationales that reference specific turn numbers.
- 9BS/BA degree or equivalent experience in fields like linguistics, journalism, computer science, or a related analytical discipline.
- 10Prior experience in data annotation, AI quality evaluation, or content moderation is strongly preferred.
Who Should Apply
The ideal candidate is fluent in Spanish, comfortable sharing their personal Google data, and genuinely curious about how AI uses real user context. The role suits someone who can work independently, write clear evaluation rationales, and commit to a daily 4-hour overlap with PST. People who are uncomfortable with privacy trade-offs or expect a structured in-office environment will find this role less suitable. Candidates often get rejected when their Spanish writing falls short of the project's bar, or when they miss the 24-hour assessment deadline. Another common miss is underestimating the PST overlap requirement and then being unavailable during the required hours.
Salary Insight
Pay is $15 per hour for a 3-month contractor engagement.
Location
Required Skills
Application Tip
Before the assessment, practice writing side-by-side comparisons that reference specific turn numbers and call out Grounding errors or overnarrating. Also be ready to confirm your 4-hour PST overlap and explain how your personal Google account setup will support authentic testing.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst (Personalization) - Portuguese
You will evaluate a new personalization feature for Gemini by testing how it draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own life and judge whether the model's responses are grounded, integrated, and genuinely helpful. Your daily routine involves comparing two model answers side-by-side and writing rationales that reference specific turns. The role demands Portuguese fluency, a willingness to use your personal Google account, and full-time availability aligned to global coverage.

Turing
VerifiedAI Quality Analyst (Personalization) - Polish
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

Turing
VerifiedAI Quality Analyst (Personalization) - Turkish
In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

Turing
VerifiedAI Quality Analyst - English
You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Turing
VerifiedAI Quality Analyst (Personalization) - German
An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

Turing
VerifiedAI Quality Analyst (Personalization) - Arabic
You will evaluate a new personalization feature in Gemini. You will design prompts that draw on your own life and judge how well the model uses data from past conversations, Gmail, Google Search, and YouTube to tailor responses. Each review weighs Grounding, Integration, and helpfulness. The project focuses on Arabic, so strong reading and writing skills in that language are required. Work is remote and requires a 4-hour overlap with Pacific Time each day.

