
AI Quality Analyst (Personalization) - Vietnamese
Overview
This role puts you inside the evaluation loop for a new personalization feature in Gemini. You will design multi-turn prompts (1-5 turns) that draw on your own Google account data, including past chats, Gmail, Search, and YouTube history. Your core job is to judge whether the model's personalized responses are relevant, grounded in real evidence, and natural, using dimensions like Grounding and Integration. Work is conducted in Vietnamese, comparing two model outputs side-by-side and writing rationales that cite exact conversation turns. The engagement lasts 3 months and pays $15 per hour, with a daily commitment that includes PST overlap.
What You'll Do8
- 1Design and run multi-turn chat prompts that require the model to use personal data from your Google account, with 1-5 turns per evaluation.
- 2Decide whether each generated response matches the intent behind your starting prompt and whether personalization was applied correctly.
- 3Inspect responses for Grounding failures, such as claims about you that are not supported by evidence or that amount to hallucinations.
- 4Evaluate Integration quality, watching for personal details that sound natural rather than forced or overnarrated.
- 5Compare two model responses side-by-side and rank them on overall helpfulness, ease of use, and enjoyment.
- 6Write concise, well-structured rationales for each side-by-side ranking, referencing the exact turns where strengths or issues appear.
- 7Open the model's Debug Info to verify that chat summaries and the correct personal data sources were used to generate the response.
- 8Delete every evaluation conversation after you finish to prevent it from contaminating future chat history and degrading future personalized results.
Requirements11
- 1Near-native proficiency in Vietnamese, with strong reading and writing skills, since Vietnamese is the focus language.
- 2Willingness to log in with your primary personal Google account and enable personal data sources such as Gmail, Search, and YouTube.
- 3Full-time availability in your local time zone, including at least 4 hours of daily overlap with Pacific Time (up to 40 hours per week).
- 4Analytical thinking strong enough to evaluate nuanced, ambiguous AI responses and judge personalization quality.
- 5Experience designing creative, multi-turn starting prompts that draw on personal context.
- 6Familiarity with personalization concepts, including the ability to detect incorrect personalization, poor inferences, and forced connections.
- 7Meticulous attention to detail when reviewing Side-by-Side (SxS) model responses, catching subtle differences in naturalness and overnarrating.
- 8Excellent written communication and collaboration skills, with the ability to write clear rationales that reference specific turn numbers and provide constructive feedback and detailed annotations.
- 9Self-motivated and able to work independently, with a desktop or laptop and a reliable internet connection.
- 10A BS/BA degree or equivalent experience in a relevant analytical field such as linguistics, journalism, computer science, law, policy, or ethics.
- 11Preference for prior experience in data annotation, AI quality evaluation, content moderation, or a related role.
Who Should Apply
The right candidate for this role is fluent in Vietnamese, comfortable giving the model access to real personal Google data, and curious about how AI personalization should behave. They write rationales that are precise and evidence-based, pointing to specific turns in the conversation rather than vague impressions. This role is less suitable for people who want simple, repeatable annotation tasks or who are not comfortable using a primary personal account. Candidates often get rejected when they fail to complete the timed assessment within 24 hours, or when their rationales do not reference specific turns and show a clear grasp of grounding and integration.
Salary Insight
Contract rate of $15 per hour for a 3-month engagement.
Location
Required Skills
Application Tip
During the assessment, write rationales that cite exact turn numbers and name the specific personalization failure (for example, a forced connection to a Gmail thread). Show that you can separate natural integration from overnarrating, and that you understand how Grounding works.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst - English
This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Thai
Turing is assembling a global team of Thai-speaking quality analysts to evaluate a new Gemini personalization feature. Analysts design 1-5 turn prompts from their own experience and test how well the model uses data from Gmail, Google Search, and YouTube. Each day involves stack-ranking two model responses side-by-side, scoring for grounding, integration, and helpfulness, and writing rationales that reference specific turn numbers. The role requires using a primary personal Google account, full-time availability in your local time zone, and a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Korean
This role evaluates a new personalization feature for Gemini that draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own personal context and then judge how well the model grounds responses in real evidence rather than inferences. The work includes side-by-side (SxS) ranking of two model outputs, writing structured rationales that cite specific turn numbers, and checking Debug Info to verify chat summaries and data sources. Korean is the focus language, so reading and writing fluency in Korean is required. The position is remote, full-time with at least 4 hours of daily overlap with PST, and runs as a 3-month contractor engagement.

Turing
VerifiedAI Quality Analyst - English
You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Turing
VerifiedAI Quality Analyst (Personalization) - Indonesian
This role sits on a global team that checks how well Gemini tailors responses using a user's own history, including past chats, Gmail, Google Search, and YouTube activity. As an AI Quality Analyst, you craft personal, multi-turn prompts in Indonesian and then judge whether the model grounded its replies in your stated context. You also rank two model responses side by side to decide which is more helpful and natural. The work is remote, but you must keep a set schedule that overlaps with PST hours.

Turing
VerifiedAI Quality Analyst (Personalization) -Japanese
Gemini's new personalization feature depends on human judgment to know if it works. You will create and test multi-turn prompts in Japanese using your own Google data from Gmail, Search, and YouTube. Each task involves comparing two model outputs and deciding which one handles grounding, integration, and helpfulness better. The role runs for 3 months as a contractor position, with a minimum daily commitment of 4 hours and a required overlap with Pacific Time.

