
AI Quality Analyst (Gemini) - Chinese
Overview
You will evaluate a new personalization feature in Gemini that uses your past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The role blends creative prompt design with analytical review. You will write multi-turn prompts based on personal context and then score two model responses side-by-side on Grounding, Integration, and Helpfulness. The project centers on Chinese-language content, so strong reading and writing ability in Chinese is required.
What You'll Do6
- 1Create and run multi-turn conversation prompts of 1-5 turns that force the model to use your personal data and experiences.
- 2Check model responses for Grounding errors to see whether claims about you come from real evidence rather than faulty inferences or hallucinations.
- 3Compare two model outputs side-by-side and rank which one is more helpful, easier to use, and more enjoyable.
- 4Write clear, structured explanations for your rankings, citing specific conversation turns where the model did well or poorly.
- 5Pull and inspect the model's Debug Info to confirm it used the right chat summaries and data sources.
- 6Delete evaluation conversations after each task so they do not contaminate your future chat history.
Requirements8
- 1Native or near-native reading and writing ability in Chinese, since the project focuses on Chinese-language outputs.
- 2Experience designing creative, multi-turn prompts that rely on personal context to challenge the model.
- 3Strong understanding of personalization concepts, including identifying incorrect personalization, weak inferences, and forced connections.
- 4Meticulous attention to detail when comparing side-by-side (SxS) responses and catching subtle differences in naturalness and overnarrating.
- 5Excellent written communication skills for writing clear, concise rationales that reference specific turn numbers.
- 6Ability to provide constructive feedback and detailed annotations.
- 7A BS/BA degree or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or a related analytical field.
- 8Prior experience in data annotation, AI quality evaluation, content moderation, or a related role is preferred.
Who Should Apply
The ideal candidate is an analytical evaluator with strong Chinese reading and writing skills and a background in data annotation, AI quality evaluation, or content moderation. This person should enjoy designing multi-turn prompts from personal experience and defending their rankings with clear, turn-specific rationale. The role demands meticulous comparison of side-by-side outputs and disciplined deletion of evaluation chats after each session. Candidates who prefer open-ended research or who find repetitive comparative scoring tedious are a poor fit. Common rejection reasons include weak rationales that fail to reference specific turns, and prompts that lack meaningful personal context to stress-test the model.
Salary Insight
The offered rate for this project is $15 per hour.
Location
Required Skills
Application Tip
For the assessment, build each prompt from a concrete personal experience, like a recent trip or a specific email thread, and then compare the two model responses by referencing exact turn numbers and explaining your reasoning in terms of Grounding, Integration, and Helpfulness.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst - English
You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Turing
VerifiedAI Quality Analyst - English
This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Polish
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

Turing
VerifiedAI Quality Analyst (Personalization) - Turkish
In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

Turing
VerifiedAI Quality Analyst (Personalization) - Korean
This role evaluates a new personalization feature for Gemini that draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own personal context and then judge how well the model grounds responses in real evidence rather than inferences. The work includes side-by-side (SxS) ranking of two model outputs, writing structured rationales that cite specific turn numbers, and checking Debug Info to verify chat summaries and data sources. Korean is the focus language, so reading and writing fluency in Korean is required. The position is remote, full-time with at least 4 hours of daily overlap with PST, and runs as a 3-month contractor engagement.

Turing
VerifiedAI Quality Analyst (Personalization) - German
An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

