
AI Quality Analyst (Personalization) - Korean
Overview
This role evaluates a new personalization feature for Gemini that draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own personal context and then judge how well the model grounds responses in real evidence rather than inferences. The work includes side-by-side (SxS) ranking of two model outputs, writing structured rationales that cite specific turn numbers, and checking Debug Info to verify chat summaries and data sources. Korean is the focus language, so reading and writing fluency in Korean is required. The position is remote, full-time with at least 4 hours of daily overlap with PST, and runs as a 3-month contractor engagement.
What You'll Do8
- 1Craft and run multi-turn conversational prompts (1-5 turns) that force the model to use your personal information and experiences.
- 2Judge whether each response matches the intent set in your starting prompt and whether the personalization fits that intent.
- 3Flag grounding failures where the model makes claims about you that lack evidence or rely on flawed inferences.
- 4Check how well the response weaves personal data into the conversation without robotic overnarrating.
- 5Compare two model responses side-by-side (SxS) and rank which one is more helpful, easy to use, and enjoyable.
- 6Write clear rationales for your rankings, citing specific turn numbers where issues or strengths appear.
- 7Inspect Debug Info to verify that chat summaries and data sources were used as expected.
- 8Delete evaluation conversations after each session to keep your personal chat history clean.
Requirements10
- 1Native-level reading and writing in Korean, since Korean is the project's focus language.
- 2Willingness to use your primary personal Google account and turn on personal data sources for realistic evaluation.
- 3Availability for full-time work in your local time zone, with at least 4 hours of overlap with PST per day.
- 4Strong analytical skill for evaluating ambiguous AI responses and judging personalization quality.
- 5Experience creating creative, multi-turn prompts based on personal context to test model capabilities.
- 6Solid understanding of personalization concepts, including spotting incorrect personalization and forced connections.
- 7Sharp attention to detail when comparing side-by-side responses for subtle differences in naturalness.
- 8Excellent written communication for producing concise, structured rationales with turn-number references.
- 9BS/BA degree or equivalent experience in a relevant analytical field such as linguistics, journalism, or computer science.
- 10Prior experience in data annotation, AI quality evaluation, or content moderation is preferred.
Who Should Apply
The ideal candidate combines native-level Korean fluency with a strong habit of analytical writing and prompt design. This person should feel comfortable using their personal Google data for work and can commit to a fixed full-time schedule with PST overlap. The role is less suitable for anyone unwilling to use a personal account, since test accounts won't work for genuine personalization checks, and it will frustrate people who prefer a standard 9-to-5 local schedule rather than rotating global shifts. Candidates often get rejected when their written rationales stay vague, lacking specific turn references, or when their sample prompts show they don't understand how to trigger grounded personal responses. Attention to detail in spotting subtle overnarrating also makes or breaks an application.
Salary Insight
Offered rate is $15 per hour.
Location
Required Skills
Application Tip
In your assessment, write rationales that cite specific turn numbers and call out grounding failures. Show you can tell a forced personal connection from a natural one, and use concrete examples from your own prompt designs to demonstrate Korean fluency and analytical precision.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst - English
This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Turkish
In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

Turing
VerifiedAI Quality Analyst (Personalization) - Polish
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

Turing
VerifiedAI Quality Analyst (Personalization) -Japanese
Gemini's new personalization feature depends on human judgment to know if it works. You will create and test multi-turn prompts in Japanese using your own Google data from Gmail, Search, and YouTube. Each task involves comparing two model outputs and deciding which one handles grounding, integration, and helpfulness better. The role runs for 3 months as a contractor position, with a minimum daily commitment of 4 hours and a required overlap with Pacific Time.

Turing
VerifiedAI Quality Analyst (Personalization) - Vietnamese
This role puts you inside the evaluation loop for a new personalization feature in Gemini. You will design multi-turn prompts (1-5 turns) that draw on your own Google account data, including past chats, Gmail, Search, and YouTube history. Your core job is to judge whether the model's personalized responses are relevant, grounded in real evidence, and natural, using dimensions like Grounding and Integration. Work is conducted in Vietnamese, comparing two model outputs side-by-side and writing rationales that cite exact conversation turns. The engagement lasts 3 months and pays $15 per hour, with a daily commitment that includes PST overlap.

Turing
VerifiedAI Quality Analyst (Personalization) - Russian
You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

