
AI Quality Analyst (Personalization) -Japanese
Overview
Gemini's new personalization feature depends on human judgment to know if it works. You will create and test multi-turn prompts in Japanese using your own Google data from Gmail, Search, and YouTube. Each task involves comparing two model outputs and deciding which one handles grounding, integration, and helpfulness better. The role runs for 3 months as a contractor position, with a minimum daily commitment of 4 hours and a required overlap with Pacific Time.
What You'll Do8
- 1Craft and run multi-turn conversation prompts (1-5 turns) that force the model to draw on your personal Google activity and experiences.
- 2Judge whether each response uses your intent and personal history in a relevant way, flagging cases where personalization is missing or misapplied.
- 3Inspect outputs for grounding failures, such as claims about you that rely on weak inferences or hallucinations.
- 4Rate how well personal data is woven into the response, watching for robotic over-narration or forced connections.
- 5Compare two model outputs side-by-side and rank them by helpfulness, usability, and overall feel.
- 6Write concise rationales that cite specific turn numbers to justify each ranking.
- 7Pull and verify debug data to confirm the model relied on the right chat summaries and enabled data sources.
- 8Delete evaluation chats after each session to keep your personal history clean and uncontaminated.
Requirements11
- 1Fluency in written and spoken Japanese, with the ability to craft natural prompts and rationales in the language.
- 2Willingness to use your personal Google account as the test environment and enable personalization data sources like Gmail and YouTube history.
- 3Availability to work at least 4 hours per day, up to 20 hours per week, with 4 hours overlapping Pacific Time.
- 4Strong analytical thinking to evaluate ambiguous AI outputs and spot subtle flaws in personalization.
- 5Experience designing creative multi-turn prompts that test a model's use of personal context.
- 6Understanding of personalization quality dimensions, including grounding, integration, and helpfulness.
- 7Exceptional attention to detail when reviewing side-by-side model responses and noticing differences in naturalness and over-narration.
- 8Excellent written communication, able to produce clear, structured rationales with specific turn references.
- 9Ability to work independently in a remote setting with a reliable desktop or laptop and internet connection.
- 10BS/BA degree or equivalent experience in a relevant analytical field such as linguistics, computer science, journalism, or policy.
- 11Prior experience in data annotation, AI evaluation, or content moderation is preferred.
Who Should Apply
The ideal candidate is fluent in Japanese, comfortable using their personal Google account for evaluation work, and able to write clear, structured rationales that cite specific turns. This person has prior experience judging AI output or annotating data and can spot subtle differences between two model responses. The role is less suitable for anyone uncomfortable sharing their personal Google data or unable to maintain daily overlap with Pacific Time. Candidates often get rejected when their assessment rationales are vague, when they miss grounding errors, or when they fail to complete the evaluation within the 24-hour window.
Salary Insight
The offered rate for this project is $15 per hour.
Location
Required Skills
Application Tip
Showcase your Japanese writing skills by submitting a short sample rationale in Japanese, and state that you can commit to the required PST overlap and are open to enabling personal data sources on your own Google account.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst - English
This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Turkish
In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

Turing
VerifiedAI Quality Analyst (Personalization) - Polish
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

Turing
VerifiedAI Quality Analyst (Personalization) - Korean
This role evaluates a new personalization feature for Gemini that draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own personal context and then judge how well the model grounds responses in real evidence rather than inferences. The work includes side-by-side (SxS) ranking of two model outputs, writing structured rationales that cite specific turn numbers, and checking Debug Info to verify chat summaries and data sources. Korean is the focus language, so reading and writing fluency in Korean is required. The position is remote, full-time with at least 4 hours of daily overlap with PST, and runs as a 3-month contractor engagement.

Turing
VerifiedAI Quality Analyst (Personalization) - Russian
You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Vietnamese
This role puts you inside the evaluation loop for a new personalization feature in Gemini. You will design multi-turn prompts (1-5 turns) that draw on your own Google account data, including past chats, Gmail, Search, and YouTube history. Your core job is to judge whether the model's personalized responses are relevant, grounded in real evidence, and natural, using dimensions like Grounding and Integration. Work is conducted in Vietnamese, comparing two model outputs side-by-side and writing rationales that cite exact conversation turns. The engagement lasts 3 months and pays $15 per hour, with a daily commitment that includes PST overlap.

