
AI Quality Analyst (Personalization) - Turkish
Overview
In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.
What You'll Do8
- 1Craft multi-turn conversation prompts (1-5 turns) that force the model to draw on your personal Google data, including past chats, Gmail, and search history.
- 2Judge whether each response uses the personal context you provided, and note any missed or misapplied personalization.
- 3Check responses for Grounding errors, such as unsupported claims about you, flawed inferences, or hallucinations.
- 4Assess how well personal data is integrated into the response, flagging robotic 'overnarration' or forced connections.
- 5Compare two model outputs side-by-side (SxS) and rank them on overall helpfulness, usability, and enjoyment.
- 6Write clear rationales for your rankings, referencing specific turn numbers and concrete examples.
- 7Inspect and verify the model's "Debug Info" to confirm that chat summaries and data sources were correctly retrieved.
- 8Delete evaluation conversations after each task to keep personal data from contaminating your future chat history.
Requirements11
- 1High proficiency in reading and writing Turkish (the project's focus language).
- 2Willingness to use your primary personal Google account and enable personal data sources like Gmail, Search, and YouTube history.
- 3Full-time availability: 30-40 hours per week, with at least 4 hours overlapping with Pacific Time (PST).
- 4Demonstrated ability to evaluate nuanced, ambiguous AI responses, with a focus on personalization quality.
- 5Experience designing creative, multi-turn prompts from personal context to stress-test AI capabilities.
- 6Strong grasp of personalization concepts, including spotting incorrect personalization, poor inferences, and forced connections.
- 7Meticulous attention to detail when comparing side-by-side model responses for naturalness and overnarrating.
- 8Excellent written communication, with the ability to write structured rationales that reference specific turn numbers.
- 9Bachelor's degree or equivalent experience in a relevant field such as linguistics, computer science, journalism, law, or policy.
- 10Prior experience in data annotation, AI quality evaluation, or content moderation is a significant plus.
- 11Reliable desktop or laptop and stable internet connection.
Who Should Apply
The ideal candidate has strong Turkish reading and writing skills, experience with AI quality evaluation or data annotation, and a proven ability to design creative multi-turn prompts. This role is less suitable for people who are uncomfortable using their personal Google account for work or who cannot commit to at least 30 hours per week with 4 hours of PST overlap. Many candidates get rejected because their Turkish proficiency does not meet the bar, or because they fail to provide clear, turn-specific rationales in the assessment. A strong application also depends on your willingness to enable personal data sources and your ability to spot subtle differences in model responses.
Salary Insight
The project pays $15 per hour. The engagement runs for 3 months and requires 30-40 hours per week.
Location
Required Skills
Application Tip
When preparing for the assessment, write a sample SxS comparison of two AI responses that references specific turn numbers and quotes exact phrases to demonstrate your attention to detail and analytical clarity.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst - English
This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Polish
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

Turing
VerifiedAI Quality Analyst (Personalization) - Arabic
You will evaluate a new personalization feature in Gemini. You will design prompts that draw on your own life and judge how well the model uses data from past conversations, Gmail, Google Search, and YouTube to tailor responses. Each review weighs Grounding, Integration, and helpfulness. The project focuses on Arabic, so strong reading and writing skills in that language are required. Work is remote and requires a 4-hour overlap with Pacific Time each day.

Turing
VerifiedAI Quality Analyst (Personalization) - Spanish
Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

Turing
VerifiedAI Quality Analyst (Personalization) -Japanese
Gemini's new personalization feature depends on human judgment to know if it works. You will create and test multi-turn prompts in Japanese using your own Google data from Gmail, Search, and YouTube. Each task involves comparing two model outputs and deciding which one handles grounding, integration, and helpfulness better. The role runs for 3 months as a contractor position, with a minimum daily commitment of 4 hours and a required overlap with Pacific Time.

Turing
VerifiedAI Quality Analyst (Personalization) - German
An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

