
AI Quality Analyst (Personalization) - Polish
Overview
Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.
What You'll Do8
- 1Build and run multi-turn conversation prompts (1-5 turns) that force the model to incorporate your personal data and experiences.
- 2Check whether each response respects your original intent and applies personalization in the right context.
- 3Inspect responses for Grounding failures, such as unsupported claims about you, flawed inferences, or hallucinations.
- 4Judge Integration by noticing whether personal details are woven in without awkwardness instead of sounding like robotic overexplaining.
- 5Rank two model outputs against each other side-by-side, deciding which is more helpful, easier to use, and more enjoyable.
- 6Draft concise, defensible rationales for your rankings, citing specific turns where issues or strengths appear.
- 7Pull and verify Debug Info to confirm the model actually used your chat summaries and connected data sources.
- 8Delete evaluation conversations right after scoring so they don't contaminate your own chat history.
Requirements10
- 1Read and write Polish at a near-native level; the project's focus language is Polish.
- 2Use your primary personal Google account, not a test account, and allow access to personal data sources for authentic evaluation.
- 3Maintain full-time availability in your local time zone, with at least 4 hours of overlap with PST, since the team runs 24/7.
- 4Evaluate nuanced and ambiguous AI outputs while judging personalization quality.
- 5Design creative, multi-turn prompts rooted in personal context to stress-test the model.
- 6Spot personalization errors like poor inferences, forced connections, and unnatural integration of personal data.
- 7Compare Side-by-Side (SxS) responses and identify subtle differences in naturalness and overnarrating.
- 8Write clear, structured rationales that reference exact turn numbers and specific conversation details.
- 9Provide constructive feedback and detailed annotations on model behavior.
- 10Hold a BS/BA degree or equivalent experience in a field like linguistics, journalism, computer science, or another analytical discipline; prior work in data annotation or AI quality evaluation is a plus.
Who Should Apply
The ideal candidate is comfortable mixing creative prompt-writing with strict analytical evaluation. They already use their personal Google account for everyday work and are willing to let Gemini, Gmail, Search, and YouTube activity inform test prompts. This role is less suitable for people who prefer structured, scripted tasks or who hesitate to share personal data with an AI system. Candidates often score low when their rationales are vague instead of referencing specific turn numbers, or when they miss the 24-hour window to complete the assessment. Strong Polish writing alone won't carry the application; evaluators need to see evidence-backed reasoning.
Salary Insight
Paid $20 per hour for a 1-month contractor engagement, with a commitment of at least 4 hours per day and up to 40 hours per week.
Location
Required Skills
Application Tip
When completing the assessment, treat it like a real SxS ranking task: cite specific turn numbers for both strengths and flaws, and keep your rationale concise while still defending the ranking.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Turing
VerifiedAI Quality Analyst (Personalization) - Russian
You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

Turing
VerifiedAI Quality Analyst - English
This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Turing
VerifiedAI Quality Analyst (Personalization) - Turkish
In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

Turing
VerifiedAI Quality Analyst - English
You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Turing
VerifiedAI Quality Analyst (Personalization) - Spanish
Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

Turing
VerifiedAI Quality Analyst (Personalization) - German
An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

