Turing
TuringVerified listing
Remote

AI Quality Analyst (Personalization) - German

15–15/hr
Remote
Posted August 17, 2026
contract

Overview

An AI Quality Analyst on this project evaluates Gemini's new personalization feature by testing how well it uses data from past conversations, Gmail, Google Search, and YouTube activity. Analysts design multi-turn prompts grounded in their own experiences, then judge the responses for Grounding, Integration, and Helpfulness. The work is remote, contract-based, and focused on German-language quality. German reading and writing skills at a professional level are required, along with agreement to connect a primary personal Google account.

What You'll Do8

  • 1Design and run multi-turn conversational prompts, usually 1-5 turns, that require the model to use your personal information and experiences.
  • 2Rank two model responses side by side to decide which is more helpful, easy to use, and enjoyable.
  • 3Write clear, defensible rationales for your rankings and cite the exact turn numbers where strengths or issues appear.
  • 4Check whether personalized claims about you are grounded in evidence, and flag hallucinations or flawed inferences.
  • 5Judge whether personal data is woven into the response in a natural way and without robotic overnarrating.
  • 6Pull and verify Debug Info from the model to confirm that chat summaries and data sources were used correctly.
  • 7Delete each evaluation conversation after use to keep your future chat history clean.
  • 8Provide constructive feedback and detailed annotations to support model quality improvements.

Requirements11

  • 1Professional-level German reading and writing, since German is the focus language for this project.
  • 2Willingness to use your primary personal Google account and enable personal data sources for authentic evaluation.
  • 3Full-time availability, including at least 4 hours per day and up to 40 hours per week, with a 4-hour overlap with PST.
  • 4Strong analytical skills for assessing nuanced and ambiguous AI responses.
  • 5Experience designing creative, multi-turn prompts that draw on personal context.
  • 6Ability to spot incorrect personalization, poor inferences, and forced connections.
  • 7Attention to detail for detecting subtle differences in Side-by-Side model responses, including naturalness and overnarrating.
  • 8Excellent written communication, with the ability to write clear, structured rationales and reference specific turn numbers.
  • 9BS/BA degree or equivalent experience in a relevant analytical field, such as Policy, Law, Ethics, Linguistics, Journalism, or Computer Science.
  • 10Prior data annotation, AI quality evaluation, or content moderation experience is preferred.
  • 11Reliable desktop/laptop and a stable internet connection.

Who Should Apply

Candidates who combine strong German reading and writing skills with genuine interest in how AI personalization behaves will fit well here. The role suits people who can construct creative multi-turn prompts from their own life and then judge the output with a critical eye. The role is less suitable for anyone uncomfortable with connecting their primary personal Google account or working a strict PST-overlap schedule. Common rejection reasons include weak German writing in the assessment and vague rationales that do not reference specific turn numbers. Slow submission of the 24-hour assessment can also take a qualified candidate out of the running.

Salary Insight

The project pays $15 per hour.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

geminigermangmailgoogle searchyoutubeprompt engineeringmulti-turn promptsside-by-side evaluationai quality evaluationdata annotationcontent moderationpersonalizationdomain-specific languages

Application Tip

Before applying, draft a German-language evaluation sample that compares two model responses and references specific turn numbers. That directly shows the written communication and attention to detail this role demands.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

16d agoRemotecontract

AI Quality Analyst (Personalization) - Polish

Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

20–20/hr
GeminiGmailSearch+12 more
Turing

Turing

21d agoRemotecontract

AI Quality Analyst (Personalization) - Turkish

In this contract role, you will evaluate a new personalization feature for Gemini. You will measure how well the model incorporates signals from your past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. You will design multi-turn prompts based on your own experiences and then judge the model's output on dimensions like Grounding, Integration, and Helpfulness. The work demands both creative prompt design and disciplined analytical review of model responses.

15–15/hr
GeminiGmailSearch+14 more
Turing

Turing

20d agoRemotecontract

AI Quality Analyst (Personalization) - Spanish

Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

15–15/hr
SpanishGeminiGmail+12 more
Turing

Turing

23d agoRemotecontract

AI Quality Analyst - English

This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Competitive salary
GeminiGmailSearch+10 more
Turing

Turing

6h agoRemotecontract
Hot

AI Quality Analyst - English

You will judge how well Gemini draws on your personal history, including past chats, Gmail, Google Search, and YouTube activity, to craft relevant replies. The job centers on a new personalization feature, so your own Google account becomes the test bed. Each evaluation asks you to compare two model responses side by side and decide which one feels more natural and useful. You also classify weak Grounding, awkward Integration, and judge the overall Helpfulness of each output. Prompt design is a core part of the work: you create multi-turn scenarios from your own experiences to push the model.

Competitive salary
GeminiPersonalizationPrompt Engineering+15 more
Turing

Turing

22d agoRemotecontract

AI Quality Analyst (Personalization) - Arabic

You will evaluate a new personalization feature in Gemini. You will design prompts that draw on your own life and judge how well the model uses data from past conversations, Gmail, Google Search, and YouTube to tailor responses. Each review weighs Grounding, Integration, and helpfulness. The project focuses on Arabic, so strong reading and writing skills in that language are required. Work is remote and requires a 4-hour overlap with Pacific Time each day.

15–15/hr
ArabicGeminiGmail+10 more