Turing
TuringVerified listing
Remote

AI Quality Analyst - Portuguese (Portugal)

15–15/hr
Remote
Posted August 23, 2026
contract

Overview

Turing hires an AI Quality Analyst to evaluate a new Gemini personalization feature that pulls from your past conversations, Gmail, Google Search, and YouTube activity. You will write multi-turn prompts in Portuguese from your own experiences and judge how well the model responds. Each review weighs personalization quality across Grounding, Integration, and Helpfulness. The contractor role runs for 3 months, pays $15 per hour, and asks for 4 hours of daily overlap with PST.

What You'll Do8

  • 1Design and run multi-turn conversation prompts of 1-5 turns that push the AI to use your personal information and experiences.
  • 2Judge each model response against your original intent and decide whether the personalization matches your starting prompt.
  • 3Check responses for Grounding problems, verifying that claims about you have supporting evidence rather than flawed inferences or hallucinations.
  • 4Assess Integration quality to see whether the model weaves personal data naturally into the reply without robotic overnarrating.
  • 5Compare two model responses side-by-side and stack-rank them based on which is more helpful, easy to use, and enjoyable.
  • 6Write clear, defensible rationales for your rankings and cite the specific turns where issues or strengths appear.
  • 7Extract and verify the model's Debug Info to confirm it used the right chat summaries and data sources.
  • 8Delete evaluation conversations after each task to keep your chat history clean and prevent data leakage.

Requirements10

  • 1High-level reading and writing proficiency in Portuguese (Portugal).
  • 2Willingness to use your primary personal Google account and enable personal data sources for authentic evaluations.
  • 3Full-time availability in your local time zone with 4 hours of overlap with PST each day.
  • 4Strong analytical skills to evaluate nuanced and ambiguous AI responses with a focus on personalization.
  • 5Experience crafting creative, multi-turn starting prompts based on personal context to stress-test model behavior.
  • 6Understanding of personalization concepts, including how to identify incorrect personalization, poor inferences, and forced connections.
  • 7Sharp attention to detail when reviewing Side-by-Side (SxS) responses to catch subtle differences in naturalness and overnarrating.
  • 8Clear written communication that produces structured ranking rationales referencing specific turn numbers.
  • 9BS/BA degree or equivalent experience in a relevant analytical field.
  • 10Prior data annotation, AI quality evaluation, or content moderation experience gives you an edge.

Who Should Apply

The right candidate has strong Portuguese reading and writing skills and a habit of noticing small differences in AI responses. This person can invent personal, multi-turn prompts and decide whether the model used their own data in a useful way. The role works poorly for people who dislike using their personal Google account or who cannot maintain 4 hours of daily overlap with PST. Common rejection reasons include prompt designs that stay too generic and rationales that skip specific turn numbers. Candidates also lose the opportunity by missing the 24-hour assessment window.

Salary Insight

Offered rate is $15 per hour.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

portuguesegeminigmailgoogle searchyoutubeprompt engineeringmulti-turn conversationside-by-side evaluationgroundingpersonalizationdata annotationai quality evaluationcontent moderationdebug infostack-rankingdomain-specific languages

Application Tip

In the assessment, build a sample multi-turn prompt around a real personal scenario, such as planning a weekend trip with details from your Gmail and Google Search history. Then in your written rationale, point to the exact turn where the model either used those details well or made a false inference. That demonstrates both prompt engineering and your ability to write precise, turn-specific justifications.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Turing

Turing

24d agoRemotecontract

AI Quality Analyst (Personalization) - Portuguese

You will evaluate a new personalization feature for Gemini by testing how it draws on past conversations, Gmail, Google Search, and YouTube activity. You will design multi-turn prompts from your own life and judge whether the model's responses are grounded, integrated, and genuinely helpful. Your daily routine involves comparing two model answers side-by-side and writing rationales that reference specific turns. The role demands Portuguese fluency, a willingness to use your personal Google account, and full-time availability aligned to global coverage.

15–15/hr
GeminiGmailSearch+11 more
Turing

Turing

24d agoRemotecontract

AI Quality Analyst - English

This role puts you inside Gemini's personalization quality loop. You will design multi-turn prompts that draw on your own Google activity, including Gmail, Search, and YouTube history, then judge whether the model's responses are grounded, integrated, and helpful. The work involves side-by-side SxS evaluations, writing rationales that reference specific turns, and verifying debug info to confirm the model used your data correctly. This 3-month contractor engagement requires at least 4 hours per day with a 4-hour overlap with PST.

Competitive salary
GeminiGmailSearch+10 more
Turing

Turing

16d agoRemotecontract

AI Quality Analyst (Personalization) - Polish

Evaluate a new Gemini personalization feature by testing how well the model draws on past conversations, Gmail, Google Search, and YouTube activity to tailor responses. You will design multi-turn prompts from your own personal experiences, then score outputs on Grounding, Integration, and Helpfulness. The position is a contractor role that requires Polish reading and writing proficiency and at least 4 hours of daily overlap with Pacific Time. Work happens remotely on your own device, and every evaluation conversation must be deleted afterward to keep your personal history clean.

20–20/hr
GeminiGmailSearch+12 more
Turing

Turing

20d agoRemotecontract

AI Quality Analyst (Personalization) - Spanish

Evaluators in this role test Gemini's new personalization feature by measuring how well the model uses past conversations, Gmail, Google Search, and YouTube activity to tailor responses. The work blends creative prompt writing with structured quality review. Each evaluation prompt starts from the evaluator's own experiences and runs 1-5 turns. You will score responses on Grounding, Integration, and Helpfulness, then compare them side-by-side and justify each ranking. The project centers on Spanish, and you must use your personal Google account for authentic data access.

15–15/hr
SpanishGeminiGmail+12 more
Turing

Turing

12d agoRemotecontract

AI Quality Analyst (Personalization) - Vietnamese

This role puts you inside the evaluation loop for a new personalization feature in Gemini. You will design multi-turn prompts (1-5 turns) that draw on your own Google account data, including past chats, Gmail, Search, and YouTube history. Your core job is to judge whether the model's personalized responses are relevant, grounded in real evidence, and natural, using dimensions like Grounding and Integration. Work is conducted in Vietnamese, comparing two model outputs side-by-side and writing rationales that cite exact conversation turns. The engagement lasts 3 months and pays $15 per hour, with a daily commitment that includes PST overlap.

15–15/hr
VietnameseGeminiGmail+8 more
Turing

Turing

22d agoRemotecontract

AI Quality Analyst (Personalization) - Russian

You will evaluate how well Gemini uses personal data from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. The work draws on your own experiences to craft multi-turn prompts, then rates the model on grounding, integration, and helpfulness. You will compare side-by-side outputs, write structured rationales, and verify that the model pulled from the right data sources. This is a remote contractor role requiring Russian fluency and daily overlap with PST.

15–15/hr
RussianGeminiGmail+12 more