
PDF Annotation & Transcription Experts – Gujarati | $12.68/hr Remote
Overview
This role builds training data for document AI in Gujarati, one of the scripts parsing models stumble on. You find real PDF pages from sources such as newspapers, textbooks, exam papers, forms, and menus, then annotate every meaningful region, set reading order, and transcribe text verbatim in Gujarati script. Humans author every annotation, and the project does not use model-generated output. Every task gets a complete second review from another Gujarati expert, so sustained attention to detail matters more than raw speed. The position is remote and open only to contributors in India, the United States, Canada, or Western Europe.
What You'll Do8
- 1Locate a publicly accessible Gujarati PDF in an assigned document type, confirm it has at least one multimodal element such as a table, figure, diagram, or handwriting, and record the source.
- 2Identify every meaningful region on the page and draw bounds around components: title, section headings, paragraphs, lists, tables, figures, captions, formulas, questions, answer fields, and similar items.
- 3Assign each bounded region a component type and a reading-order index that follows the page's actual reading sequence.
- 4For any region that is part of a figure or table, attach it to that parent component with a parent component identifier.
- 5Transcribe all text exactly as printed or handwritten in Gujarati script; mark any region that is not legible instead of attempting a guess.
- 6Record page metadata such as language, document type, source, page dimensions, and flags for tables, formulas, and handwriting.
- 7Review a colleague's completed task end to end, checking region boundaries, component types, reading order, and transcription accuracy.
- 8As you gain experience, serve as the second-expert reviewer for new annotations.
Requirements6
- 1Native fluency in Gujarati with full command of the Gujarati script, diacritics, and conjunct forms.
- 2Document-focused professional background in annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
- 3Consistent application of a shared component taxonomy across hundreds of pages, without improvising rules per document.
- 4Comfort with dense and unpredictable layouts, including multi-column newspapers, exam papers, and handwritten forms.
- 5Precision at character level; a single wrong diacritic is a defect you catch before submission.
- 6Bonus if you have prior work in regional-language AI data, MTPE, bilingual QA, OCR correction, typesetting, digitization, or Unicode normalization.
Who Should Apply
The person who thrives here treats Gujarati script as a craft and reads page structure with intention. If you are a native Gujarati speaker with professional document experience in transcription, proofreading, or localization, you will fit this role. This role is less suited for people who want high-volume speed or who prefer transliteration to writing in the original script. Candidates often lose the role on sample tasks that contain one diacritic error, or when they cannot explain how reading order should move across a multi-column spread. Another quick reject is submitting PDFs that look identical to existing corpus templates, because the project needs layout diversity.
Salary Insight
The pay is $12.68/hr. Accuracy is the primary review metric, and handling time comes second, so the project rewards deliberate work over throughput. At this rate, the role sits in the usual band for remote annotation tasks that require a specialist language like Gujarati.
Location
Required Skills
Application Tip
Attach a one-page sample that shows your annotation process on a real Gujarati PDF: component boxes, types, and reading order. Note any work you have done with Unicode normalization, Gujarati input methods, OCR correction, or handwritten text, since those exact specialties make your profile stand out.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedPDF Annotation & Transcription Experts – Bengali
Remote contributors in India, the US, Canada, and Western Europe are wanted for a Bengali document annotation and transcription project. The task is to take real public PDFs and produce a complete structural map of each page, then transcribe every text region character-for-character in Bengali script. The data deliberately includes the layouts parsing models choke on: handwriting, dense columns, tables, and mixed-script pages from newspapers, exams, and everyday documents. Everything is manual, and every task gets rechecked by a second Bengali expert.

Mercor
VerifiedPDF Annotation & Transcription Experts – Odia
Odia is one of the languages that document AI handles poorly, and this project builds training data to close that gap. Each task starts with a public PDF page, often something with handwriting, dense columns, or tables, and ends with a structured component map plus a word-for-word transcription in Odia script. Pages come from newspapers, exam papers, forms, brochures, menus, and other everyday sources, so the corpus reflects how Odia documents actually look. Annotators decide how regions are bounded, typed, and ordered, and a second Odia expert reviews every submitted page.

Mercor
VerifiedPDF Annotation & Transcription Experts – Malayalam
Every page in this project starts as a real Malayalam PDF, often a newspaper spread, textbook page, or exam sheet with handwritten parts and multi-column layouts. You turn each page into a complete structure map: every meaningful region gets a component label and a reading-order index, and every text region gets a faithful transcription in Malayalam script. That output becomes training data for vision-language models that still stumble on Indic scripts. The whole pipeline is human-authored, with no parsing models generating labels or text. Accuracy is judged at character level, so one wrong diacritic counts as a defect.

Mercor
VerifiedPDF Annotation & Transcription Experts – Telugu
Every Telugu PDF page you process turns into structured training data for document AI. The project targets the material parsing models handle worst, including handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages, and it covers Telugu plus four other Indic scripts, Japanese, and Korean. You will find real newspapers, textbooks, exams, flyers, forms, manuals, menus, brochures, notices, and worksheets, then tag each region, set reading order, and transcribe text in Telugu script. All output is human-authored, so your character-level precision shapes the final dataset. Remote work pays $12.68 per hour and is open to contributors based in India, the United States, Canada, or Western Europe.

Mercor
VerifiedPDF Annotation & Transcription Experts – Japanese
Document understanding models often break down on scripts with underrepresented training data, so this project supplies that data for Japanese within a broader set that includes Korean and five Indic scripts. You will take real public PDFs and turn each page into a structural map that captures every meaningful region in the correct reading order. The corpus focuses on the hard material: handwriting, vertical text, multi-column newspapers, tables, diagrams, and mixed-script documents pulled from textbooks, exams, flyers, forms, manuals, menus, notices, and worksheets. A person assigns every component type and reading-order index, and a person writes every transcription in Japanese script. A second Japanese expert reviews each finished task end to end.

Mercor
VerifiedPDF Annotation & Transcription Experts – Korean
This remote role feeds document AI models with the Korean PDF pages that typical parsers do not handle well. You will source one public PDF in an assigned document category, then annotate its full structure: each title, heading, paragraph, table, figure, caption, formula, and answer field gets a component type and a reading-order index. With structure in place, you transcribe every text region character for character in Hangul, including handwritten notes and hanja. A second Korean expert reviews every delivered page, and mistakes trigger a redo, so exactness matters more than volume.

