
PDF Annotation & Transcription Experts – Bengali | $12.68/hr Remote
Overview
Remote contributors in India, the US, Canada, and Western Europe are wanted for a Bengali document annotation and transcription project. The task is to take real public PDFs and produce a complete structural map of each page, then transcribe every text region character-for-character in Bengali script. The data deliberately includes the layouts parsing models choke on: handwriting, dense columns, tables, and mixed-script pages from newspapers, exams, and everyday documents. Everything is manual, and every task gets rechecked by a second Bengali expert.
What You'll Do7
- 1Locate a publicly accessible Bengali PDF that fits a given document type and contains at least one image, table, diagram, or handwritten element, and record its source.
- 2Mark every meaningful region on the page, such as title, headings, paragraphs, lists, tables, figures, captions, formulas, question fields, and answer fields, and classify each one.
- 3Assign each region a reading-order index that shows how a real person reads the page.
- 4Link each region to its parent figure or table using a component identifier.
- 5Transcribe every text region exactly as it appears, including handwriting, and flag any area where the source is simply unreadable.
- 6Add page metadata for language, document type, source URL, page dimensions, and flags for tables, formulas, and handwriting.
- 7Review another contributor's completed tasks end to end; experienced annotators take on review duties.
Requirements6
- 1Native-level fluency in Bengali, including reading and writing Bengali script with all diacritics and conjuncts.
- 2Demonstrated experience with document-based work: annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
- 3Character-level accuracy; a missing diacritic is a defect, so the work demands careful attention to detail.
- 4Ability to apply a fixed taxonomy consistently across hundreds of pages instead of improvising.
- 5Comfort with unusual layouts: multi-column newspapers, exam sheets, handwritten forms.
- 6Familiarity with Unicode normalization, Bengali input methods, or OCR correction (nice to have).
Who Should Apply
The right fit is a native Bengali speaker with a background in document annotation, transcription, proofreading, or language QA, and a habit of following an annotation taxonomy exactly. If you're not comfortable with handwritten text, old newsprint, or complex tables, or if you'd rather paraphrase than transcribe letter-for-letter, this role will frustrate you. Rejection most often happens when applicants show no proof of working with Bengali script, or when they treat the source loosely, for example by transliterating into Latin script. It also happens when people miss legibility flags and then guess at unreadable content.
Salary Insight
The listed rate is $12.68/hr for this remote contract work. Because this is an hourly engagement with no benefits, confirm the expected weekly hours and whether the rate also covers training time and review tasks. Given the location pool includes the US and Canada, this rate sits on the lower end, so it makes sense to check volume before applying.
Location
Required Skills
Application Tip
Submit a sample transcription of a messy Bengali PDF, ideally a newspaper page with columns or a handwritten form, and list the document tools you use for viewing and marking regions. Mention any past annotation, proofreading, or localization work in terms of page counts or projects, and state that you can handle Unicode-based Bengali input.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedPDF Annotation & Transcription Experts – Gujarati
This role builds training data for document AI in Gujarati, one of the scripts parsing models stumble on. You find real PDF pages from sources such as newspapers, textbooks, exam papers, forms, and menus, then annotate every meaningful region, set reading order, and transcribe text verbatim in Gujarati script. Humans author every annotation, and the project does not use model-generated output. Every task gets a complete second review from another Gujarati expert, so sustained attention to detail matters more than raw speed. The position is remote and open only to contributors in India, the United States, Canada, or Western Europe.

Mercor
VerifiedPDF Annotation & Transcription Experts – Odia
Odia is one of the languages that document AI handles poorly, and this project builds training data to close that gap. Each task starts with a public PDF page, often something with handwriting, dense columns, or tables, and ends with a structured component map plus a word-for-word transcription in Odia script. Pages come from newspapers, exam papers, forms, brochures, menus, and other everyday sources, so the corpus reflects how Odia documents actually look. Annotators decide how regions are bounded, typed, and ordered, and a second Odia expert reviews every submitted page.

Mercor
VerifiedPDF Annotation & Transcription Experts – Malayalam
Every page in this project starts as a real Malayalam PDF, often a newspaper spread, textbook page, or exam sheet with handwritten parts and multi-column layouts. You turn each page into a complete structure map: every meaningful region gets a component label and a reading-order index, and every text region gets a faithful transcription in Malayalam script. That output becomes training data for vision-language models that still stumble on Indic scripts. The whole pipeline is human-authored, with no parsing models generating labels or text. Accuracy is judged at character level, so one wrong diacritic counts as a defect.

Mercor
VerifiedPDF Annotation & Transcription Experts – Korean
This remote role feeds document AI models with the Korean PDF pages that typical parsers do not handle well. You will source one public PDF in an assigned document category, then annotate its full structure: each title, heading, paragraph, table, figure, caption, formula, and answer field gets a component type and a reading-order index. With structure in place, you transcribe every text region character for character in Hangul, including handwritten notes and hanja. A second Korean expert reviews every delivered page, and mistakes trigger a redo, so exactness matters more than volume.

Mercor
VerifiedPDF Annotation & Transcription Experts – Telugu
Every Telugu PDF page you process turns into structured training data for document AI. The project targets the material parsing models handle worst, including handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages, and it covers Telugu plus four other Indic scripts, Japanese, and Korean. You will find real newspapers, textbooks, exams, flyers, forms, manuals, menus, brochures, notices, and worksheets, then tag each region, set reading order, and transcribe text in Telugu script. All output is human-authored, so your character-level precision shapes the final dataset. Remote work pays $12.68 per hour and is open to contributors based in India, the United States, Canada, or Western Europe.

Micro1
VerifiedBengali Language Expert
This remote contractor role puts your Bengali language expertise to work training next-generation AI systems. You'll help AI models learn to understand, generate, and process Bengali accurately by reviewing, correcting, and creating high-quality language data. No prior AI experience is required—your deep knowledge of Bengali is what matters most.

