Mercor
MercorVerified listing
Remote

PDF Annotation & Transcription Experts – Japanese | $33.58/hr Remote

33.58–33.58/hr
Remote
Posted September 2, 2026
hourly
114 openings

Overview

Document understanding models often break down on scripts with underrepresented training data, so this project supplies that data for Japanese within a broader set that includes Korean and five Indic scripts. You will take real public PDFs and turn each page into a structural map that captures every meaningful region in the correct reading order. The corpus focuses on the hard material: handwriting, vertical text, multi-column newspapers, tables, diagrams, and mixed-script documents pulled from textbooks, exams, flyers, forms, manuals, menus, notices, and worksheets. A person assigns every component type and reading-order index, and a person writes every transcription in Japanese script. A second Japanese expert reviews each finished task end to end.

What You'll Do6

  • 1Locate a public Japanese PDF in a designated category, confirm it includes at least one multimodal element such as an image, table, diagram, or handwriting, and record where you found the file.
  • 2Label each meaningful region of a page, from titles and section headings to paragraphs, lists, tables, figures, diagrams, captions, formulas, questions, and answer fields, then assign a component type and a reading-order index.
  • 3Attach each labeled region to its parent figure or table using a component identifier so the relationships between page elements stay intact.
  • 4Transcribe every text string in the original Japanese characters, including kanji, hiragana, katakana, furigana, and handwriting, and flag any region that is not legible.
  • 5Capture page-level metadata such as language, document type, source URL, page dimensions, and whether the page contains tables, formulas, or handwriting.
  • 6Check another contributor's complete task end to end and move into review duties once your own work reaches the project's quality bar.

Requirements10

  • 1Native-level Japanese with full command of kanji, hiragana, katakana, furigana, and kanji variant forms.
  • 2Must be based in Japan, South Korea, India, the United States, Canada, or Western Europe.
  • 3Professional experience in translation, interpretation, journalism, transcription, editing, or another document-heavy discipline.
  • 4Character-level accuracy, because a single wrong character turns a completed task into a rejection.
  • 5A systematic method for applying the same region taxonomy across hundreds of different page layouts.
  • 6Comfort with vertical text, multi-column newspaper layouts, exam papers, and handwritten forms.
  • 7Experience with AI training data annotation, labeling, grading, or bilingual evaluation is preferred.
  • 8Experience with MTPE, subtitling, bilingual QA, OCR correction, or post-editing is preferred.
  • 9Knowledge of Unicode normalization, Japanese input methods, full-width and half-width forms, and kanji variants is preferred.
  • 10Ability to deliver transcription in Japanese script rather than romanized text.

Who Should Apply

People already doing document-heavy Japanese work will find this role familiar. Translation, transcription, editorial, and journalism backgrounds line up best when the work has included handwriting or unusual layouts. The role is less suitable for those who prefer clean typed documents, because hard-to-parse layouts are the whole point of the corpus. Candidates score low when they cannot show native kanji and furigana proficiency, or when their experience is limited to conversational translation and general data entry.

Salary Insight

The assignment pays $33.58 per hour. That rate is set for native-level Japanese contributors working as independent contractors, which differs from a salaried local hire. The project tracks handling time, but accuracy remains the first measure of a task that is accepted.

Location

Typeremote
LocationRemote
Eligible countriesJapan, South Korea, India, United States, Canada +18 more
This is a remote position

Required Skills

japanesekanjihiraganakatakanafuriganatranscriptionannotationpdfreading orderdocument layout analysismetadata annotationjapanese input methodsunicode normalizationtypesettingcopy editingproofreadingpost-editingsubtitlingbilingual qaocrai training data

Application Tip

Show a concrete trace of your accuracy: describe one time you spotted a kanji variant error or transcribed ambiguous handwriting, and state whether your task passed a second-expert review on the first pass. This gives reviewers a measurable signal of the character-level discipline the role demands.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

2d agoRemotehourly

PDF Annotation & Transcription Experts – Korean

This remote role feeds document AI models with the Korean PDF pages that typical parsers do not handle well. You will source one public PDF in an assigned document category, then annotate its full structure: each title, heading, paragraph, table, figure, caption, formula, and answer field gets a component type and a reading-order index. With structure in place, you transcribe every text region character for character in Hangul, including handwritten notes and hanja. A second Korean expert reviews every delivered page, and mistakes trigger a redo, so exactness matters more than volume.

33.58–33.58/hr
· 114 openings
KoreanHangulHanja+16 more
Mercor

Mercor

2d agoRemotehourly

PDF Annotation & Transcription Experts – Odia

Odia is one of the languages that document AI handles poorly, and this project builds training data to close that gap. Each task starts with a public PDF page, often something with handwriting, dense columns, or tables, and ends with a structured component map plus a word-for-word transcription in Odia script. Pages come from newspapers, exam papers, forms, brochures, menus, and other everyday sources, so the corpus reflects how Odia documents actually look. Annotators decide how regions are bounded, typed, and ordered, and a second Odia expert reviews every submitted page.

12.68–12.68/hr
· 120 openings
OdiaOdia ScriptAnnotation+12 more
Mercor

Mercor

2d agoRemotehourly

PDF Annotation & Transcription Experts – Telugu

Every Telugu PDF page you process turns into structured training data for document AI. The project targets the material parsing models handle worst, including handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages, and it covers Telugu plus four other Indic scripts, Japanese, and Korean. You will find real newspapers, textbooks, exams, flyers, forms, manuals, menus, brochures, notices, and worksheets, then tag each region, set reading order, and transcribe text in Telugu script. All output is human-authored, so your character-level precision shapes the final dataset. Remote work pays $12.68 per hour and is open to contributors based in India, the United States, Canada, or Western Europe.

12.68–12.68/hr
· 192 openings
TeluguTelugu ScriptTranscription+19 more
Mercor

Mercor

2d agoRemotehourly

PDF Annotation & Transcription Experts – Gujarati

This role builds training data for document AI in Gujarati, one of the scripts parsing models stumble on. You find real PDF pages from sources such as newspapers, textbooks, exam papers, forms, and menus, then annotate every meaningful region, set reading order, and transcribe text verbatim in Gujarati script. Humans author every annotation, and the project does not use model-generated output. Every task gets a complete second review from another Gujarati expert, so sustained attention to detail matters more than raw speed. The position is remote and open only to contributors in India, the United States, Canada, or Western Europe.

12.68–12.68/hr
· 192 openings
GujaratiGujarati ScriptPdf Annotation+15 more
Mercor

Mercor

2d agoRemotehourly

PDF Annotation & Transcription Experts – Malayalam

Every page in this project starts as a real Malayalam PDF, often a newspaper spread, textbook page, or exam sheet with handwritten parts and multi-column layouts. You turn each page into a complete structure map: every meaningful region gets a component label and a reading-order index, and every text region gets a faithful transcription in Malayalam script. That output becomes training data for vision-language models that still stumble on Indic scripts. The whole pipeline is human-authored, with no parsing models generating labels or text. Accuracy is judged at character level, so one wrong diacritic counts as a defect.

12.68–12.68/hr
· 192 openings
MalayalamMalayalam ScriptTranscription+17 more
Mercor

Mercor

2d agoRemotehourly

PDF Annotation & Transcription Experts – Bengali

Remote contributors in India, the US, Canada, and Western Europe are wanted for a Bengali document annotation and transcription project. The task is to take real public PDFs and produce a complete structural map of each page, then transcribe every text region character-for-character in Bengali script. The data deliberately includes the layouts parsing models choke on: handwriting, dense columns, tables, and mixed-script pages from newspapers, exams, and everyday documents. Everything is manual, and every task gets rechecked by a second Bengali expert.

12.68–12.68/hr
· 192 openings
BengaliBengali ScriptTranscription+17 more