Skip to content
← Back to job listings

Multilingual Data Contributors: PDF Collection for AI Training

jobgether · US

RemoteExternal listingcontract5 days ago

About The Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Multilingual Data Contributors: PDF Collection for AI Training based in the United States.

This remote contract opportunity is designed for multilingual contributors who can help improve AI language technologies through high-quality document <sourcing.You> will identify public, legally usable PDF documents in your designated language and contribute them to an AI research project.Your work will help support the development of more capable text recognition and generation systems across multiple languages.The role is fully asynchronous, allowing you to complete sourcing and submission activities <independently.You> will be responsible for checking document quality, licensing eligibility, and basic metadata before <submission.It> is particularly suited to researchers, data contributors, annotators, language specialists, and detail-oriented online researchers.Fluent reading ability in Telugu, Odia, Gujarati, Malayalam, Japanese, or Korean is required.

Accountabilities

  • Source high-quality public PDF documents in your designated language using online archives, public records, open-source repositories, and other appropriate sources.
  • Verify that each document is legally usable and meets applicable public-domain or open-source licensing requirements.
  • Review documents for relevance, quality, and compliance with the project’s sourcing guidelines.
  • Download and securely upload approved documents through the designated research platform.
  • Provide basic descriptive metadata and other required information for each submitted document.
  • Work independently and asynchronously while following project instructions and quality standards.
  • Respond to review feedback and make adjustments to submissions when required.
  • Maintain careful attention to copyright, licensing, document quality, and submission accuracy throughout the process.

Requirements

  • Fluent reading comprehension in Telugu, Odia, Gujarati, Malayalam, Japanese, or Korean.
  • Comfortable searching for, evaluating, downloading, and organizing digital documents online.
  • Basic understanding of public-domain, open-source, and copyright or licensing principles.
  • Strong attention to detail and the ability to consistently follow document sourcing and quality requirements.
  • Ability to independently research information and make sound judgments about document suitability.
  • Comfortable working asynchronously and managing assigned tasks without close supervision.
  • Experience in research, data annotation, document collection, language-related work, digital archives, or online research is beneficial.
  • Reliable computer and internet connection capable of downloading and uploading PDF files.
  • Ability to provide accurate basic metadata and descriptive information for submitted documents.

Benefits

  • Compensation: $150 per completed task.
  • Fully remote contract opportunity.
  • Flexible, asynchronous work that can be completed independently.
  • Opportunity to contribute to the development and improvement of multilingual AI technologies.
  • Suitable for researchers, language contributors, data annotators, and other individuals looking for project-based AI work.
  • Opportunity to gain practical experience contributing human-generated data to machine learning research.
  • No specific academic degree requirement is stated; relevant research, language, or data experience is valued.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing