# Japanese Documents Dataset (PDF) *This dataset contains a curated collection of Japanese-language documents in PDF format. The corpus includes textbooks, research papers, news articles, public-doma
# olmOCR-mix-1025 olmOCR-mix-1025 is a dataset of ~270,000 PDF pages which have been OCRed into plain-text in a natural reading order using gpt-4.1 and a special prompting strategy that preserves any