INTRODUCTION: This dataset is based on 2017 Sci-Hub logs for the Russian Federation, enriched by CrossRef metadata and Unpaywall data. The findings are published separately. ================
A refined version of ArXiv dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Langua