scilons/SciLaD-all-json-v1
收藏资源简介:
SciLaD(科学语言数据集)是一个新颖的大规模科学语言数据集,完全基于开源框架和公开可用的数据源构建。它包含两个主要部分:一个经过精心策划的英文分片,涵盖超过1000万篇科学出版物;以及一个多语言、未经过滤的TEI XML分片,包含超过3500万篇出版物。此外,该数据集还提供了用于生成SciLaD的可扩展流水线。在本仓库中,我们分享了完整的多语言JSON版本数据集,该版本通过使用grobid-client-python从TEI-XML转换而来,并基于Grobid 0.8.1处理。我们计划在每次新Grobid发布时更新数据集。数据集由Scilons Project策划,支持多语言处理。
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. In this repository, we share the full, multilanguage JSON version of the dataset, transformed from TEI-XML using grobid-client-python and processed with Grobid 0.8.1, with plans to release updates at each new Grobid release. The dataset is curated by Scilons Project and supports multilanguage NLP.



