Linear Next基准使用了一系列高质量的数据集,包括用于通用语言建模任务的DCLM-pro数据集、涵盖多个教育领域内容的Cosmopedia-v2和Fineweb-edu数据集、用于代码理解和生成任务的The Stack v2数据集、包含数学内容和问题的Finemath数据集,以及专注于逻辑推理和问题解决的Natural Reasoning数据集。
It is somewhat helpful to use well-understood and commonly used standard datasets so that the findings can be quickly evaluated. Nevertheless, most of the corpora datasets are prepared for NLP tasks.
Text files of different size and structure. More precisely, we selected random data from the Gutenberg dataset. This artefact contains five different datasets with random text files (i.e. e-books in