Zyda
收藏资源简介:
Zyda是由Zyphra创建的一个包含1.3万亿tokens的大型语言模型预训练数据集。该数据集整合了多个高质量的开源数据集,通过严格的过滤和去重过程,确保数据质量。Zyda不仅在性能上超越了其他开源数据集,如Dolma和RefinedWeb,还显著提高了Pythia系列模型的表现。数据集的创建过程包括综合多个数据源、应用先进的过滤和去重技术。Zyda的应用领域广泛,主要用于提升大型语言模型的训练效果,解决模型在资源分配和性能优化方面的问题。
Zyda is a large language model pre-training dataset with 1.3 trillion tokens, developed by Zyphra. This dataset integrates multiple high-quality open-source datasets, and employs strict filtering and deduplication procedures to guarantee data quality. Zyda not only outperforms other open-source datasets such as Dolma and RefinedWeb, but also significantly boosts the performance of the Pythia series of models. The creation process of Zyda involves integrating multiple data sources and applying advanced filtering and deduplication technologies. The dataset has a wide range of application scenarios, and is primarily used to enhance the training efficacy of large language models and resolve issues concerning model resource allocation and performance optimization.

- 1Zyda: A 1.3T Dataset for Open Language ModelingZyphra · 2024年



