WebOrganizer
收藏资源简介:
WebOrganizer是一个旨在组织和优化预训练数据的数据集。该数据集由普林斯顿语言与智能实验室和艾伦人工智能研究所共同创建,通过对CommonCrawl语料库中的网页内容进行分类,将其划分为24个主题和格式领域。数据集涵盖了广泛的互联网内容,大小达到数万亿个tokens,为研究者和开发者提供了一个深入了解和优化预训练数据的工具。WebOrganizer通过构建丰富的二维数据结构,允许灵活地上采样或下采样领域,为数据标注和下游任务优化提供了丰富的可能性。
WebOrganizer is a dataset dedicated to organizing and optimizing pre-training data. It was co-created by the Princeton Language and Intelligence Lab and the Allen Institute for AI. By classifying web content from the CommonCrawl corpus, it divides the content into 24 thematic and format domains. The dataset covers a wide range of Internet content, with a scale of trillions of tokens, providing researchers and developers with a tool for in-depth understanding and optimization of pre-training data. WebOrganizer constructs a rich two-dimensional data structure, enabling flexible upsampling or downsampling of domains, which offers abundant possibilities for data annotation and downstream task optimization.




