M2D2
收藏资源简介:
M2D2是由东京大学和华盛顿大学合作开发的大规模多域语言建模数据集,包含85亿个令牌,覆盖145个领域,主要从维基百科和Semantic Scholar提取。数据集通过维基百科和ArXiv的分类学进行组织,形成两级层次结构,便于研究领域间的关系及其对模型适应性的影响。M2D2特别适用于研究语言模型在不同领域间的转移学习,尤其是在细粒度领域间的适应性。该数据集的开发旨在解决语言模型在多样化数据分布下的性能问题,并探索如何有效适应新领域。
M2D2 is a large-scale multi-domain language modeling dataset co-developed by the University of Tokyo and the University of Washington. It comprises 8.5 billion tokens and covers 145 domains, with data primarily extracted from Wikipedia and Semantic Scholar. The dataset is organized using the taxonomies of Wikipedia and ArXiv to form a two-level hierarchical structure, which facilitates research on the relationships between domains and their impacts on model adaptability. M2D2 is particularly suitable for investigating transfer learning of language models across different domains, especially adaptation between fine-grained domains. This dataset was developed to address the performance issues of language models under diverse data distributions and to explore effective approaches for adapting to new domains.

- 1M2D2: A Massively Multi-domain Language Modeling Dataset东京大学 · 2022年



