MULTILEGALPILE
收藏资源简介:
MULTILEGALPILE是一个包含24种语言和17个司法管辖区的689GB大规模多语言法律语料库。该数据集由伯尔尼大学等机构创建,涵盖了多种法律数据源,支持在公平使用原则下预训练NLP模型。数据集主要采用宽松许可,适用于法律领域的文本分析和理解。MULTILEGALPILE的应用领域包括法律文本的预训练模型开发,旨在提高法律语言处理的准确性和效率,特别是在多语言和跨司法管辖区的法律文本分析中。
MULTILEGALPILE is a 689 GB large-scale multilingual legal corpus encompassing 24 languages and 17 jurisdictions. Developed by institutions including the University of Bern, this dataset covers diverse legal data sources and supports pre-training of NLP models under the fair use principle. It primarily adopts permissive licenses and is applicable to text analysis and understanding tasks in the legal domain. The application scenarios of MULTILEGALPILE include the development of pre-trained models for legal texts, aiming to improve the accuracy and efficiency of legal language processing, especially in multilingual and cross-jurisdictional legal text analysis.




